06

Twenty activations from a benchmark audit

Evidence · GovernanceEstée Lauder Companies
UX ResearchExperience StrategyProduct & Design LeadershipAI & Insight Systems
The problem

A corporate-mandated UX audit had benchmarked the brand sites against the Baymard research library. What came back was a benchmark, not a plan: 650 guidelines and a pile of performance scores, with no way for a brand team to know which ones applied to them or what to change on Monday.

My role

Led the qualitative insights work that turned the benchmark into an activation playbook. Which twenty guidelines mattered most for our site patterns, what each one meant inside our own CMS, and how we would know whether it worked.

Approach
  • Filtered 650+ guidelines to the twenty with the most impact on our patterns
  • Scored every one for level of effort so teams could sequence rather than stall
  • Audited seven global brands against all twenty and published the matrix
  • Wrote the CMS route for each change, not the principle behind it
Data sources
  • Baymard Institute research library
  • Fullstory behavioral analytics
  • Brand-by-brand site audits
Activations
20
Prioritized from the full guideline library
Brands audited
7
Every brand scored against every activation
Source library
650+
Design guidelines behind 130,000 hours of usability research
Effort tiers
4
Low through high, so a team with no engineering time could still move
Borrowed evidence is cheap and usually wasted. The work was not finding the research. It was making somebody else's research specific enough to act on.
/ Compliance matrix
06 / Baymard activations

Every brand, every activation,
on one page

A column that is red all the way down is a platform problem. A row that is red across is a brand problem. Putting both on one grid is what stopped the conversation from being about whose fault it was.

Cross-sell placement
CTA scope
Description structure
Parent-category links
Mobile list height
Shipping near buy
Aveda
Bobbi Brown
Clinique
Estée Lauder
La Mer
Origins
MAC
Key
● meets the guideline  ·  ◐ partial  ·  ○ in violation
Two failures were universal. No brand linked cross-sells back to their parent category, and every brand had mobile list items taller than half the screen. Those stopped being brand tickets and became platform work.
/ The activation tracker
06 / Baymard activations

Twenty rows. Each one written
as a CMS instruction.

The tracker was the deliverable. Level of effort on the left so a team could sequence, the guideline in the middle, and the actual route through our own tooling on the right. Nobody had to interpret a principle.

Activation tracker, low through medium effort tiers: eight guidelines with level of effort, the insight, and the CMS steps for each
Tiers 1 & 2Eight activations a brand could ship with no engineering allocation at all. Content structure, cross-sell placement, shipping information near add-to-bag.Open full size
Activation tracker, medium through medium-high effort tiers: nine guidelines with level of effort, the insight, and the CMS steps for each
Tier 3Nine activations needing a content decision behind them rather than just a field change. Cross-sell relevance, navigation chunking, product list counts.Open full size
Activation tracker, high effort tier: three guidelines routed to engineering with recommendations for each
Tier 4Three activations that could only be solved in the platform, routed to engineering with the recommendation already written. Overlay behaviour, mobile zoom, product comparison.Open full size
Splitting the list by effort was the decision that made it move. A single ranked list of twenty would have stalled on the first item that needed a developer.
/ Evidence and measurement
06 / Baymard activations

Each activation carried its own
proof and its own success test

Every guideline was written up the same way. What the behaviour was, an example of the pattern working somewhere else, our own version of the failure, and the specific change. Then a defined move from violated to adhered, with behavioural metrics attached before any of it shipped.

Measurement framework: the adherence scale moving from violated high and violated low to adhered high and adhered low, alongside the Fullstory behavioural metrics tracked before and after each change
Success defined firstAdherence graded on a four-point scale, with frustration events, time on site, engagement and bounce tracked in Fullstory before and after. Reported by activation rather than by brand.Open full size
A single activation written up in full: the insight, CMS action items, further items requiring development, a working example, a failing example with a participant quote, and the recommendation
One activation in fullThe low-effort cross-sell placement guideline. Insight, CMS steps, what needs a developer, a pattern that works, a pattern that fails, and a participant explaining why she left the page.Open full size
Defining the measure before the change is what separates an activation from a recommendation. It also meant a brand team could disagree with the priority without disagreeing about the evidence.
/ What it produced
06 / Baymard activations

An audit says where you stand.
This said what to change, in what
order, and how you would know.

Measurement was defined before any of the work started, which is the part audits usually skip.
What to change

Each activation written as a CMS instruction rather than a principle. Move the routine cross-sell above reviews. Break any text block over five lines into headed sections. Put shipping information below add-to-bag.

In what order

Four effort tiers. The low tier was entirely CMS work, so brands with no engineering allocation could still ship against it in the same quarter.

How you would know

Frustration events, rage and dead clicks, time on site, engagement, and bounce, tracked in Fullstory before and after each change and reported by activation rather than by brand.

Somebody else had already done the research. The contribution was making it actionable inside our own tools.