13

When the analytics said win
and the sessions said no

Mixed methods · UsabilityEstée Lauder
UX ResearchExperience StrategyProduct & Design LeadershipAI & Insight Systems
Methods
Unmoderated usability testing Analytics review Severity grading Pilot testing Competitive review
Analytics reviewed first, then sessions run against what the numbers claimed.
Inputs and tools
UserTesting Google Analytics Nielsen Norman heuristics 12 sessions, 2 arms 7-task think-aloud script
I wrote the script and the recommendations. Moderation was shared with a researcher on the product side.
The problem

A mobile-only AR skin diagnostic had just launched in North America. The user scans their face, the tool identifies skin concerns, and it returns a personalized product routine. The launch numbers were excellent and nobody had watched anyone use it. Product wanted to know what to build in v2.

My role

Wrote the study and the task script, shared moderation with a researcher on the product side, and authored the recommendations and the prioritized requirements that went to the product team.

Methodology
  • Analytics reviewed first, then twelve remote sessions run against what the numbers claimed
  • Two pilot sessions, one per device, before any real session ran
  • Ten sessions total, split across mobile and desktop entry
  • Seven-task script, think-aloud, unmoderated on UserTesting
  • Findings graded on a five-level severity scale tied to Nielsen Norman heuristics
Why five per arm

Not a budget compromise. The plan was three rounds of five: find the surface problems, fix them, then test again for the deeper ones. Multiple small rounds change a product; one large round documents it.

Study methodology showing the post-launch position, remote user testing method, platform, and audience composition
/ Insight
13 / Skin diagnostic

The personalization tool was not personalizing, and the metrics could not see it.

What the numbers said
  • Three quarters of people who started the scan finished it
  • Conversion roughly double the mobile baseline
  • Time on site around five times the site average
  • Revenue per user roughly two and a half times the mobile average
What the sessions said
  • Every participant in the mobile arm received an identical recommended routine
  • The same defect appeared in the separate quiz-based finder, returning the same results to everyone
  • Participants could not work out how to trigger the capture, because there is no button
  • Five recommended products read as an upsell rather than a routine
A diagnostic that returns the same answer to everyone still converts, right up until someone compares two results.
/ Reading the analytics honestly
13 / Skin diagnostic

The lift was real. The causal
story attached to it was not.

These figures were being read internally as evidence that the tool drove conversion. It is worth saying plainly what they can and cannot support, because the gap between those two things is where the study earned its keep.

Completion
75%
Of people who launched the scan and finished it
Conversion
2x
Against the mobile baseline, mobile only
Time on site
5x
Against the site average session
Revenue per user
2.5x
Against the mobile average
The selection problem

People who choose to scan their face for a skincare routine are already further along than the average visitor. Comparing them to the site baseline compares intent, not the tool. The honest read is that the diagnostic identifies high-intent shoppers extremely well. Whether it creates them is a different question and this data cannot answer it.

What it would take to know

A holdout: the same high-intent segment, half routed to the diagnostic and half to the existing quiz, measured on the same window. Without that, every one of these multiples is a description of who showed up.

Analytics summary showing completion rate, conversion rate, time on site, and revenue per user against baselines
The analytics pull that preceded the sessions. Reviewing it first is what made the identical-routine finding land as a defect rather than a curiosity: the numbers had already convinced everyone the tool was working.
/ How it was run
13 / Skin diagnostic

Twelve sessions, seven tasks,
five severity levels

The tool was mobile-only, but people reach a mobile-only tool from a desktop browser too, so the sample was split to see both. The findings below come from the mobile arm unless stated otherwise, which is where the substantive failures were.

Sessions
12
Two pilots plus ten tests
Arms
2
Mobile and desktop, five participants each
Tasks
7
Expectation, capture, results, recommendations, surrounding content
Severity levels
5
Positive through catastrophic, each with a fix threshold
The seven-task test script, opening with an expectation-setting question before the tool is launched
The script. Task one asks what people expect the tool to do before they touch it, which is what produced the finding that they wanted to be asked about their skin rather than photographed.
Five-level severity ranking from positive through cosmetic, minor, major, and catastrophic
The severity scale, with each level carrying its own fix threshold. Grading this way is what let a product team sequence the backlog without renegotiating every item.
Study objectives and the nine specific things the sessions were watching for
Objectives and the nine watch-fors. Worth noting honestly: what this deck labelled a hypothesis was an objective. The Baymard search audit is where I was writing genuinely falsifiable hypotheses in the same period.
Attribution

This was a two-person study. I wrote the script and the recommendations; session moderation was shared with a researcher embedded on the product side. Participant imagery from the sessions is deliberately excluded from this page.

Why the pilots mattered

Running one pilot per device caught script problems before ten sessions were spent on them. It is the cheapest quality control in unmoderated research and the step most often skipped.

/ Where it broke
13 / Skin diagnostic

The camera step was the wall

Everything upstream of the scan was a minor issue. The scan itself was major, and it was major for a reason that only shows up on a phone: the interface had removed the one control people were looking for.

The tool entry point on mobile with the get started call to action
Entry point, graded minor. People expected to be asked about their skin type and current routine first, and expected a prompt to take a photo. Nothing on the screen tells them a face scan is what happens next.
The selfie tips screen listing three preparation instructions before the scan
Selfie tips, graded minor. The instruction to use good lighting was the sticking point: near a window, overhead, bathroom. Guidance that assumes shared understanding of a term nobody defines.
Major · No capture control

There is no shutter button. The photo fires on a countdown once the face is centered closely enough. Participants tapped the screen repeatedly trying to find the trigger, then recentered, then tapped again.

Major · Invisible readiness state

Three indicators are supposed to turn green to signal the system is ready. In session they were not noticeable. The affordance existed and did not communicate.

Major · Wrong camera mode

At least one participant expected to use portrait mode rather than the front-facing camera, which is a reasonable read of what a beauty brand would ask for and produces an unusable input.

Distance, framing, and lighting were all things the interface knew and would not say.
/ The results problem
13 / Skin diagnostic

People believed the diagnosis
and rejected the prescription.

The analysis screen was the one part of the flow graded positive. Participants found the concerns credible, liked that each one named a cause in a single sentence, and engaged with the severity scales. Then the routine arrived and the trust went with it.

01 Five products is not a routine

Participants read a five-product recommendation as a sales target rather than advice. Several said outright they would not replace an entire regimen at once, and those with sensitive skin named a real risk in doing so. The recommendation that followed was two or three products, with samples offered as the low-commitment entry.

02 The routine did not read as ordered

The recommended products and the daily regimen were separate sections, so nothing connected step one to the cleanser. The page looked like a product grid rather than a sequence. Merging them was the structural fix.

03 Nothing to do next

No add-to-bag on the results, and regimen images that did not link through to the products they showed. The tool ended on a wall.

The finding that outranked all of them

Every participant in the mobile arm received the same recommended routine. Five different faces, five different stated concerns, one identical output. The same defect was present in the quiz-based skincare finder, which also returned identical results regardless of input.

Two personalization surfaces, both shipped, both returning a constant. Neither had been caught because both were converting.

Accuracy has a floor
One participant's freckles were classified as dark spots. Credibility in a diagnostic is not a feature, it is the whole product.
Two tools, one job, no distinction
Participants could not tell the scan-based tool from the quiz-based one, and expected the scan to include the quiz's questions.
The ingredient glossary existed and the product pages did not link to it, so the education people wanted was one tap away and unreachable.
The scan was believable. What it recommended was not, and that order of failure is the expensive one.
/ What went to product
13 / Skin diagnostic

Findings sorted into a backlog,
not a list of complaints

Every observation was attached to the screen it came from, graded for severity, and written as a change someone could pick up. Defects were flagged separately from design recommendations, because they have different owners and different urgency.

Prioritized requirements grouped by screen: entry point, selfie tips, capture, and recommended routine
The requirements as handed over, grouped by screen so each card maps to an owner. Defects are marked as bugs and separated from the experience recommendations around them.
Prioritized requirements for the daily regimen and the supporting content sections
The second half: sequencing the regimen, linking imagery to product pages, and reordering the supporting content so guidance comes before additional merchandising.
A competitor's four-step how it works explainer for the same underlying AR tool
A competitor running the same underlying AR technology explains it in four steps with images before anyone opens the camera. Pulling this in gave the copy recommendation a working precedent instead of a preference, which is what got it built.
/ What I would do differently
13 / Skin diagnostic

Three honest limits

This should have run pre-launch

The study sits in the post-launch column of its own roadmap. A defect returning identical results to every user is a pre-launch QA finding, and it took a usability study months after release to surface it. The lesson was not about the method, it was about when the method gets invited.

Only round one ever ran

The design called for three iterative rounds. Round one found the surface problems and the deeper questions never got tested, so what is here is the first pass of a plan that was built to go further.

The accuracy question stayed open

Participants raised whether lighting was driving their results, and freckles being read as dark spots suggests they were onto something. That is a computer vision validation question, not a usability one, and it needed a study of its own that I would scope now rather than hand back as a note.