Rebuilding an association site on eight weeks of evidence
A global medical association with content spread across six domains and nine conference sub-brands. The site was not broken. It was built for one audience, and everyone else was left to work it out. Thirteen distinct audiences surfaced in discovery. Six of them had no dedicated journey.
Project-lead UX researcher across discovery, analytics synthesis, survey design, interviews, and information architecture, on a timeline compressed from six months to two.
- Coded all 51 proposed navigation destinations by the audience each one actually serves
- Synthesized 12 months of behavioral data as a screening layer
- Ran a global survey, then weighted interview recruiting toward the voices it missed
- GA4 and internal site search
- Lyssna survey, 107 responses across 31 countries
- Moderated interviews
I was going to leave before
the research finished.
Interviews would not close before my contract ended, and the architecture was heading into wireframes with an agency and client team who had never sat in a session.
So every artifact had to answer a question I would not be there to answer: how confident are we in this, and what would change our mind?
Three evidence streams,
each held to one job
None of them got to do another one's. This went on the first slide of the analytics readout, before a single finding.
Behavioral data
What people do, where they get stuck, and what they search for that is not there.
Why they did it. Motivation and context do not show up in a dashboard.
Global survey
What people say they need, ranked and weighted across roles.
What people actually do. It also misses the least engaged, who do not answer surveys.
Moderated interviews
Why. The context, the workarounds, and what the failure actually looks like.
Prevalence. An interview finding is never a percentage.
Three layers, so confidence travels with the decision
We shipped the architecture as a working hypothesis rather than a recommendation, and made the confidence behind every line of it visible on the page.
Verdict before recommendation

Keep protects what already works from being redesigned anyway. Verify names the open question and what would answer it, so silence never reads as a yes. Change is reserved for where behavior alone is enough to act.
IF / THEN flags

Every unvalidated assumption shipped with its own test. A members-only tool with a brand name that tells you nothing: if people recognize it, the label works alone. If not, it needs a plain description or nobody finds it.
Triangulation status

Confirmed means independent voices agree, including at least one who does not personally need it. Building is more than one voice but not yet enough to earn structure. New signal is tracked, waiting on more voices.
Where the two methods met,
and where they disagreed
Twelve months of behavioral data and a global survey, held against each other rather than reported side by side. Verbatims are attributed by role and region only.







The contract was compressed from six months to two, and was likely to end before qualitative work finished. Every deliverable had to be defensible by someone who was not in the room when the decision was made, which is why the rejected options, the unresolved questions, and the evidence tier all stay visible on the page.
Engagement by area,
against a 23% site average
The ask was aggregation,
not more content
Decisions, each traceable
to its evidence
Nothing in the research forced a top-level restructure. All the pressure was inside the existing sections.
The evidence was thinnest
exactly where the stakes
were highest
Reported as qualitative signal and a recruiting pool. Never as a percentage. So I weighted interview recruiting toward the people the survey missed: non-members, patients, caregivers, and participants outside high-income settings.
Each participant was selected against a documented gap, with the reason recorded. A rationale, not a list of whoever answered first.
The handoff was a deliverable,
not an afterthought
The test: can a designer who never sat in an interview tell how confident the research was, and know what would change the answer? Most handoffs answer neither.
Confidence per decision, visible without asking
Open questions travel with the artifact
Written so a non-researcher can run a session
Remaining interviews keep filling documented gaps
Positioned to become the content style guide
- Recruit the thin segment first, not last. The voices carrying the most weight landed last because I scheduled as people became available
- Put the evidence tier in the architecture from version one, as a column in the sitemap rather than an annotation added after
Synthetic panel calibration
A protocol I designed off the back of this project. I have not run it, and it was not client work. The open problem here is the thin patient sample: synthetic panels are sold as the fix for exactly that, and are least trustworthy exactly there.
- A. Prompt-only. A role and some demographics in a prompt, which is what most tools on the market actually do
- B. Evidence-grounded. Grounded in real behavioral data and published research, the version people claim works better
- C. Real respondents. The survey responses, kept completely separate. The answer key
Measured on directional agreement, spread, whether it volunteers criticism, texture of answers, and accuracy by group. A against B is the interesting comparison, because it tells you whether grounding buys anything. Predictions get written down first.