Skip to main content

How Merlin works

A defined audience, modelled one respondent at a time.

Audience research describes who the respondents are. A language model supplies the wider context needed to interpret a question that was never in the original survey. Merlin puts the question to each modelled respondent separately, then combines their answers into a distribution.

The seed data comes from leading survey companies, and each persona is constructed from one real respondent's answers — their demographics and context — so that the population being modelled is a specific, documented one rather than a general impression of a group.

  1. Input — audience research

    Seed data about real people defines who is in the audience: demographics, attitudes, behaviours, media, purchase history.

  2. Context — language model

    Wider world knowledge lets a persona reason about a question the original survey never asked.

  3. Output — a distribution

    Many modelled answers, showing the pattern and the spread rather than a single view.

Synthetic audience: an AI representation of your real-world audience, shown as a crowd with individual respondents labelled by segment

A synthetic audience is a model of a real-world audience — a crowd of individually modelled respondents, each carrying their own profile.

The mechanics of personas, seed data and audiences are covered in Concepts; the terms themselves are in the Glossary.

Why not just ask a generic LLM?

Merlin uses language models. That is not the part that makes it research. The difference is the audience data underneath it, the fact that respondents are asked separately, and the evaluation against that specific audience.

Generic LLMMerlin
Starting pointBroad training data and your promptResearch about a defined audience
OutputUsually one plausible answerA distribution from many modelled respondents
Audience testNone, unless you build onePredictions compared with withheld real answers

Ask a general-purpose model "how would UK music streamers react to a £2 price rise?" and you get one answer, shaped by whatever the model has absorbed about music streamers in general and by how you phrased the question. There is no defined population behind it, no spread, and nothing to check it against.

Merlin answers the same question by asking every persona in a named audience — built from a specific dataset, fielded on a specific date — and reporting how many said what. That is a result you can interrogate: you can cut it by segment, ask why in a focus group, and check how well the question sits against the data the audience was built from.

This difference has been measured rather than asserted. Electric Twin ran both approaches against the Gallup World Poll — asking ChatGPT directly to predict survey response distributions, and running the same questions through the Electric Twin engine — and scored each on distributional accuracy.

Two charts comparing accuracy per question category: ChatGPT's results are widely spread from below 0.4 up to 1, while Electric Twin's cluster tightly near 1

The same questions, both methods. Electric Twin wins across every question category and on average across all countries — and the tighter clustering is the bigger story.

Two things to take from it. Electric Twin is more accurate on average, across every category and country. And its predictions vary far less: ChatGPT's accuracy swings from under 0.4 to near 1 depending on the question, while Electric Twin's sit in a narrow band near the top.

That second point is the one that matters in practice. A method that is sometimes excellent and sometimes badly wrong, with no way to tell which you got, is not usable for a decision. Consistency is what makes a result something you can plan around — which is also why it is one of the two signals behind every confidence grade.

The model is not the whole method

Three things follow from this, and they are worth holding on to when explaining Merlin:

  • The audience is the unit of quality, not the model. A well-defined audience built from strong seed data will outperform a vague one on the same underlying model. See Audiences.
  • The spread is part of the answer. A 55/45 split and a 95/5 split mean very different things, and a single generated opinion would hide the difference.
  • Every result carries its own grade. Confidence grades tell you how far the question you asked sits from the data the audience was built on, and how stable the answer was across re-runs.

Next