Skip to main content

Confidence grades

Every survey result in Merlin comes with a grade — a quick read on how much weight to put on what you're seeing. This page explains what the grade means, how to interpret it, and what to do when a result is less reliable than you'd hoped.

At a glance

Each result card shows the grade as a letter — A, for example. Click it to see the two signals behind the grade and, where relevant, what to change.

GradeWhat it meansWhat to do
AVery high confidence — your audience data covers the topic and the model answered consistently across re-runs.Use it as it stands.
BConfident — a good combination of coverage and consistency.Good to act on. The sub-scores are worth a glance, but they don't need changing.
CSlightly more caution — this doesn't necessarily mean a less accurate result, but we are slightly less confident.Look to the sub-scores — a better-phrased question or audience change can lift the confidence before you rely on it.

The two signals behind the grade

The grade is built from two underlying signals. Clicking the grade letter shows both, each as its own letter grade, so you can see why a result landed where it did.

Source alignment answers the question "Is this question close to the audience data we have?" It compares the question you asked against the survey data your audience was built from. A question about TV streaming asked against a media-consumption audience aligns closely; a question about, say, foreign policy asked against that same audience aligns poorly.

Consistency answers the question "Did the model give the same answer when we asked it more than once?" Merlin runs every question several times behind the scenes. If the model returns broadly the same distribution each time, consistency is strong. If results swing around between runs, it is weak. Consistency is the strongest individual predictor of accuracy in our testing.

The two sub-scores are graded on different scales, so read each on its own terms rather than comparing the two letters directly. Consistency is held to a higher bar than source alignment, so the same letter can reflect a different underlying strength on each.

The combined grade comes from how the two signals interact. You can think of it as a 3 × 3 grid:

Source alignment ASource alignment BSource alignment C
Consistency AAAB
Consistency BABC
Consistency CBCC

Reading the grade

For a single-audience result, click the grade letter to see, top to bottom:

  • The gradeA, B, or C
  • The two sub-scores — Source alignment and Consistency, each as its own letter grade with a plain-English explainer
  • One line of advice, when a signal needs attention

The advice is specific about which lever to pull:

When source alignment is weak

Source alignment could be improved — try a different audience for this topic.

When consistency is weak

Consistency could be improved — try rephrasing or constraining the question.

Multi-segment results

When a result card breaks a question down across multiple segments, the headline grade shows the weakest segment rather than an average. One unreliable segment is therefore never hidden behind two good ones.

Click the headline grade to open a table of segments, ordered weakest first, each with its own grade. The header shows the range across segments — for example C – A when they disagree, or a single grade when they all agree. Expand any row to see that segment's two sub-scores and any advice specific to it.

What the grade does not tell you

A few things worth being explicit about:

  • It is not a measure of how interesting the result is. An A result can still be a boring one, and a C result can still be directionally useful.
  • It does not replace your judgement about whether the right audience was built in the first place. The grade tells you whether a question lands well against the audience you have defined. Whether that audience is the right one for the decision you are making is a separate question, and one Merlin does not yet score for you.
  • The grade is a band, not a per-result margin of error. Two results with the same grade share the same band even though their underlying sub-scores differ.
  • It does not apply to free-form chat answers or open-ended qualitative outputs. Grades cover structured survey and tool results only. Qualitative output is evaluated — focus groups and debates are scored against real polling data — but that happens at the platform level, not per result. See How we know it works.

Those boundaries are the same ones described in Presenting a result.

Exporting results

Grades appear only in Merlin and aren't yet included in exports. Adding them is on the roadmap. When you quote a result outside the product, carry the grade across by hand — the methodology line has a place for it.

Next