Confidence grades
Every survey result in Merlin comes with a grade — a quick read on how much weight to put on what you're seeing. This page explains what the grade means, how to interpret it, and what to do when a result is less reliable than you'd hoped.
At a glance
Each result card shows the grade as a letter — A, for example. Click it to see the two signals behind the grade and, where relevant, what to change.
| Grade | What it means | What to do |
|---|---|---|
| A | Very high confidence — your audience data covers the topic and the model answered consistently across re-runs. | Use it as it stands. |
| B | Confident — a good combination of coverage and consistency. | Good to act on. The sub-scores are worth a glance, but they don't need changing. |
| C | Slightly more caution — this doesn't necessarily mean a less accurate result, but we are slightly less confident. | Look to the sub-scores — a better-phrased question or audience change can lift the confidence before you rely on it. |
The two signals behind the grade
The grade is built from two underlying signals. Clicking the grade letter shows both, each as its own letter grade, so you can see why a result landed where it did.
Source alignment answers the question "Is this question close to the audience data we have?" It compares the question you asked against the survey data your audience was built from. A question about TV streaming asked against a media-consumption audience aligns closely; a question about, say, foreign policy asked against that same audience aligns poorly.
Consistency answers the question "Did the model give the same answer when we asked it more than once?" Merlin runs every question several times behind the scenes. If the model returns broadly the same distribution each time, consistency is strong. If results swing around between runs, it is weak. Consistency is the strongest individual predictor of accuracy in our testing.
The two sub-scores are graded on different scales, so read each on its own terms rather than comparing the two letters directly. Consistency is held to a higher bar than source alignment, so the same letter can reflect a different underlying strength on each.
The combined grade comes from how the two signals interact. You can think of it as a 3 × 3 grid:
| Source alignment A | Source alignment B | Source alignment C | |
|---|---|---|---|
| Consistency A | A | A | B |
| Consistency B | A | B | C |
| Consistency C | B | C | C |
Reading the grade
For a single-audience result, click the grade letter to see, top to bottom:
- The grade — A, B, or C
- The two sub-scores — Source alignment and Consistency, each as its own letter grade with a plain-English explainer
- One line of advice, when a signal needs attention
The advice is specific about which lever to pull:
Source alignment could be improved — try a different audience for this topic.
Consistency could be improved — try rephrasing or constraining the question.
Multi-segment results
When a result card breaks a question down across multiple segments, the headline grade shows the weakest segment rather than an average. One unreliable segment is therefore never hidden behind two good ones.
Click the headline grade to open a table of segments, ordered weakest first, each with its own grade. The header shows the range across segments — for example C – A when they disagree, or a single grade when they all agree. Expand any row to see that segment's two sub-scores and any advice specific to it.
What the grade does not tell you
A few things worth being explicit about:
- It is not a measure of how interesting the result is. An A result can still be a boring one, and a C result can still be directionally useful.
- It does not replace your judgement about whether the right audience was built in the first place. The grade tells you whether a question lands well against the audience you have defined. Whether that audience is the right one for the decision you are making is a separate question, and one Merlin does not yet score for you.
- The grade is a band, not a per-result margin of error. Two results with the same grade share the same band even though their underlying sub-scores differ.
- It does not apply to free-form chat answers or open-ended qualitative outputs. Grades cover structured survey and tool results only. Qualitative output is evaluated — focus groups and debates are scored against real polling data — but that happens at the platform level, not per result. See How we know it works.
Those boundaries are the same ones described in Presenting a result.
Exporting results
Grades appear only in Merlin and aren't yet included in exports. Adding them is on the roadmap. When you quote a result outside the product, carry the grade across by hand — the methodology line has a place for it.
Next
- Presenting a result — wording that holds up when you share a result