The Observatory × Dialogue AI | Sampling & Saturation
May 2026
Reference Guide
How Many People.
How Good the Data.
Sampling strategy and interview quality measurement for AI-moderated research. Field-tested and by the books.
The Numbers at a Glance
Qual sweet spot
4–10
per segment
Quant sweet spot
300–800
per segment
√n rule
cost to halve uncertainty
Quality metric
IQS
3-component score
When to Stop Adding People
1 / 7
Qual vs. Quant: The Benchmarks
Two different stopping rules. Qual asks: am I still learning? Quant asks: how tight is my interval? The benchmarks below are field-tested across 15 years of studies and consistent with standard margin-of-error guidelines.
Segment = any subgroup you need to analyze separately: a demographic (men 18–35), a behavior (weekly users), a market (Germany), an attitude (skeptics). If the decision depends on seeing the difference between two groups, each group is a segment.
Qual
AI-moderated IDIs
Per segment
Quant
Panel or proprietary surveys, AI-fielded
1–3
"done" / passing grade
Minimum viable
50–150
not great, but people do it
4–10
solid, this is the sweet spot
Recommended
300–800
academic quality from here down
10+
luxurious, overlaps with quant
Ideal
800–1,200
ideal for gen pop segments
15+
agencies love this scale
Overkill?
1,200+
🤠 but expensive or just extra
Dialogue's edge: AI moderation makes qual's "luxurious" tier the new default, and makes quant's sweet spot reachable at panel pricing.
Double click: what the qual literature actually says
Interviews per segmentWhat the literature saysWhat you get
3–6
Passing grade
Guest et al. (2006): 70% of themes by interview 6.
Field Methods, 18(1), 59–82
You've heard most of the topics. You haven't understood any of them deeply. Directional signal only.
9–12
Sweet spot
Guest et al.: 92% of themes by 12. Hennink et al. (2017): code saturation at 9.
Qualitative Health Research, 27(4), 591–608
You've heard it all (code saturation). You know the landscape. You may not yet understand the texture, the contradictions, the meaning underneath.
12–16
Luxurious
Hennink et al.: meaning saturation begins around 16. Hagaman & Wutich (2017): sufficient for homogeneous groups.
Field Methods, 29(1), 23–41
For a single, relatively homogeneous segment: understanding deepens. You're hearing the same things, but catching the nuance, the exceptions, the tension.
16–24
Meaning saturation
Hennink et al.: meaning saturation at 16–24.
"Code saturation may indicate when researchers have 'heard it all,' but meaning saturation is needed to 'understand it all.'"
Full meaning saturation. You understand not just what people say but why, and where the contradictions live.
20–40
Cross-cultural
Hagaman & Wutich: metathemes (themes that hold across subgroups) took up to 35–39 interviews within a single site, and 20–40+ across sites. Multi-site or cross-cultural work. This is what it takes to see what's universal vs. what's site-specific.
Malterud et al. (2016) add a useful nuance: information power. The more information a sample holds relevant to the study, the fewer participants you need. It cuts both ways: n inflates precisely when the aim is broad, the sample is heterogeneous, or theory is absent. Quality of dialogue is literally one of their five dimensions, which means better moderation can move the saturation point earlier.
Double click: the quant margin-of-error math
From sample size to precision: what margin of error you get at a given n. Standard 95% confidence, worst-case proportion (p = 0.50).
Sample sizeMargin of errorIn plain language
50
±14%
If 60% said yes, true answer is somewhere between 46–74%. Very rough.
100
±10%
Directional. You know which way it leans.
200
±7%
Starting to stabilize.
400
±5%
Standard threshold for most applied research.
600
±4%
Solid. Where a lot of good product research lives.
1,000
±3%
Tight. Academic-quality precision.
2,500
±2%
Very precise. Rarely needed for product decisions.
From target precision to required sample: what it costs to get tighter.
"I want to be within..."I need (per segment)Cost multiplier vs. ±5%
±10%
~100
0.25×
±7%
~200
0.5×
±5%
~400
1× (baseline)
±4%
~600
1.5×
±3%
~1,000
2.5×
±2%
~2,500
6.25×
±1%
~10,000
25×
The punchline: Going from ±5% to ±3% costs 2.5× more people. Going from ±3% to ±2% costs another 2.5× on top of that. The question is always: does that extra precision change the decision you'd make?
Niche audiences
"Niche" is a catch-all for three different recruiting problems: rare (big wave surfers, edge-case power users), hard-to-reach (VIPs, C-suite), and expensive (ultrawealthy, specialists who bill by the hour). They share one reality: you may never hit the numbers above. Recruit opportunistically, aim for the gen pop benchmarks as a ceiling, and make every conversation count. When sample size is constrained, interview quality carries the study.
Diminishing Returns and Sample Size
Adding people to a study helps. Then it helps less. Then it barely helps at all. In qual, you stop discovering new things. In quant, your confidence interval stops shrinking. The curve flattens either way. Knowing where it flattens is the whole game.
Side by Side
Qual Only
Quant Only
Qualitative Saturation themes
Each figure = a theme. One interview surfaces many.
Interviews
n=0
Survey Precision respondents
Each figure = a respondent. Simulated: 60% said "yes."
Respondents
n=0
Qualitative Saturation
Each figure = a theme or pattern. One interview surfaces many.
No data. Every theme is undiscovered.
Interviews
n=0
03
passing
6
70%
9
codes
12
92%
16
meaning
24
full
40
Core
0%
Variations
0%
Edge cases
0%
Meaning
0%
Survey Precision
Each figure = a respondent. Simulated survey: 60% of the population would say "yes." As you sample more, your estimate tightens around that truth.
No data yet.
Respondents
n=0
050
rough
100
±10%
400
±5%
600
±4%
1k
±3%
2.5k
±2%
The left is about discovery. The right is about precision. Different currencies, same lesson: past a certain point, more people don't change what you know. They just make you more sure of what you already knew. Unless you change the question.
Double click: the four things saturation can't see
Code saturation asks one question: have I heard this theme before? It is deliberately blind to four things that only appear at scale. This is where large-N qual earns its keep.
Blind spotWhat saturation misses
Prevalence
Saturation tells you a pattern exists. It can't tell you whether it's 9% of the population or 60%. How common a theme is, is a quantity of the qualitative, and it's invisible at n=15.
Segment structure
Clusters found at small n are hypotheses, not findings. Whether there are four segments or seven, and whether the boundaries hold, is a large-N question.
Rare-but-decisive cases
The participant whose entire decision logic is governed by one formative story. At n=15 that's noise. At n=400 it's a segment you can size and name.
Interaction effects
Does the pattern differ by subgroup? Does the say-do gap change with context? You need cells, and cells need n.
Saturation answers "have I heard it all?" These four answer "what is the population actually like?" Different questions, and the second set has always been priced out of qual. AI moderation changes the price.
Double click: qualitative saturation curves (static reference)
Qual: discovery curve. Percentage of themes and codes identified as interviews accumulate. Two curves: code saturation (you've heard the range of topics) flattens earlier than meaning saturation (you understand what they mean).
100% 92% 70% 0% 0 6 12 16 24 interviews per segment % themes discovered
Code saturation
Meaning saturation
Based on Guest et al. (2006) and Hennink et al. (2017)
Double click: quantitative margin-of-error reference
From sample size to precision: what margin of error you get at a given n. Standard 95% confidence, worst-case proportion (p = 0.50).
Sample sizeMargin of errorIn plain language
50
±14%
If 60% said yes, true answer is somewhere between 46–74%. Very rough.
100
±10%
Directional. You know which way it leans.
400
±5%
Standard threshold for most applied research.
600
±4%
Solid. Where a lot of good product research lives.
1,000
±3%
Tight. Academic-quality precision.
2,500
±2%
Very precise. Rarely needed for product decisions.
From target precision to required sample: what it costs to get tighter.
"I want to be within..."I need (per segment)Cost multiplier vs. ±5%
±10%
~100
0.25×
±5%
~400
1× (baseline)
±4%
~600
1.5×
±3%
~1,000
2.5×
±2%
~2,500
6.25×
±1%
~10,000
25×
The punchline: Going from ±5% to ±3% costs 2.5× more people. Going from ±3% to ±2% costs another 2.5× on top of that. The question is always: does that extra precision change the decision you'd make?
Interview Quality Score (IQS)
Every interview scored. Is this a participant problem, a moderator problem, or a data problem? Here's what a typical 100-interview study looks like.
Sample study · n = 100
4.2 / 5
Signal: 4.0 · Moderator: 4.5 · Usability: 4.6
5 Exceptional (40)
4 Strong (40)
3 Adequate (12)
2 Thin (2)
1 Near-empty (1)
Flagged (1)
Dead/void (4)
*Example only. Hover each cell for simulated participant details.
The weights are tunable per study. A brand perception study weights signal depth higher. A cross-market study weights data usability higher. The point is the decomposition: when a study goes sideways, you know where it went sideways.
Double click: how IQS is scored
🔍
Participant Signal Depth
Weight: 50%
Did this person give us something real? Scored on specificity of examples, unprompted elaboration, internal contradiction and complexity, and behavioral evidence: what they did, not just what they believe. Idiosyncrasy (strange, personal detail a model wouldn't generate) is the hardest dimension to fake and the strongest signal of authenticity.
🤖
Moderator Effectiveness
Weight: 30%
Did the AI do its job? Scored on probe quality (following up on signal-rich moments), confidence integrity (distinguishing participant-originated meaning from prompted agreement), pacing, and recovery from thin answers. Every session is logged. Every session is scoreable.
Data Usability
Weight: 20%
Can we trust this? Transcript complete, participant in-market, translation intact, no authenticity red flags. Authenticity signals use two layers (textual and behavioral) to catch AI-assisted responses. For cross-market: out-of-market detection, translation degradation, panel self-awareness tagging.
IQS = (Signal × 0.5) + (Moderator × 0.3) + (Usability × 0.2)
Study typeSignalModUsabilityWhy
Standard qual50%30%20%Signal is king.
Cross-market40%25%35%Translation and in-market verification carry more risk.
Brand perception55%30%15%Emotional depth matters more than data hygiene.
Concept testing45%35%20%Probing reactions to stimuli is critical.
Niche audience50%20%30%Fewer interviews. Can't afford to lose data.
The Memes (Because Why Not)
The sampling Drake
😐
"We need 2,000 respondents per segment for statistical significance"
😎
"We need enough to stop learning new things, and the budget to actually get there"
Stages of sample size enlightenment
😶
n = 3 per segment"Done! Ship it!"
🤔
n = 8 per segment with AI moderation"Now we're actually hearing the contradictions"
😎
n = 400 quant + n = 8 qual per segment"Mixed methods, still under the old budget"
🤠
n = 1,200 quant because "more is better""Buying decimals, not insight"
🔥
IQS on every interview, weights tuned per study"We don't just count the people. We grade the conversations."
The moderation chat log
Client
How do you ensure quality at scale?
Agency (traditional)
We have experienced moderators
Client
Can you prove that?
Agency
... they're very experienced
Dialogue AI
Every session scored. Signal depth: 7.8/10. Moderator effectiveness: 8.2/10. Data usability: 100%. Here's the per-interview breakdown.
The quality assurance Drake
😐
"Trust me, the data is good. I read all 40 transcripts myself"
😎
"Every interview scored on 3 dimensions. Study scored 7.6/10. Here's what dragged it down."
The Pitch in One Scroll
The problem. Most teams run too few interviews (budget wall) or too many surveys (precision theater). The first misses the real story. The second buys decimals.
What changes. AI moderation doesn't move the saturation point. It moves the budget line. The math is the same. You can just afford to get there now.
What we measure. IQS: Interview Quality Score. Signal depth (did the participant give us something real?), moderator effectiveness (did the AI do its job?), data usability (can we trust this?). Decomposable, benchmarkable, auditable. Nobody else scores this way.
The short version. More people. Better conversations. A number that proves it.