Abstract
New commercial offerings of AI-moderated qualitative interviews — the sort researchers might get with a traditional focus group — promise the depth of in-depth, in-person interviews at the speed and scale of a modern online survey. Yet the performance of such solutions remains poorly understood. In a randomized experiment (n = 3,160) comparing AI-moderated focus-group-style interviews with single respondents to standard written open-ends, we find that the AI interviews delivered substantially richer evidence — e.g., 4.8 times as many words in free-response questions — but for a narrower and non-representative slice of the sample. Completion rates for the survey fell from 99.4% to 40.5%, with drop-off concentrated among groups that aren’t comfortable with AI tools and/or provided shorter responses in open-ended questions — groups that are of interest for researchers and likely contain hard-to-reach respondents. We thus observe positive biases in questions about AI (+7 percentage points on AI optimism), as well as modest tilts on questions about politics and economics (a +6-point Republican lean, with 2024 vote choice statistically unchanged, and 4 points fewer six-figure households) among those who completed the AI interviews vs those in the AI-group who dropped off. While AI-moderated interviews can unlock unusually rich qualitative data, these results show why researchers should not treat them as a drop-in replacement for standard survey instruments without accounting for the systematic, non-random attrition they introduce.
{{inline-form}}
1. Why this matters now
Traditional in-depth interviewing is labor-intensive and expensive, and burdensome for respondents. A platform that lets respondents interact with an AI moderator at their leisure largely solves these problems, and the increasing number of such platforms is quickly changing the economics of qualitative research.
As Verasight has done with other emerging survey research technologies — such as the burgeoning sector of AI-generated survey microdata — we endeavored to verify such approaches for the broader human research community. The recent AAPOR task force report on artificial intelligence in survey research inspired us to think more broadly about how we could contribute to our field’s understanding of how AI is impacting our industry, and AI-moderated online qualitative work looked like a good candidate for a scientific study.
AI-moderated interviewing is now beyond an emerging category of research method, and vendors are already selling results to buyers. But that growth has outpaced the empirical evidence on the subject: there’s real enthusiasm about the method, but its limitations haven’t been fully explored or candidly discussed. This piece asks how well the empirical record actually supports vendors’ claims.
Below, we document the strategy and results of an experiment Verasight conducted jointly on our online panel and with Outset, a leading platform for AI-moderated interviews.
2. Experiment design
We fielded a randomized experiment on the Verasight panel between May 22 and May 31, 2026. All 3,160 participants began a standard Qualtrics survey which measured several pre-treatment covariates — AI use, comfort, and attitudes; tolerance for open-ended questions; and format preferences. Respondents were then randomized to one of two arms containing, alternatively, (a) open-ended questions covering, or (b), an AI-moderated video/audio interview, over the same four topics: a person’s most important electoral issue, their family finances, their biggest concern about AI, and the thing they see as AI’s biggest opportunity.
- Experimental condition A: Open-ended text entry on qualtrics (n = 1,740). Respondents saw four standard open-ended questions in Qualtrics.
- B: An AI-moderated video or audio interview on Outset (n = 1,420). Redirected to an AI-moderated video interview on the Outset platform, in which an AI avatar asked the same four questions aloud and probed with adaptive follow-ups.
The study was conducted from May 22 – May 31, 2026. Our experimental design enables us to explore (in Section 3) how much the AI-moderated interview causes increased respondent dropoff, (Section 4) what types of people drop off, (Section 5) what insights are enabled by AI-moderated interviews that are otherwise inaccessible to researchers, and (Section 6) public reactions to AI-moderated interviews.
3. AI-moderated interviews cause 60-point increase in dropoff
We observe significantly higher rates of respondent dropoff in the AI condition of our experiment. Among panelists who had already agreed to participate in the study, 99.4% of those in condition A completed the written survey, while just 40.5% of respondents given condition B completed the AI interview — a deficit of 58.9 percentage points (95% CI: 56.3 to 61.5 pp, p < 0.001). This is a significant, multiplicative increase in unit non-completion, with a multiplicative gap between arm A and arm B of 94.1x.

Most of the loss occurs at the hand-off between methods. Among respondents assigned to the AI interview, 43.9% spoke any recorded words at all, 40.7% answered all four questions, and 40.5% were recorded as complete interviews. In other words, most attrition happens before the interview meaningfully begins (at the redirect to the external interview platform) rather than in the middle of the conversation. Of the 43.9% of respondents who started their conversation on the AI platform, 92.1% went on to a record a completed interview.

This may suggest more of a technical limitation of the AI-interview platform, rather than avoidance by respondents of the AI survey interface. Several technical problems could explain the low response rates: browsers can fail when redirecting someone across websites (for instance, from Verasight’s online survey platform to the AI-interview software), a user’s microphone can be disabled, a respondent may be in a public space and not able to complete the interview, etc. This is worth more investigation in the future.
4. Attrition is not random
Attrition of this scale would be manageable — or at least non-problematic for estimates — if it were random; the AI condition would simply yield a smaller, but representative, sample of all respondents. But it is not random.
Completion of an AI interview was significantly associated with positive attitudes toward AI, willingness to give verbose responses, and preferences for interview format — all measured before randomization.

Holding the other covariates fixed, enjoying typed open-ended questions was associated with a 13.0 percentage-point higher completion rate; trusting AI companies to protect privacy with +9.6 points; and using AI monthly or more with +8.7 points. Reporting that one tries to skip open-ended questions was associated with a 20.2 point lower completion rate, and ranking the typed-paragraph format first with 14.0 points lower.
A word on the nulls: Puzzlingly at first appearance, the two most AI-specific preference measures — general comfort with AI (+2.1 pp, p = 0.57) and ranking the AI-avatar format first among research formats (-0.5 pp, p = 0.90) — predict essentially nothing once the rest of the model is held fixed. (Comfort and use frequency are correlated at r = 0.75, so use frequency absorbs most of what “comfort” measures.)
These are only descriptive associations, not causal effects, but the pattern might suggest selection into completing the AI interview is not primarily about liking AI, but perhaps about mode fit — willingness to speak, to elaborate, and to tolerate an externally hosted, AI-moderated video experience.
Nevertheless, the resulting evidence pool is therefore not just smaller, it is not representative:

Based on these results, we would infer that using an AI interviewer to measure attitudes toward AI selects on the outcome. Within the AI condition — a contrast that cannot be attributed to noise in randomization — agreement that AI will improve daily life runs 12.5 percentage points higher among completers than among non-completers (95% CI: 7.3 to 17.6 pp, p < 0.001). That selection leaves the completer pool 7.4 points more AI-optimistic than the full sample assigned to the AI condition. A researcher who read only the AI-interview transcripts would, for example, conclude the public is more AI-optimistic than it is.
The demographics and politics of AI-interview completers
We also compare the demographic and political composition of AI-interview completers directly against the non-completers from the same condition. Because both groups come from the same randomized arm, any difference is selection, not noise in the randomization:

Demographically, completers are more statistically significantly likely than non-completers to be men (+9.7 pp) and Black (+9.0 pp), and less likely to be Hispanic, seniors, degree-holders, or from six-figure households.
Politically, the tilt is real and potentially consequential. Respondents who got the AI-interview condition and completed it lean 6.2 points more Republican than non-completers, with fewer pure independents (party composition p = 0.019), than those respondents who got the AI-interview condition and dropped out. Additionally, 2024 vote choice is polarized in the pro-Trump direction among completers (+4.5 pp Trump, p = 0.33).
Attitudinal biases also persist after weighting the completer pool to be demographically and politically representative of the sample again. The weighted proportions are still meaningfully biased:

On the headline AI-optimism measure, roughly three-quarters of the bias survives a full demographic adjustment. Demographic weighting helps at the margin and is worth doing; it does not, however, make this evidence pool representative.
5. But the depth gains are significant
The strongest advantage of AI-moderated interviews is the amount of usable material they generate. And, as with traditional focus groups, these depth-first insights do not necessarily need to be representative to be useful.
For example, when people were asked how their family’s finances were doing, respondents in the written condition typically gave answers like these, reproduced verbatim — the median written answer on this question was 2 words:
“We are getting by .” — “Just ok it could be better” — “It is currently satisfactory, but we don’t have much leeway.”
Here is the same question in the AI condition (one interview; the final follow-up exchange is omitted for space, and everything shown is reproduced verbatim, transcription quirks included):
Interviewer: How would you describe your family’s current financial situation, and how are you feeling about it right now?
Participant: Right now my family has a lot of bills, so do I. We get by and have a little bit of extra money each month.
Interviewer: What factors contribute most to how you feel about your financial situation right now? Is it more about income, expenses, job security, or something else?
Participant: My financial situation is more about the income and how much I get each month through Social Security
Interviewer: How does your current financial situation compare to how you felt about it a year ago?
Participant: A year ago I had less money, over $300 less money and had so many bills to where I barely had any money to spend. or any money to save
The written answers and the initial spoken answer say roughly the same thing: we get by. But the additional information about income source, year-over-year change in dollars, and deferred dental care only exist because of the ability to ask follow-up questions.
Quantifying this depth is obviously an oversimplification of the potential value of such data, but below we count the number of words given by a respondent grouped by condition. For control respondents, we calculate the number of words entered into the text entry box. For respondents in the AI arm, we parse full interview length (which mixed interviewer questions and responses) for participants who completed the survey, and for dropoffs we count them as producing zero words (this was the denominator is every assigned respondent, not every completer, which would produce a skewed picture of the efficacy of the AI interview).
Even with the zeroing out of dropouts, the AI condition surprises. It produced 4.8 times as many respondent words per assigned respondent (128 vs. 27; bootstrap 95% CI for the ratio: 4.4–5.3). Among those who answered at all — a descriptive, post-randomization comparison — the AI condition averaged 291 words against 27.

Even accounting for artifacts of speech (which carries fillers, repetitions, and hedges that typed text does not), respondents in the AI condition gave more information in their responses. Filler tokens make up 4.7% of spoken tokens versus 0.5% of written tokens. Even limiting to unique words, those in the AI condition gave longer answers:

The depth comes almost entirely from follow-up probing from Outset’s AI system. Initial spoken answers alone yield “just” 1.8x the written arm per assigned respondent (people find speaking easier than writing), whereas the additional probes increase the multiplier to 4.8x. Thus the additional probing accounts for roughly 80% of the cross-mode gap. Within clean completed interviews, follow-ups carry 62.4% of respondent words.

But is the added material genuinely informative, or just longer?
To answer that, we designed an automated rubric that scores each answer for markers of usable evidence rather than for volume: concrete details, specific examples, personal context, and causal explanations (the respondent saying not just what they think but why).
Using this data we calculated a few metrics of quality. First, we constructed a conservative metric built specifically not to reward verbosity: an answer scores well for containing these markers, not for being long. According to this metric the AI condition produced 1.8x the information per assigned respondent (a gap of 5.4 quality points, 95% CI: 4.6–6.2).
Then we assembled two cruder rubric-based measures that do incorporate word counts: a ‘proxy usable-evidence score’ (3.9x), based on answer length weighted by the presence of the same evidence markers, and a count of high-usable answers (14.9x), based on the number of a respondent’s answers that clear a high-usability threshold on that score.
This informational yield is core to the economic case for AI-assisted interviews, but ultimately in the eye of the researcher. Because of survey dropoff, one clean AI interview required assigning 2.9 control respondents, versus 1.01 in the written condition. Depending on answers asked and how insights are used, the higher sampling costs may or may not be worth it.
6. Respondents who finish AI interviews rate them positively
Experience measures exist only for respondents who finished the AI interview, so the results in this section are completer-only diagnostics, not treatment effects.
Unsurprisingly, completers of the AI interview tended to enjoy it:

Among clean completers, 86% rated the overall experience 4 or 5 out of 5 and 91% rated it easy to complete. Format preference was more mixed: 45% preferred the AI interview, 29% the written format, 22% had no preference, and 3% were unsure.
Many respondents also told us they were more honest with a machine interviewer (34%) than they would have been in a standard survey (3%). Roughly 63% said they would be the same amount of honest regardless of interviewer. So one upside of the AI-moderated interviewer is that it could produce more accurate measures of quantities subject to social desirability bias. Additional research is necessary to determine if this is true.
7. AI-moderated interviews a good option for depth — on the right questions
“Are AI interviews better or worse than surveys?” is, in our view, the wrong question for researchers to be asking about this technique. Our experiment found that, due primarily to attrition when moving off the original survey platform, AI-moderated instruments do not even produce comparable respondent pools to traditional online questionnaires. And the attitudinal biases of the completing sample can persist despite demographic weighting.
The better question to ask is when do AI interviews add value, for whom, and at what cost to completion and representativeness? We find that:
- On value: the depth gain is real and survives attrition. The AI interviews produced roughly 3 to 5 times the raw material and 1.8x the information per assigned respondent on our length-neutral quality score, with about 80% of the gap attributable to the AI’s follow-up probing (Section 5).
- On cost: we did not measure interviewing costs directly, but the experiment shows one clean AI interview required assigning 2.9 respondents versus 1.01 in the written condition, and 16% of AI completes were fraud-flagged — both multipliers on the mode’s effective price per usable transcript (Sections 3 and 5). This would impact price linearly by CPI.
- On representativeness: completion fell from 99.4% to 40.5%, the evidence pool sits 7.4 points more AI-optimistic than the sample it was drawn from, and demographic weighting closes only 12% to 35% of that tilt. The mode retains a selected pool — more male, more Black, less Hispanic, fewer seniors and degree-holders, and above all more AI-comfortable and more willing to elaborate. Completers differ from non-completers by 12 to 17 points on pre-treatment AI attitudes (Section 4). Like other qualitative methods, the evidence suggests AI-moderated interviews are not inherently representative, even when care is taken to give systems a representative sample. (Sections 3 and 4). These demographic biases may be combatted by targeted sampling.
A simple decision tree for which mode to use may look like:
- Use AI-moderated interviews for depth, discovery, mechanism, respondents’ own language, segmentation, and hypothesis generation — wherever the scale advantage over human-moderated interviewing is decisive and a biased completer pool is an acceptable price for transcripts.
- Keep written and mixed-mode designs for coverage, low-burden participation, and representative point estimates. The written mode retained 99.4% of assigned respondents; no qualitative richness compensates for losing six in ten people non-randomly when the estimand is a population share.
Finally, a few words on limitations and future work. One big potential caveat to our findings is that respondents were not told in advance that they may be sent to a live, AI-moderated video or audio interview as part of an otherwise standard online survey. Some of the loss at the platform hand-off (Section 3) may reflect respondents who, at the moment of redirect, simply were not in the setting or mood for completing a spoken interview. We can imagine that offering these respondents a larger incentive could improve the drop-off rate.
Future studies should test whether telling respondents up front what condition B entails, so they can choose to start it only when they're in a suitable place to do so, and/or paying a higher incentive for the higher-burden arm, narrows this completion gap. That would tell us how much of the drop-off documented here is a fixable artifact of study design rather than an inherent cost of AI-moderated interviewing.
How to cite this report
Morris, G. Elliott, Saeideh Bakhshi, Benjamin Leff, Jake Rothschild, and Joey Marshall. 2026. "Depth at a Cost: A Randomized Comparison of AI-Moderated Interviews and Written Survey Open-Ends." Verasight.
