Ray Poynter, 2 October 2026
Note, this is written in my personal capacity and may not reflect the views of any other organisation.
A valuable study, read too broadly
Pew Research Center’s new report, Can AI Stand In for Human Survey-Takers? Not Really, published on 30 September 2026, is one of the most careful public tests of synthetic respondents we have seen to date. Everyone in the insights industry should read it.
However, much of the commentary is treating it as a broader verdict. I am seeing lots of comments in the style of ‘Synthetic data has been tested, and it failed.’ My view is that the study answers a narrower question. It shows where one particular approach to synthetic polling breaks down. That is useful knowledge, but it does not settle the wider question of whether synthetic data works and, if so, where it works.
What Pew did
Pew took three waves of its American Trends Panel from early 2026 and created an AI “digital twin” for each panellist who answered them.
Each twin was an off-the-shelf large language model, mainly Claude Opus 4.6, prompted with that person’s demographics, 71 attitudinal variables from Pew’s 2025 political typology survey, and a model-written “expert reflection” on the profile. The twin then answered the same questionnaire, in the same order, with the same randomisation.
Across nearly 300 questions, the synthetic results differed from the human results by about 12 percentage points on average. The twins stereotyped subgroups, avoided extreme answer options, rarely said “not sure”, ‘knew’ far more than real people, and missed shifts in opinion on current events. Changing the underlying model also changed the errors, with GPT-5.1 producing a different picture of the public from Claude Opus 4.6.
These are real findings, and anyone selling synthetic polling should take them seriously.
A diagnostic evaluation with a narrower remit
Pew conducted an evaluation. It is a rigorous evaluation of an important use of synthetic data. They looked at whether AI-generated digital twins can reproduce answers from a high-quality probability panel on topical public-opinion questions.
The key is understanding what that evaluation was designed to tell us.
Pew chose the digital-twin approach partly because it lets researchers see where a model succeeds or struggles. For example, which subgroups it misrepresents, which questions cause problems, and what sorts of errors occur. Pew also notes research suggesting that approaches designed to estimate aggregate results directly can produce better aggregate accuracy.
A broader evaluation would ask additional questions. For a specific research task, how accurate is the best available approach? How does it compare with the realistic alternatives? And is it accurate enough for the decision being made?
Pew’s design therefore works rather like a microscope. It tells us a great deal about where this particular implementation develops cracks. It tells us less about how the best available synthetic approach would perform on a defined commercial task.
Pew’s own reservations
Pew is admirably clear about the boundaries of what it did. They include:
- One use of synthetic data. The study looks at AI acting as a survey-taker. It explicitly excludes other uses, such as imputing missing values.
- One broad class of research problem. It focuses on questions of public importance, benchmarked against a high-quality probability panel.
- Today’s tools. Pew notes that newer models and future innovations in synthetic sample construction could produce smaller errors. The models they used are already very out of date.
- Constrained configuration testing. Time and budget meant testing settings sequentially rather than evaluating every possible combination.
- Cost-limited scale. The first wave used a subsample of 6,700 of 8,512 panellists, and the number of replicated questions was limited by budget.
- Mostly political and entirely American. Most questions concerned US politics and current affairs.
- A design that favours diagnosis. Pew acknowledges that some aggregate approaches tend to be more accurate and chose twins partly because they reveal individual and subgroup failure modes.
- A learning exercise. Pew says it has no current or future plans to use AI to generate survey results. The project was designed to understand the method rather than identify a production solution.
None of these points is a criticism of Pew. These are boundaries Pew places around its own work. The problem comes when commentary extends the conclusions beyond them.
Clean data, but not rich in every relevant domain
Some commentary suggests that if Pew, with all its panel data, cannot make digital twins work, nobody can. I think that conclusion goes too far.
The Pew profiles were strong on demographics, politics and social attitudes. Each twin had a standard demographic profile plus 71 attitudinal variables, along with an expert reflection generated from that information.
But the data were much thinner in some of the other areas the twins were asked about. When questions concerned topics such as sleep, foreign travel or sport, there was relatively little person-specific information available. The model therefore had more scope to fill gaps using patterns learned elsewhere, including stereotypes.
Compare that with some of the datasets being explored elsewhere in the industry:
- Commercial panels where members may have completed hundreds of surveys over several years.
- Online communities containing years of discussions, diaries and repeated measures on the same people.
- Toluna’s HarmonAIze personas, drawing on anonymised first-party information from its large global panel.
- The Columbia Twin-2K-500 dataset, containing more than 500 answers per person.
- Simile’s work with CVS Health, drawing on almost 3 million consented responses from more than 400,000 people.
There is an important caution here. The Columbia researchers found that having more than 500 previous answers per person improved performance only modestly compared with using no personal information. Richness alone is clearly no guarantee.
Pew’s own experiments nevertheless showed that adding relevant information helped. With GPT-5.1, average error fell from about 16 points using the standard profile to 13.7 using the extended profile and 13.1 after adding the expert reflection.
The richness question therefore remains open and important. Pew’s data are perhaps better described as clean and strong in particular domains, rather than comprehensively rich across people’s lives.
Expert pollsters, different expertise
The Pew team are outstanding survey methodologists and pollsters. Their judgement on sampling, weighting and questionnaire design is among the best in the world, and the transparency of this report is exemplary.
Building synthetic solutions introduces some different choices. These can include deciding between personas, twins and aggregate approaches; fine-tuning models on proprietary data; calibrating outputs against known benchmarks; correcting known biases such as compressed variance; providing current context; and validating outputs against the decisions a client needs to make.
Pew used off-the-shelf models with prompting rather than fine-tuning or post-hoc calibration. Its pipeline also did not provide explicit current-event grounding beyond the information already available to the model.
That is a sensible and informative methodological design. However, it does not include some of the techniques used by specialist synthetic-data providers.
The same standard of evidence needs to apply in the other direction. Commercial providers have an incentive to demonstrate that their systems work, and much of their methodology is proprietary. Supplier claims about improved performance therefore need independent validation rather than being accepted simply because the system is more sophisticated.
An analogy might help. If a world-class team of statisticians built an online survey platform from scratch and found serious problems, we would learn a great deal about the pitfalls of building survey platforms. We would still want independent evidence about what experienced platform providers achieve before reaching conclusions about online research as a whole.
What the study does tell us
Pew’s findings line up with other independent work, including Verasight’s synthetic sampling studies and the Columbia digital twins research. It would therefore be a mistake to dismiss the results as simply a consequence of Pew’s implementation.
The study provides a useful checklist of failure modes that should be tested in any synthetic offer:
- stereotyping;
- flattened variance;
- too few “not sure” responses;
- excessive knowledge;
- weakness around new events;
- sensitivity to the underlying model.
It also raises an interesting distinction between getting the level right and getting the pattern right.
On abortion, for example, the synthetic share favouring legality was close to the human result, but the extreme response categories largely disappeared. Partisan directions were usually reproduced, although some gaps were exaggerated.
Pew primarily reports average absolute error. Depending on the research decision, correlation, ranking or direction agreement may sometimes matter as much as absolute levels.
What a broader evaluation would include
A broader evaluation of synthetic data would:
- Define the use case and the decision the research must support.
- Compare several approaches, including aggregate models, calibrated personas and fine-tuned twins where appropriate.
- Compare synthetic approaches with the realistic alternatives for that decision, such as human samples, historical data or conventional modelling.
- Use the richest appropriate and ethically available data.
- Define evaluation criteria in advance, including level accuracy, direction, ranking and correlation.
- Validate on held-out questions or data so that systems are not optimised against the answers on which they will eventually be judged.
- Test across different categories, including commercial topics where synthetic approaches are already being used.
- Where possible, use independent validation rather than relying solely on supplier testing.
Pew has published the comparison toplines, which makes some additional analysis possible. Let’s hope researchers take advantage of that.
Closing thoughts
The Pew study is a gift to the industry. It is careful, transparent and full of lessons.
It provides strong evidence that a well-designed, off-the-shelf, prompt-based digital-twin approach is currently a poor substitute for probability polling on topical national issues. That is an important finding.
The bigger question remains open. Where, and with which methods, is synthetic data good enough for the decisions people use it for?
Answering that requires evaluations designed around specific use cases, employing strong competing approaches and judged against the precision the decision actually requires. It also requires independent evidence about commercial systems, rather than assuming that greater sophistication automatically means better answers.
For now, the Pew study gives us an excellent diagnosis of some important failure modes and a demanding benchmark against which other approaches should be tested.
The debate about synthetic data should remain open and evidence-led.
Sources
- Pew Research Center: Can AI Stand In for Human Survey-Takers? Not Really
- Pew Research Center: Methodology
- Verasight: The Limits of Synthetic Samples in Survey Research
- Peng et al.: Digital twins are funhouse mirrors: Five systematic distortions
- CVS Health: How CVS Health test-drives better care experiences using generative agents
- Toluna: Synthetic Personas
- Emporia Research: Beyond the Buzz: 4 Core Ways Vendors Build “Synthetic” Survey Respondents
