Executive Summary
This article does not ask how accurate an AI is when it answers a survey in a person's place. It follows the question one step back: which tier that accuracy was established at. Oded Netzer of Columbia Business School and Rajan Sambandam, president of the American market research firm TRC Insights, published a paper in September that took roughly a hundred attitude questions and measured the same synthetic respondent with four different yardsticks. The four numbers said four different things.
The heaviest finding fits inside one paragraph. Twins given nothing but demographics, age and gender and the like, barely recovered any individual's answers. Across all 108 questions the individual-level correlation never reached 0.5. Yet the topline those same twins produced, once everyone was averaged together, landed closer to the human topline than the condition that had also been handed five attitude questions from the same sector. The two numbers come from the same data at the same moment and both are true. The abstract names what that pairing does: aggregate measures hold up even when the model was told very little, and in holding up they mask the complete absence of respondent-level differentiation.
So the question to put to a vendor of synthetic respondents is not what the accuracy percentage is. It is which of four families that number came from, which of three tiers your own decision sits at, and whether the two are the same tier. When they are not, "it matched well" remains true and stops having anything to do with your decision. No clause in the international norms already governing this industry requires anyone to answer that question yet.
How closely a twin told only age and gender tracked one person's answers
Across 108 questions, none of them reached 0.5
Topline error, and the condition given less information is the smaller one
Same data, same moment, two numbers
Questions the screen clears for a twin to answer
The other 62 are flagged as needing new human data
A Korean panel's cell-level error against a baseline that uses no information
Uncalibrated, the synthetic panel came out behind the baseline
Two Measures, Same Data, Opposite Directions
A word first, because it carries two unrelated meanings. The "digital twin" in this article is not the kind that rebuilds a factory line or an organ inside a computer so engineers can run it. It means getting a large language model to answer a survey in the place of one particular person, and it also means the fake respondent that results. The industry calls these synthetic respondents, synthetic consumers, synthetic panels. All three name the same object from slightly different angles, and this article reaches for whichever fits the sentence. The mirror-image problem, where a survey put to humans quietly fills up with answers that no human wrote, we covered separately in You Asked People. Increasingly, AI Is Answering.
The starting point is a paper posted in September. It has two authors who stand in very different places. One is Oded Netzer, a professor at Columbia Business School. The other is Rajan Sambandam, president of TRC Insights, a market research firm that sells human sample for a living. The empirical material comes from a large 2025 survey run by the Center for Customer-Based Execution and Strategy at Rice University's Jones Graduate School of Business, and the firm that fielded it was that same company. Better to say the conflict out loud: a paper about the limits of synthetic responses is co-authored by someone whose business those responses would eat. The paper raises the structure itself in its introduction. Optimists have commercial incentives, it notes, and so do pessimists, since a great deal of the loudest skepticism originates with firms whose business is human sample.
The survey asked customers what they value across 18 sectors. There were 3,063 respondents, but no one rated all 18. Each person took up to three sectors, which leaves roughly 500 actual respondents per sector. Six attitude questions per sector makes 108 in all. From this material the researchers built two conditions side by side. In one, the twin was told only demographics, age and gender among them. In the other, the twin was also handed the five remaining attitude questions from the same sector. Then both conditions answered the same questions again.
Put the results in one table and the two measures walk in opposite directions. Individual-level correlation climbed sharply when the twin was given more to work with. The topline error, computed by averaging everyone into a per-question mean and comparing that mean against the human one, was smaller in the condition that had been told less.
| What the twin was given | Individual correlation | Questions above 0.5 | Topline error |
|---|---|---|---|
| Demographics only | 0.08 | 0 of 108 | 0.39 |
| Demographics plus 5 attitude questions from the same sector | 0.57 | 74.1% | 0.48 |
Netzer and Sambandam (2026), Table 1 and section 5.2. The correlation captures how well the twin recovered people's high-and-low ordering, question by question; the topline error is how far the per-question mean landed from the human mean. Lower error is better.
Read the top row. Given your age and your gender, the twin recovered essentially none of you as an individual. A correlation of 0.08 is close to zero, and all 108 questions sat below 0.5. Yet the per-question means that same twin generated were 0.39 away from the human means, and that is smaller than the 0.48 produced by the condition with five extra attitude questions in hand. More information pushed the topline further off.
In the demographics-only condition, the average individual-level correlation between twin and human answers was 0.08, and not one of the 108 questions cleared a correlation of 0.5. Yet the aggregate MAE in that condition was 0.39, slightly better than in the demographics-plus-attitudes condition. A twin that knows nothing about a person can still produce a topline a client would accept. Netzer and Sambandam (2026), section 5.2
The mismatch is not peculiar to this one survey. As Netzer and Sambandam read the literature, the "75% accurate" figure most often cited in twin research rests on the same structure. In the Twin-2K-500 mega-study, twins prompted with a person's entire 500-question record hit an accuracy of 0.748. Twins given no individual information at all scored 0.734, and twins given only 14 demographic variables scored 0.746, which the paper describes as statistically indistinguishable from the full persona. Essentially the whole headline number is already available with very little data, or with none.
Those three numbers sit on top of each other because of the scale, the paper explains. On a bounded scale of five or seven points, simply guessing the midpoint for everyone lands individual-level accuracy near 0.75, since no answer can be far from the middle. In the same comparison, random guessing already establishes a floor of 0.629. What 0.748 reports, then, is not its own magnitude but its height above that floor, and the height is a bit over one point above a condition that was told nothing whatsoever.
The question, in other words, is not how accurate a twin is. It is which yardstick produced the number. The next section takes those four apart one at a time.
Four Things We Call Accuracy
The paper's second contribution is to split one word into four. Different studies measured different things and all of them called it accuracy, which is how the field ended up with a spread running from near-perfect to near-chance. At one end sits a report that a generative model's ratings of 464 moral scenarios tracked the average human judgment at 0.95. At the other end sit the 0.748 and the 0.734 from the previous section. A reader can work through the same literature and conclude that twins are 95% accurate, or 75% accurate, or barely better than chance, and all three conclusions have published evidence behind them. Much of that spread may come not from one method being better than another but from differences in what was being measured.
2.1Four Yardsticks, Three Tiers
The dividing line between the four is simple enough. Are people and twins paired up one by one, or pooled before the comparison? To that the paper adds a second axis aimed at buyers: which tier the decision you are about to make actually sits at. The table below puts both on one screen.
| Family | What it measures | Decisions this family alone can carry |
|---|---|---|
| 1. Cross-question correspondence | Whether summary values, usually means, track each other across questions, scenarios and studies | None |
| 2. Question-level distributional comparison | Whether the synthetic and human distributions overlap within one question | Tier 1 — concept screening, average willingness to pay, demand curves, average treatment effects |
| 3. Matched individual accuracy | One person's answer against that person's twin, one to one | Tier 3 — personalization, twin-as-panelist designs, respondent-level imputation |
| 4a. Within-person, across questions | The shape one person's answers make as they move across questions | Supporting |
| 4b. Within-question, across people | Who answers higher and who answers lower inside a single question | Tier 2 — segmentation, targeting, positioning, differentiated pricing |
Netzer and Sambandam (2026), section 3. Tier 2 asks for Family 4b sorting plus Family 2 comparisons computed separately within each segment.
The word "none" in the first row is the most important cell in the table. A simulator that hands every synthetic respondent the identical answer, with pure population priors and zero individuation, can still approach a cross-item correlation of 1. It only has to get the rough ordering of question means right. A Family 1 number therefore supports a Family 1 claim and nothing else, and it starts to mislead the moment the conversation turns to individuals or segments.
Information also travels in one direction only. Get the individuals right and every aggregate follows for free. Matching aggregates implies nothing at all about individuals. Which is why the paper asks buyers for one sentence: make the vendor say which tier the accuracy claim was established at.
2.2The Floor Nobody Subtracts
There is one more place where accuracy figures inflate. People asked the same question twice do not answer the same way twice. So a practice grew up of dividing twin accuracy by human test–retest reliability to produce a "relative accuracy." Acknowledging the ceiling is the right instinct. But a ratio to the ceiling hides the distance from the floor as well.
As Netzer and Sambandam set it out, the study of interview-grounded generative agents is the illustration. On held-out General Social Survey items those agents scored a raw 65.7%, which becomes 82.6% once divided by the participants' own two-week self-consistency of 79.5%. The second figure is the one that gets quoted. In the same study, agents given nothing but demographics reach 74% on that identical normalized scale. The distance to the ceiling was measured; the distance from the floor was not. The correction the paper recommends raises both ends at once: subtract the floor from the accuracy, subtract the floor from the test–retest figure, then divide. Retest data is missing more often than not, and even the floor correction on its own moves the number.
Correlation coefficients carry a related trap. A correlation is invariant to rescaling, so it says nothing about magnitude. In an example the two authors cite, a study predicting treatment effects across 70 experiments reported a correlation of 0.85 while the predicted effects ran roughly twice the true sizes, and comparing treatment arms within a single experiment dropped the correlation to 0.39. That study has since appeared in Nature, where the published version reports 70 preregistered experiments, 469 treatment effects and 119,330 participants. The direction is unchanged: high correlation, overstated effect sizes.
Set the floor and the ceiling properly and a number can also move the other way. In the paper's own empirical work, the 0.57 correlation from the attitudes condition looks thin against a perfect 1. But the highest correlation this data could in principle yield is 0.736, and 0.57 is 77% of that. Where you plant the yardstick decides whether the same value reads as 57% or as 77%.
One last trap concerns the precision attached to these numbers. With synthetic respondents, how many times you query the model is a design choice, not a sampling constraint, so conventional standard errors shrink toward zero as the query count rises. A confidence interval can be made as narrow as you like around an estimate that is simply wrong, and the paper calls this a dangerous illusion of precision. Tight error bars on a synthetic panel's results may not be telling you that the sample was large.
2.3Where People Differ Is Where Twins Break
Once the four families are separated, a pattern repeats across studies: twins that do well on averages come apart on differences between people. Brand, Israeli and Ngwe, the work Netzer and Sambandam build on, report that even a fine-tuned model approximates population means while failing to recover differences across income, gender and political groups, and sometimes inverting them. This, the paper writes, is where the gap between what vendors claim and what has been demonstrated is at its widest.
It shows up in two shapes. One is compression: twins answer within a narrower range than the humans they stand in for. The other is that accuracy itself varies by person, running higher for people who are more educated, higher-income and politically moderate. Neither pattern surfaces in a mean accuracy figure. For anyone building a targeting application this is worse than being uniformly inaccurate, because the tool is most wrong exactly where segment comparisons get made.
Why the range narrows has been measured on the training-data side. We took that mechanism apart in Bigger Models Still Wrote With Less Variety Than Their Training Data. Only the conclusion is needed here. Recovering differences between people requires grounding data about people, and there is no shortcut from a population prior to individual or segment differences.
Pull the section together and the paper arrives somewhere specific. Accuracy is not a property of the twin. It is a property of the twin-and-question pair. A twin that reconstructs someone's brand preferences well may know nothing whatever about that same person's politics. A single accuracy figure attached to a construction method is therefore unreadable unless it also states how far the evaluation questions sat from the grounding data. Leakage lands in the same spot. Widely used survey scales, textbook judgment-bias items, long-running social survey questions are all likely to be in the training corpus already, and a score earned on those inflates the score you should expect on a new question.
The next section's method follows directly from that conclusion. If the distance between a question and the grounding data governs accuracy, and if that distance can be measured before any human answers arrive, then whether this twin should be trusted with this question can be settled before the question is ever asked.
A Screen That Needs No Answer Key
The place the paper is willing to give synthetic responses is narrow and concrete. It is the question that occurs to you only after fieldwork has closed, the one the authors call the forgotten question. Going back to people means opening the field again, and the cost of that usually means the question just stays buried. Why not build twins from the answers already collected and let them fill in that one item? Plausible, and it hits a wall immediately. There is no answer key for that question, so there is no way to check whether the twins got it right.
The third contribution goes around the wall. Take what the twin produced and measure, with a random forest, how much of it can be reconstructed from the inputs the twin was given. A high value means the twin's answers are meaningfully a function of that respondent's own data. A low value means the answers do not vary systematically with the respondent data at all, in which case the model pulled its answer out of its own general knowledge, and the output should not be trusted as a simulation of these particular respondents. Not one human answer enters the calculation. It runs without ever collecting ground truth for the forgotten question.
Below is what happened when the diagnostic was applied to all 108 questions and the threshold was raised step by step, with scores recomputed on whatever survived. The second column has to be read alongside the rest.
| Screen | Questions retained | Mean correlation | Share below 0.5 | Topline error |
|---|---|---|---|---|
| No screen | 108 (100%) | 0.57 | 25.9% | 0.48 |
| Diagnostic above 0.5 | 94 (87%) | 0.59 | 18.1% | 0.48 |
| Diagnostic above 0.6 | 72 (67%) | 0.62 | 11.1% | 0.42 |
| Diagnostic above 0.7 | 46 (43%) | 0.65 | 4.3% | 0.32 |
Netzer and Sambandam (2026), Table 2. The mean correlation averages the individual-level correlations of whatever questions survived the screen; the error is the absolute deviation of the per-question mean.
The sentence "poorly answered questions fell from 25.9% to 4.3%" comes from the top and bottom rows of that table. Set the two side by side and it reads as though the twin got roughly six times better, which is not what happened. The 4.3% is a share of the 46 questions that survived out of 108. The other 62 did not disappear; they were flagged as requiring new human data. What the diagnostic did was not improve the twin but decide in advance which questions never to put to it. The paper says as much: this is a go/no-go decision about whether to pose the question at all, not a choice between two estimators.
The good thing about the table is that it prints the price next to the gain. Beside every column of rising scores sits the column where the question count falls below half. Nothing became more accurate; the region where accuracy could not be guaranteed was cut away, and the size of what got cut is the cost of the method.
The paper also states the diagnostic's own limits first. The measure is a necessary but not a sufficient condition, since a model can make good use of the data it was given and still predict the outcome incorrectly. In practice the diagnostic correlates with individual-level accuracy at 0.61. The two point the same way without measuring the same thing.
One addition is worth recording. The researchers put the same judgment to humans. Sixteen marketing researchers at TRC Insights saw 18 of the 108 questions and rated, from 0 to 100, how confident they were that a twin could recover respondents' answers. Human confidence and the machine diagnostic agreed at 0.87. The conclusion is not that the machine beat the humans. It is that each can substitute for the other when the other is expensive, and each can validate the other when the stakes warrant both. The authors attach their own caveat: with only 18 questions behind it, the expert figure is imprecisely estimated and should be taken with caution.
From the forgotten question the paper takes one more step. The same logic applies at a larger and more valuable target, which the authors call the forgotten step. Academic research and national statistics work a question through several stages in sequence. Commercial research almost never does. Budget and calendar compress a study into one stage, occasionally two. Qualitative exploration shrinks or vanishes, price ranges get set by instinct instead of by test, wording goes into the field without a pretest, and the attribute list is settled in a meeting, not from data. What gets cut is not the final study the client is paying for. It is the preparatory work that would have made that final study better.
That, the paper argues, is where synthetic responses belong. Screening twenty claim wordings down to five, bounding a plausible price range, checking whether an attribute list is reasonable, pressure-testing a questionnaire for order and framing effects. Low stakes, high repetition, and directionally right is good enough, which happens to suit a tool that gets direction right and magnitude wrong. When a pretest is off, the human study that follows catches it, so the error never reaches the decision. Framed this way, synthetic responses are not competing with human data for the same budget. They restore steps the budget had already eliminated.
One figure deserves to be carried across exactly. The abstract says the screen raises mean correlation by 15%, while the two values in the table, 0.57 and 0.65, work out to 14.0%. The paper does not account for the difference. Computation from unrounded values looks like the likely explanation, but we could not confirm it, so this article keeps the two apart and uses the table.
The Questions People Split On Fall Out
The table in the previous section says how many questions dropped out. What this article weighs most is which ones. The paper names them in one paragraph of the body and adds a line in a footnote. Read either passage on its own and you will misread it, so both go here.
Looking into the specific questions that the R² > 0.7 screen rejects, we indeed see that the least answerable items are systematically those least anchored in the attitudinal data provided (in our data, most notably the diversity, equity, and inclusion importance items, whose twin answers had low actual accuracy across categories and unreliable R² accuracy measures). Netzer and Sambandam (2026), section 4.3.1
The two outlier questions with very high R² but low correlations are questions about safety and diversity, equity, and inclusion, which may reflect the stereotyping and ideology bias reported in Peng et al. (2026). Netzer and Sambandam (2026), footnote 10
The two passages do not contradict each other. The same family of questions splits two ways. Most of it falls out at the screen, and a piece of it clears the screen with a high score and then turns out to be wrong anyway. The place where the diagnostic works least well overlaps with the place where the twin performs worst.
Why those questions in particular? Items asking how much someone values diversity, or safety, are the items on which people diverge most sharply. Two respondents of the same age, the same income and the same consumption habits in a sector can answer them in opposite directions. And that divergence is not captured well by the neighbouring attitude questions the twin was handed. The questions on which people differ most are the questions where a twin fails to reproduce the differences between them.
Compare that with a case where the method works and the contrast is sharp. When the forgotten question is, say, the importance of streaming-service affordability, the twin's answer is driven overwhelmingly by what that respondent said about adjacent attributes in the same sector, content selection and streaming quality, with demographics contributing little. That is what the method looks like when it works. It interpolates something unasked from the things already asked. Where there are no neighbours to interpolate from, it has nothing to do.
A buyer gets value out of this fact by turning it over. "We ran a pre-screen" is not grounds for reassurance; it is grounds for asking which questions failed. If the rejected list overlaps with the questions your decision actually turns on, that study needs to go back to human respondents. Whether to use synthetic responses at all comes after that.
We have watched synthetic respondents break on identity-laden questions once before, in Synthetic Respondents Given Two Identities Used Only One in August. What we covered there was a persona given more than one identity and a model that reflected only one of them. This paper's data points the same direction by a different route.
Asked in Korean, the Panel Invented Differences
Everything so far is American material, collected in English. What happens when the same exercise runs in Korean, and in a consumer domain that moves fast? A preprint posted in July took up that spot. Its two authors, based at Seoul Cyber University and Sungkyunkwan University, have not been through peer review. The manuscript is laid out in a journal template with the article identifier still blank. The numbers in this section should be read as preprint numbers.
What they tested is the Korean synthetic persona dataset NVIDIA released. The researchers conditioned two models on those personas, built a virtual panel of roughly 8,000 for each, stratified the panels by sex and age, and had them answer eight media service usage indicators. The answer key is the Korea Media Panel Survey run by the Korea Information Society Development Institute, a longitudinal study that has returned to the same households and individuals every year since 2010; its 2024 wave covered 4,006 households and 8,693 people. The paper uses the 2024 and 2025 waves.
The row to stop on is the third. It sets the error from predicting each sex-by-age cell separately against the error from doing no such thing at all, writing a single weighted grand mean into every cell.
| Condition | Gemini | EXAONE |
|---|---|---|
| Overall error (2024) | 17.1pp | 15.0pp |
| Sex-by-age cell error | 18.9pp | 15.9pp |
| Grand-mean baseline, no cells at all | 11.6pp (both models) | |
| Spread of error across cells | 52.4pp | 36.2pp |
| After age-regression calibration | 8.6pp | 6.7pp |
| Estimating directly from the real data used to calibrate | 3.6pp (both models) | |
Kim and Cho (2026) preprint, Table 7 and body text. Units are percentage points and lower is better. The grand-mean baseline writes one weighted overall mean into every cell across all eight indicators. Excluding the teen cell, where the synthetic age band maps worst onto the real population, the spread of error across cells falls to 42.2pp for Gemini and 32.6pp for EXAONE; the size and ordering of the other metrics hold.
Compare the third row against the second. The group structure the personas produced was worse than using no group structure at all. The paper's explanation is direct: personas impose incorrect variation, differences drawn from stereotype that are not in the population. This is a different order of problem from the earlier sections. Erasing differences between people loses information. Manufacturing differences that were never there persuades you that you have gained information you do not have.
Fairness requires reading the caveat alongside it. The researchers also ran a condition that stripped the narrative out of the same personas and supplied only sex and age, and the cell-level errors there came out higher, at 22.3pp and 22.8pp. The narrative attached to a persona was not contributing nothing. It simply did not contribute enough to catch the condition that drew no groups at all.
The clearest case is the indicator for AI use. Both models inflated usage the younger the respondent, and the inflation tapers smoothly down the age bands. Gemini overstated the teen band by 71 points and the band of 70 and over by 2. It is the folk belief that younger people use AI more, laid evenly across every age bracket. The error also runs the other way in places: the same model put the OTT usage rate of people in their sixties 64 points too low. On that indicator the paper records separately that synthetic endorsement collapsed when the item was phrased with the short term "OTT service," while adding its own caveat that the wording experiment had no human split-ballot arm, so it demonstrates the model's sensitivity to framing rather than any recovery of validity against a human benchmark.
Calibration brings the error down. But what calibration consumes is real survey data, and putting that same data straight into direct estimation instead brings the error down further, to 3.6pp. Which leaves the synthetic panel a role only where real data is nearly absent or a segment is entirely unobserved. Calibration does not carry forward either: applied across a one-year gap, the error climbs back to the 14-point range. The paper calls that contemporaneous error reduction rather than forward-predictive validity.
Why Korea, and why consumer behaviour? The authors do not read their result as contradicting the favourable findings on American public opinion. Political attitudes are abundantly represented in training corpora and change slowly. Fast-moving technology-adoption behaviour in a non-English market combines thinner cultural representation with rapid drift, which is precisely the condition under which stereotype priors and literal readings of question wording take over. Synthetic panel validity, on this account, is domain- and locale-contingent rather than a general property of large language models.
The dataset under test here is the one we introduced in April in 7 Million Synthetic Personas — Korea's Path to Sovereign AI. That piece presented these personas as a complement rather than a replacement, and this paper puts numbers behind the sentence. Watch one figure. The 11.6% in the April article is how much deep persona conditioning improved accuracy in a different paper (arXiv:2509.10127); the 11.6pp in this section is the error of a baseline that used no information at all. Same digits by coincidence, different metrics from different papers, and they point opposite ways.
Finally, what this result is and is not evidence for. The personas used here are not tied to any individual's actual response record. They are closer to segment-level personas than to the third type of twin the earlier sections described. Carrying these numbers over to a commercial product built on individual-level data would therefore be wrong. One claim survives: the failure the earlier taxonomy predicts, the one that arrives when segment-level personas are asked to divide a population into segments, showed up in Korean-language data. The authors have posted their analysis code and aggregate results in a public repository, though the individual-level source data sits behind an application process at the providing institution and is excluded from the package.
The Rulebook Is Already Here
Everything demanded so far may sound like complaining from outside the industry. It is not. Market research has international norms already, and in 2025 those norms admitted synthetic data as a named category. What those norms ask for, though, stops at disclosure.
The code maintained jointly by the International Chamber of Commerce and the research industry body ESOMAR functions as this sector's constitution. The previous revision was 2016, so this is the first overhaul in nine years, and the code describes it as a significant revision emphasising ethical conduct and accountability, transparency, and the necessity for human oversight. The definitions are the striking part. "Individual" was redefined to mean a human being, specifically to separate it from synthetic and virtual personas, and synthetic data and synthetic persona each received a definition of their own. Entering the definitions means entering the range of the rules.
Four obligations follow. The text below is the code's own.
| Clause | What it requires |
|---|---|
| Article 4 (a) ii | The use of a synthetic persona for data collection must be clearly notified to the data subject at the beginning of the research |
| Article 7 (e) | The client must be informed when AI or other emerging technologies are to be used, and this includes the use of synthetic data and synthetic personas. In such situations, the extent of human oversight must be stated |
| Article 7 (g) | Upon request, researchers must allow clients to arrange for independent checks on the quality of data collection and data preparation, subject to appropriate confidentiality agreements |
| Article 9 (b) | Researchers and clients must disclose whether AI, synthetic data or other emerging techniques played a significant role in sampling, deployment, analysis or interpretation, and to what extent human oversight was involved |
ICC/ESOMAR International Code, 2025 edition. The text of the code is published free of charge. The edition it replaces dates from 2016.
Line the four up and the common element shows. Every one of them says: state it. State that a synthetic persona is in use, state the extent of human oversight, disclose the role synthetic data played when publishing, permit an independent check if one is requested. A clause requiring anyone to say which of the four families an accuracy claim was established at does not appear anywhere in the code. Even the independent check in Article 7 (g) is written against the quality of data collection and data preparation, which leaves it unclear whether working out which tier an accuracy figure holds at falls inside it.
There is one further development on the standards side. A new edition of ISO 20252, the international standard setting service requirements for market research, was issued in September 2026. The 2019 edition is withdrawn and the transition period runs three years. The new edition is reported to contain provisions requiring that the stage at which AI was used be described, that clients be informed, and that versions be recorded. The standard itself is paywalled and we did not read it, so how it treats synthetic respondents is a question this article leaves open.
6.1What Is on Sale in Korea, and the Lines the Sellers Drew
Korea already has products in this category. Not one but at least three. The third column below records only what each company has published itself, and does not imply that no unpublished validation exists.
| Company | What shipped | Validation made public |
|---|---|---|
| Opensurvey | "Talk to a synthetic consumer" in DataSpace (July), Concept Studio closed beta (August) | Distribution comparison against existing data, a separate AI verification pass that filters out internally inconsistent answers, and answers labelled by whether they rest on evidence or on inference |
| Macromill Embrain | AI synthetic survey pilot (June) | Published its own pilot results comparing four study types against a human panel |
| ManyPerson | AI opinion simulation service | We could not find any |
Start with the lines these companies drew themselves. In an April write-up of its own webinar, Opensurvey listed three risks of synthetic panels and named the first one variance collapse: a large language model renders the average person convincingly while severely under-reproducing the range of answer patterns inside a real population. The compression the earlier sections described, in other words, already has a name the company gave it. The same piece says that finding uses where 80% accuracy is enough is more realistic than expecting 99%, and the product write-up defines synthetic consumers not as a replacement for a real panel but as a first-pass screening tool at the front of the research process. That is a fairly low line to draw around your own product.
Embrain published pilot results first, in June. Comparing four types against a human panel, brand awareness and recall, actual consumption behaviour, perception, and stimulus evaluation, it found meaningful gaps between human and AI answers on brand awareness and consumption behaviour, while the gap was comparatively small on stimulus evaluation such as new product concepts and naming. A company executive described the result as a supporting tool usable for particular study types rather than something that replaces all research. The company called it a pilot test and did not break the scores out by tier, so it cannot be called validation.
Nor is validation impossible in principle. It is simply that the precedents sit with a different technology. Statistical models trained on real survey data to augment respondents have several published third-party validations behind them, involving Google, L'Oréal with a French polling institute, and YouGov among others. That is a different technique from generating an individual's answers with a large language model, which is what this article is about, and what those validations confirmed was distributional comparison throughout, which is Tier 1 in the table above.
That is the factual record. The clauses, the publication of the standard, the three companies' products and public materials are all written down in documents anyone can check. One addition: we found no Korean media coverage of this paper. That means we did not find any, not that none exists.
What follows is interpretation. What is missing right now is not the rulebook but one square between the rulebook and the paper. The code asks you to disclose that synthetic data was used. The paper asks you to disclose which tier the accuracy holds at. The first can go into a contract; the second has nowhere to go yet. Korea's products arrived between April and August, and the yardstick this article borrowed was published in September. In order of arrival, the goods came first and the yardstick second. Which means the party able to put this question into a contract today is not a standards body but the buyer. What to write into the independent check that Article 7 (g) leaves open is, for the time being at least, the purchaser's call.
Why This Matters to Pebblous
Our own position first. Pebblous is a company that validates data quality. An article arguing that validation of synthetic responses is missing tilts in our favour. So the three passages below are not about a product. They are about the shape the earlier sections' structure takes when it turns up in our own work.
7.1A Passing Score for What, Exactly
What section 1 showed is that passing on one measure does not mean passing on another. The lower topline error was produced by the twin that knew no individual at all. The two numbers came from two different yardsticks, and neither of them is wrong.
Data quality work runs into that same shape constantly. The quality score attached to a dataset is usually one number covering the whole of it: missingness, label agreement, distributional fit. A good number is good about the thing the number measures. Labels can be wrong throughout a minority class and the overall agreement rate will barely move. When training fails on that class, the failure is visible inside the class and not in the headline score. Which is why there is something to ask about a score before you ask what it is. How many scores were produced, and at what unit.
7.2Inventing a Difference Costs More Than Erasing One
In the Korean measurements from section 5, the heaviest line was not the size of the error. It was that an uncalibrated synthetic panel came out behind a grand mean that used no information at all, and that the reason was personas imposing variation the population does not have. Section 5 set the two failures against each other, and the one that manufactures a difference is the expensive one, because it hands you a structure to act on.
This distinction is almost always missing from the places where synthetic data quality gets measured. Whether the distribution resembles the original is measured; how confidently the output is wrong in directions the original never went is not. If one line could be added to a synthetic data report card, it would not be how closely this resembles the source. It would be how much structure absent from the source got written into it.
7.3The Checks You Can Ask For Have Names
What this article leaves a practitioner is not a warning but four named checks. All four are written in the paper, and all four run on data the builder already holds.
- The number next to an uninformed floor. Random guessing, the scale midpoint for everyone, an empty persona, a demographics-only condition. An accuracy figure reported without at least one of these alongside it is not measuring capability.
- Sorting respondents against one another within a single question. If the decision turns on segments, this is the floor requirement.
- Accuracy broken out by group. One mean accuracy figure will not show you which group the tool fails hardest on.
- The wrong-person twin test. Rebuild each twin from a different respondent's record and compare it against the twin built from the right one. The gap is how much this product actually knows about the individual.
What matters is that none of the four asks for new data. The families do not draw on different material; they group the same material differently. So the answer that validation is expensive does not hold, at least for these four items.
The place to ask is already inside the contract. Article 7 (g) of the international code, from section 6, says that on request researchers must let clients arrange an independent check. What that clause is written against, though, is the quality of collection and preparation. Putting the four checks above onto that inspection list is still the buyer's job.
Pebblous does not have much to advance here. Roughly this: instead of attaching one score to one dataset, we have spent time on the question of what unit to score and how many scores to produce. A DataClinic diagnosis leaves per-class and per-band results alongside the dataset-level grade. That is the whole of where this article and our work touch.
The gaps deserve recording too. The text of the new ISO 20252 edition is paywalled and we could not open it, so what section 6 says about that standard rests on a secondary summary. One Korean product states a generation ceiling of around 600 people, and whether that ceiling comes from the number of source records or from compute cost is not in any public material we found. Nor did we find Korean media coverage of this paper, which means we failed to find it rather than that none exists. Thank you for reading this far.
References
The list comes in four groups. Items 1 and 2 are the two papers this article is built on, both read in full. Item 3 is the reproducibility package for one of them. Items 4 through 12 are prior work cited by those two papers; where we could not confirm the original, the entry is marked as a secondary citation, and in the body those values were lowered into a reporting frame, "as Netzer and Sambandam read the literature." From item 13 on come the normative documents, the Korean industry material, and our own earlier articles that this one connects to.
Academic
- 1.Netzer, O., & Sambandam, R. (2026). "Synthetic Data in Marketing Research: How to Evaluate and When to Trust." arXiv:2609.13995 (v1 2026-09-12, v2 2026-09-19). Read in full. arxiv.org/abs/2609.13995
- 2.Kim, H., & Cho, K. T. (2026). "Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey." arXiv:2608.28615 (v1 2026-07-23). Preprint, not peer reviewed. arxiv.org/abs/2608.28615
- 3.Reproducibility package for the paper above. github.com/howardkim1977/persona-validation-repro (MIT), Zenodo snapshot doi 10.5281/zenodo.21397425. Individual-level source data and API keys are excluded from the package.
- 4.Ashokkumar, A., Hewitt, L., Ghezae, I., & Willer, R. (2026). "Large language models can predict the results of social science experiments." Nature 656(8126), 115–122. doi 10.1038/s41586-026-10742-x. The correlation of 0.85 in the body is the value reference 1 cites; the published version's sample size is given in section 2.2.
- 5.Brand, J., Israeli, A., & Ngwe, D. (2026). "Using LLMs for Market Research." The work reference 1 builds on. [Secondary citation — original not confirmed]
- 6.Toubia, O. et al. (2025). Twin-2K-500 mega-study. Source of the uninformed floor of 0.734. [Secondary citation]
- 7.Peng et al. (2026). Under-dispersion in twin answers and systematic accuracy differences across respondent groups. [Secondary citation]
- 8.Park et al. (2026). Interview-grounded generative agents. Raw accuracy 65.7%, 82.6% once divided by test–retest consistency. [Secondary citation]
- 9.Kinzinger & Hartmann (2026). Analysis of 2.1 million twin responses, reporting that reasoning mode lifts only one family of accuracy. [Secondary citation]
- 10.Miller (2026). The wrong-person twin test and the modal-answer floor. [Secondary citation]
- 11.Mittal, V., & Tsiros, M. (2025). Source publication for the Rice C-CUBES survey (N = 3,063). [Secondary citation]
- 12.Morris et al. (2025). Industry authors concluding that synthetic data "should not be used" as a substitute for public opinion and survey data. [Secondary citation]
Norms, standards, statistics
- 13.ICC/ESOMAR (2025). International Code on Market, Opinion and Social Research and Data Analytics. Previous revision 2016. Read in full. standards.esomar.org
- 14.ESOMAR (2025). 5 Topics of Discussion to Help Buyers of Augmented Synthetic Data. Buyer guidance. [Partial — summary confirmed, PDF not read]
- 15.ESOMAR (2024). 20 Questions to Help Buyers of AI-Based Services. [Partial]
- 16.ISO 20252:2026. Market, opinion and social research — vocabulary and service requirements. Issued September 2026, 2019 edition withdrawn, three-year transition. Contains provisions on AI use. [Partial — secondary summary, standard text paywalled]
- 17.Korea Information Society Development Institute. Korea Media Panel Survey. The 2024 wave was the 15th, covering 4,006 households and 8,693 individuals. Reference 2 uses the 2024 and 2025 waves as its answer key.
- 18.Third-party validations of Fairgen-style statistical synthetic augmentation (Google, L'Oréal with IFOP, YouGov, GIM, Dig Insights). All are holdout-sample tests, and the technique differs from the large language model approach this article covers. [Partial — via vendor-published material]
Korean industry and press
- 19.Opensurvey official blog. "Can you trust respondents made by AI?" (2026-04-29). The company's own account of three risks, variance collapse, social desirability bias and absent representativeness, plus three reliability criteria. blog.opensurvey.co.kr
- 20.Opensurvey official blog. Concept Studio introduction. Built on food diary data, up to 120 questions linked per respondent, "a first-pass screening tool at the front of research rather than a replacement for a real panel."
- 21.Electronic Times (2026-04-08). Announcement of a forthcoming AI synthetic panel. "To reduce the error of an LLM generating answers that converge on the distribution mean, we concentrated on comparing distributions against existing data."
- 22.Digital Today (2026-07-09). Launch of the synthetic consumer conversation feature in DataSpace, with answers separated into evidence-backed and inferred.
- 23.AI Times (2026-09-24). Interview with Opensurvey's CTO. Three-stage personal data handling, a ceiling of roughly 600 synthetic consumers, a separate AI verification pass, and the need to build a quality evaluation framework. aitimes.com
- 24.Macromill Embrain press release (2026-06-25, Newswire). AI synthetic survey pilot test results. Source for the third column of the table in section 6.1 and for the company executive's remark quoted in the body.
Related Pebblous articles
- 25.Pebblous. "7 Million Synthetic Personas — Korea's Path to Sovereign AI." report/nemotron-personas-korea-2026-04
- 26.Pebblous. "Synthetic Respondents Given Two Identities Used Only One." report/synthetic-persona-intersectional-collapse
- 27.Pebblous. "Bigger Models Still Wrote With Less Variety Than Their Training Data." report/llm-training-data-diversity-gap
- 28.Pebblous. "You Asked People. Increasingly, AI Is Answering." blog/ai-survey-contamination-social-science