Executive Summary

This article does not ask whether AI knows etiquette. It asks whose etiquette AI holds as its default. A paper released in late September by Kimmo Eriksson's team at the Institute for Futures Studies in Stockholm took an international survey of what people in 90 societies actually said, treated it as an answer key, and handed four frontier models the same grid to fill in. The question put to the models was not "do you think this behavior is appropriate" but "what average score would people in that society have given it?"

The sharpest result came from dropping the society name altogether and asking the same questions again. The answer that came back with no society attached sat closest, in all four models, to that same model's own picture of the United States, and closer to it than to the average of all 90 societies. Supplying the name does move the answer. What it does not do is restore even half of the real distance between societies. Asking in the language the survey itself used leaves the picture largely intact. Algeria and Saudi Arabia, which were surveyed with the same Arabic instrument, stayed fused together in all four models and in both prompt languages.

The same team's previous paper carried the opposite headline. Measured on everyday American scenarios, models estimated norms better than people did, and the discussion section of that paper filed a caveat against itself: given how far American culture travels, the models might simply be unusually good at American norms. The new study is that caveat taken out to 90 societies. The question it leaves a practitioner is therefore not which model to buy. It is whether you hold a baseline that records how your own population's answers are spread, and whether you have any way to check how far a model narrows that spread.

0.43–0.49

Share of the real distance between societies that the models reproduced

Averaged over 150 scenarios, before the sampling-noise correction

0.23–0.33

How well they identified which society is the more permissive one

Survey noise alone caps even a perfect predictor at about 0.90

0.11–0.20

Distance from the no-society answer to the model's own United States

The average of all 90 societies sits further out, at 0.19 to 0.30

1.33–1.58×

Inside one society they widen the range instead of narrowing it

One model, one scale, opposite behavior depending on the axis

1

The Same 150 Questions, Put to 90 Societies

A study like this needs an answer key before it can begin. To check whether a model knows a culture you have to be holding a table that says "people in that society actually answered this," and such a table only comes from asking them. The key used here is the Global Study of Everyday Norms. The same research team surveyed 25,422 people across 90 societies between 14 July 2023 and 31 May 2024, and published the results in Communications Psychology in 2025.

What the survey asked about was scenes, not values. Fifteen behaviors crossed with ten situations give 150 scenes, and respondents rated how appropriate each scene was in their society on a six-point scale. The paper recenters that scale on zero and reports it from −2.5 to +2.5. Laughing out loud at a funeral, kissing on a bus, taking out a phone during a job interview: each is one cell of the grid. How the grid is built matters later on. The shorthand "150 everyday behaviors" gets used a lot, but what is really there is fifteen behaviors walking through ten settings.

Axis Items
15 behaviorsarguing · laughing out loud · cursing · kissing · crying · singing · talking · flirting · listening to music on headphones · reading a newspaper · bargaining · eating · resting · shouting in anger · using a phone
10 situationsfuneral · library · workplace · job interview · restaurant · park · city sidewalk · bus · cinema · party

The first twelve behaviors and all ten situations are carried over unchanged from Gelfand and colleagues' 33-nation study of 2011; resting, shouting in anger and using a phone were added. Source: Eriksson et al. (2025), Methods.

10 situations (funeral · library · workplace …) 15 behaviors (kissing · cursing …) example: funeral × laughing loudly → rated on a −2.5 to +2.5 scale 15 × 10 = 150 scenarios
▲ Original Pebblous diagram (Fig. 1 reinterpretation) — Eriksson et al. (2025), Methods. Fifteen behaviors crossed with ten situations make the 150-scenario grid that anchors this entire study.

The questionnaire went out in 41 translations and each participant answered in their own language. Putting all 150 scenes to one person invites careless answers near the end, so each respondent received a random subset. Of the 90 × 150 = 13,500 cells, 13,468 were filled, with a median of 54 responses per cell. Most of the 32 empty cells are places where the item was never fielded: the kissing item was dropped in Kuwait and Saudi Arabia, the flirting item in Kuwait.

The paper is upfront about what kind of sample this is. It is an online convenience sample rather than a nationally representative one, and the share of students averages 70% across societies. As a check, value items included in the same questionnaire were correlated against the representative samples of the World Values Survey and European Values Study. Freedom of choice correlated at 0.73 across 73 societies, gender equality at 0.74 across 70, and belief in God at 0.85 across 74. That falls short of guaranteeing representativeness, but it is evidence that the ordering of societies has not been turned upside down.

1.1Guessing the Average Score Other People Gave

The team handed the same 150 scenes to four commercial models: GPT-5 and its successor GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. Each prompt named one society and one scene and asked for the average score respondents in that society would have given, as a single number between −2.5 and +2.5. Every cell was queried five times and only the answers that came back as numbers were averaged. Generation settings such as temperature and top-p were left unspecified; the output cap was set to 2,048 tokens.

That task definition is the premise of this article. The model was not seated as a stand-in respondent. It was examined as a source of knowledge about what the average answer in a given society would be. The synthetic-respondent validation study we covered four days ago runs the other way: there the model occupied one individual respondent's seat, and the question was "how would you answer?" Even when the two draw on the same kind of survey data, they measure different abilities. Passing one is no reason to expect a pass on the other.

Collection came in two stages. An exploratory study using GPT-5 alone was preregistered on 11 September 2025; a confirmatory study asking whether the patterns seen there reappeared in other models was preregistered on 13 April 2026, after which the remaining three models were run. Both registrations sit in public repositories. That ordering, one model first and the extension afterwards, is precisely what let the project survive the data incident described in section 7.

2

What Comes Out When No Society Is Named

The experiment the team appended at the very end produces the clearest picture in the paper. They erased the society name and put the same 150 scenes again, leaving only "how appropriate is this behavior?" With no society specified, a model has nothing to work with but whatever yardstick it carries by default. Those 150 answers were bundled into a single profile, and the distance from that profile to each of the 90 society-conditioned profiles the same model had produced was measured one by one.

The nearest neighbor was the United States. It ranked first in three of the four models and second in the fourth, at distances of 0.11 to 0.20 scale points. The hypothetical midpoint formed by averaging all 90 societies sits further away, at 0.19 to 0.30. Given no name, in other words, a model is not standing at the center of the world but over next to one particular society. Canada, Australia, New Zealand and the United Kingdom were the next closest.

Distance between societies (scale points) → No society named (the model's default) United States 0.11–0.20 Canada · Australia New Zealand · UK Average of 90 societies 0.19–0.30
▲ Original Pebblous diagram (Fig. 2 reinterpretation) — Eriksson et al. (2026), summary of results. The answer given with no society named sits closer to the United States than to the hypothetical midpoint averaged across 90 societies.

Two objections attach themselves easily to this result, and the authors blocked both with numbers. The first is that the United States is simply an unremarkable society, so anything that picks no society in particular lands there. Rank the model's 90 profiles by closeness to the mean of the others and the United States comes in between 40th and 81st. It was not selected for being typical. The second objection is that the default might resemble what Americans actually answered. Rank the surveyed profiles by the same criterion and the United States comes in between 4th and 24th. In the authors' phrasing, the resemblance is specific to the default and each model's own representation of the United States. Not the real country, but the America inside the model.

2.1A Caveat the Same Team Filed in Its Previous Paper

One companion paper has to be set alongside this before the weight of the result reads correctly. The same first author published a study on 1 September titled "Large language models outperform humans at estimating society's everyday norms." Six commercial and open-weight models were given the same kind of estimation task on 555 everyday American scenarios. The answer key was ratings collected from 550 US respondents in January 2023, and 320 human estimators were seated alongside the models. The result matched the title. The two best models posted mean absolute errors of 0.78 and 0.85, the single best human 0.92, and the average human 1.77.

At the close of that paper's discussion sits the following qualification.

"This study is limited to norms in the United States. Whether similar patterns hold across cultures remains to be tested. Due to the reach of American culture, it could be that LLMs are especially good at estimating American norms." Eriksson, Karlsson, Vartanova & Strimling (2026), Discussion.

The new paper is that sentence carried out to 90 societies. And the form in which it was confirmed is one notch stronger than the caveat anticipated. The models do not merely know American norms particularly well. Withhold the name of the society and America is what comes out.

Reading the two papers together also makes clear that "AI does not know norms" is the wrong summary. On one axis the models outperformed people; on the other they failed to separate societies. Both results come from the same team using the same approach, and they do not contradict each other. They address different axes. Asking what is more inappropriate within a society is one question; asking how much more permissive that society is than another is a different one. The models are competent on the first and collapse on the second. That is why the phrase this article uses is not "does not know" but "the default keeps winning."

3

Name the Society and Half the Gap Still Goes Missing

Naming the society does move the answer. The issue is how far. The team used two measures whose names are easy to confuse, so the first step is to tell them apart. One is compression: for a given scene, take how widely the 90 survey means are spread, take how widely the model estimates are spread, and divide the second by the first. A value of 1 means the model spread its answers as far as reality; 0.5 means it covered half the distance. The other is the cross-cultural correlation, which sets the width aside and asks whether the model at least got the ordering right about which society is more permissive.

Compression came in between 0.43 and 0.49 across the four models, averaged over the 150 scenes. Once the sampling noise inside the survey itself is stripped out and the figures recomputed, they rise slightly, to 0.47 through 0.54. Which of the two you are quoting has to be stated or the argument wobbles. "Less than half" holds for all four models only on the uncorrected values; after correction one model edges just past half. Where this article says later that half the gap does not survive, it is the uncorrected figure being referred to.

The ordering side is worse. Cross-cultural correlations stayed between 0.23 and 0.33. Broken out by scene, the share falling below 0.30 runs from 49% to 66% depending on the model, and the share with compression under 1 runs from 83% to 93%. These correlations should not be read as a report card scored from 0 to 1, though. Because of noise in which respondents happened to be sampled, roughly a fifth of the between-society variance is unpredictable from the start, and a flawless predictor would top out near 0.90. Even against that ceiling, 0.23 to 0.33 is about a third of the way.

3.1The Same Model Moves the Other Way on a Different Axis

The natural conclusion at this point is that models have a habit of flattening numbers toward the middle. On a different axis, though, the same data shows the opposite. Hold one society fixed and ask how well the model orders the 150 scenes inside it, and the correlation is 0.83 to 0.87. Within that society the model does not narrow the range of its answers relative to the survey at all; it widens it, by a factor of 1.33 to 1.58.

What was asked Within one society Between societies
Accuracy of the ordering (correlation)0.83–0.870.23–0.33
Width of the answers (ratio to survey)1.33–1.58×0.43–0.49×

Ranges across the four models. The authors twice ruled out capability comparisons between models on the grounds of mismatched inference settings, so per-model values are not listed here. Source: Eriksson et al. (2026), Table 1 and Table S12.

The authors nail this point down. What they are seeing is not a general tendency to compress numerical differences. One model, on one scale, widens or narrows its range depending on which axis the question sits on. The evidence tilts toward what is missing being society-specific information rather than expressive range. The comparison against individual respondents points the same way: measure each person's own answers against their society's mean, and on average 46% to 82% of respondents miss by more than the model does. Inside a single society, the model beats most people.

3.2The Dead End in Averaging Several Models

By this point one remedy suggests itself to any practitioner. If a single model cannot be trusted, query several and average them. The companion paper mentioned above happens to have tested exactly that, and the result was discouraging. Pool the estimates of fifteen people and the mean absolute error drops sharply, from 1.77 to 0.57, because human errors are close to independent of one another: their errors correlate at 0.06 on average. Model errors resembled each other instead, correlating at 0.43. Combining all six models therefore landed at 0.73, no better than the 0.71 of the single best model.

"Organizations seeking robust social AI cannot simply ensemble multiple models, because the models share too many blind spots." Eriksson, Karlsson, Vartanova & Strimling (2026), Discussion.

That finding was obtained on American norms and cannot be transplanted onto the 90 societies of the new paper as is. The direction it points is the same, though. Averaging across models works only when the errors differ from one another. If models raised on the same web get it wrong in the same place at the same time, collecting more ballots returns the same answer.

4

The Behaviors AI Gets Right Are the Ones People Argue About Out Loud

A single average tells you nothing about what is missing. So the team split the cross-cultural correlation apart by behavior. The fifteen behaviors did not score uniformly badly. A factor of nine separated the highest from the lowest, and the ordering followed a rule.

A behavior's correlation is the average of the scene-level correlations for the ten situations it appears in. The right-hand column holds a figure the survey collected separately: participants were asked what someone objecting to the behavior would see as the main problem, and the share choosing "vulgar" was counted for each behavior.

Behavior Cross-cultural r Vulgarity marker
Kissing0.620.55
Flirting0.580.46
Cursing0.420.46
Eating0.330.16
Listening to music on headphones0.310.10
Laughing out loud0.280.16
Singing0.280.10
Shouting in anger0.230.20
Resting0.200.11
Crying0.170.07
Arguing0.160.15
Using a phone0.120.08
Bargaining0.120.13
Reading a newspaper0.110.08
Talking0.070.13

Values for GPT-5 alone, the model used in the exploratory study. The ordering of behaviors itself agreed across the four models at 0.79 to 0.90. Source: Eriksson et al. (2026), Table S3.

Kissing scores 0.62 and talking scores 0.07. As societies change, a model roughly tracks how kissing will be viewed and barely tracks how simply holding a conversation will be viewed. Line that up against the right-hand column and a rule emerges. The behaviors with good scores are the ones people experience as vulgar. At the behavior level, the correlation between accuracy and the vulgarity marker averaged 0.88 across the four models.

That the link is not a coincidence becomes visible once the other candidates are entered alongside it. The survey offered participants three concerns to choose from: that the behavior is vulgar, that it is inconsiderate of others, that it is imprudent. Only vulgarity attached to accuracy, at partial correlations of 0.43 to 0.49. Inconsiderateness ran from −0.12 to 0.14, effectively zero, and imprudence ran from −0.27 to −0.18, pointing the other way. Dropping behaviors one at a time and recomputing leaves the conclusion standing, as does clustering the fifteen behaviors for robust standard errors.

The explanation the authors attach points at the data. Cultural differences bound up with vulgarity get discussed more explicitly in text, so they leave a stronger trace in training corpora. How a given society views kissing in public is written down in travel guides, in newspaper columns, on internet forums. How acceptable it is in that society to murmur to the person next to you in a library is, as a rule, written down by nobody. Live there and you learn it; nobody writes it.

4.1Vulgarity as the Axis That Separates Societies

The baseline survey paper adds a layer here. Its conclusion was that a large part of the norm differences among the 90 societies is explained by one axis: the priority a society gives to different moral foundations. Care, fairness and liberty on one side; loyalty, authority and purity on the other. Societies weighted toward the first are more permissive overall, and more permissive in particular about behaviors regarded as vulgar, while being less permissive about behaviors that impose on others.

The same paper also calculated how far each of the three concerns depends on the setting: the share of variance explained by situation rather than by behavior.

Concern Share riding on the situation
Vulgar9%
Inconsiderate60%
Imprudent79%

Source: Eriksson et al. (2025), two-way analysis of variance over behavior × situation.

What follows is interpretation. The two facts above are written in two different papers, and the overlay is ours. The axis that generates differences between societies is vulgarity, and vulgarity is a property that barely depends on the setting. It rides almost entirely on the name of the behavior, which makes it the axis most easily transcribed into words. Inconsiderateness and imprudence, by contrast, flip their verdict according to where the act takes place, and knowledge of that kind comes from having lived in the society. That the models salvage cross-society ordering only on the vulgar behaviors fits the picture in which they picked up the portion of the differentiating axis that happens to be written down.

We carry over the caveat the authors attach to themselves: with only fifteen behaviors, the interpretation is provisional. A correlation of 0.88 computed over a list of fifteen is not guaranteed to survive at thirty behaviors, or a hundred.

5

Does Asking in the Local Language Fix It?

The obvious reaction is that the prompts were in English, so of course the results came out that way. The team had the same thought. They selected 16 societies and rebuilt the prompts in each society's survey language: Arabic for Algeria and Saudi Arabia, Portuguese for Brazil, Simplified Chinese for China, Persian for Iran, Hebrew for Israel, Spanish for Mexico, Kinyarwanda for Rwanda, and the national language in each of France, Germany, Greece, Japan, Poland, South Korea, Turkey and Vietnam. The scene list was cut to 30 and each cell was queried twice. Korea's presence in these 16 comes up again later. The paper does not report figures for Korea on its own.

The third row of the table introduces a measure the earlier sections did not use. It records, within a society, how far the model's estimate sat from that society's survey mean on average, in points of the −2.5 to +2.5 scale. On the first two rows a higher number means a better result; on the error row a lower number does.

Measure English prompt Local-language prompt
Cross-cultural correlation0.145–0.3040.268–0.427
Compression0.41–0.470.58–0.82
Within-society error (mean absolute)0.66–0.900.64–0.84

16 societies × 30 scenes; ranges across the four models. Source: Eriksson et al. (2026), Table S9.

Read the table alone and it looks like an improvement. Compression climbed as high as 0.58 to 0.82, closing 22% to 69% of the shortfall. The error row, though, changes that. The reduction in error from English to the local language runs from 0.02 to 0.11 points, and for three of the four models the 95% interval includes zero. Only one model can be said to have improved. Compression, meanwhile, stays under 1 in all four. Even asked in the local language, the models estimate smaller differences than the survey found.

A more telling check follows. The team ran a separate crossed design, prompting Brazil, China, Germany and Turkey in each of the four languages. The design is built to tell whether the extra variation produced by local-language prompting points in the direction where real differences lie. The change in the aimed slope stayed within one standard error for all four models. Switching languages does pull the answers further apart; it does not pull them apart more accurately. Reading "prompt in the local language and it gets fixed" off the compression figures alone misses this.

5.1Two Societies Sharing a Language Never Came Apart

Two societies on the list of 16 were surveyed with the same Arabic instrument: Algeria and Saudi Arabia. Their real norm profiles came out different in the survey. Measuring how much of that difference the models recover gives slopes of −0.02 to 0.10 under English prompting and −0.19 to 0.06 under Arabic. All four models, under both languages, left the two societies undifferentiated. If prompt language is doing the work of specifying a society, then two societies sharing a language get lumped into one.

Set this result next to vendor documentation and it becomes clear why the gap has gone unnoticed. In the official system card for one of the models tested here, the section on multilingual performance runs to three sentences. It reports that the MMLU exam was handed to professional translators, rendered into other languages, and scored broadly in line with previous models. The Korean row of that table reads 0.896 and 0.854. Which is to say the model answers knowledge questions correctly when asked in Korean.

The word "culture" appears nowhere in the full text of that same card. There is a separate section on fairness and bias, but the instruments it uses are English question-answering benchmarks that measure bias in American society. When a vendor says a model is good at multiple languages, the evidence behind it is a translated knowledge test. It means the model understands Korean sentences, not that it knows what Koreans consider appropriate. The cell this paper measures is precisely the one absent from the card.

A note on scope. What we opened ourselves was one version of one vendor's card. The documents of the other two vendors were not checked, so this does not generalize to "vendors do not evaluate culture." One more thing: the card's own text says the exam was translated into 13 languages while its table carries 14 language rows. The numbers disagree inside the card, so this article does not rest any argument on the language count.

6

The Authors' Own Pre-Test of Six Objections

Results of this size invite methodological suspicion. That the answer key rests on a student-heavy sample, so the models are not to blame. That the phrasing of the question produced the answer. The paper pre-tested objections of this kind from six directions. Here they are one at a time.

First, the sample skews toward students. Split the survey participants into students and non-students, average each group separately, and the two averages correlate at 0.98. They are nearly the same answers to begin with. The team nonetheless ran a separate prompt for 54 societies asking for the average university students in that society would give. Estimates from the general prompt correlated at 0.28 with the full-sample means and 0.28 with the student means, identically. The student-targeted prompt genuinely changed the estimates, correlating at only 0.73 with the general-prompt ones, and its correlation with the student means came to 0.30. From 0.28 to 0.30.

Second, ask differently and you get different answers. Beyond the original prompt, the team ran a version asking directly how appropriate the behavior is in the named society, and a version that swapped the scale for 0 to 100. Across all twelve combinations of four models and three prompts, compression stayed below 1 and the vulgarity gradient stayed positive. Third, weighting cells by respondent count, dropping cells with fewer than twenty responses, or excluding entire societies with thin per-cell samples all leave compression between 0.43 and 0.53.

6.1A Coin Flip on Twenty Years of Change

The fourth test is on the time axis. Line up Gelfand and colleagues' early-2000s survey against this one across the 26 societies and 120 scenes they share, and you have twenty years of change. Asked to estimate that change, the models got the direction right 52% of the time, with a confidence interval running from 43% to 62%. What little correlation there was came down to essentially one scene: remove listening to music on headphones and the overall correlation flips negative while directional accuracy lands at 50%. For reference, answering "more permissive" to every single scene scores 56%.

Read alongside the baseline survey paper, that score looks worse still. The survey paper reports that the twenty-year change ran in broadly the same direction across societies, with an internal consistency of 0.94 for the change pattern across 22 societies. Arguing and flirting drew more disapproval over time; crying, eating, headphone listening and newspaper reading drew less. The direction was not scattered society by society but bunched to one side, which makes this an easy problem rather than a hard one. And the result was still a coin flip.

6.2Why This Report Avoids "It Only Knows Rich-Country Norms"

The paper contains one more correlation that would make an excellent headline. Match per-society error against the UN Human Development Index and you get −0.52 to −0.77, which reads as: the more developed the society, the better the model did. The authors then recomputed the same regression with other variables entered alongside it. They added how far that society's norm profile sits from the average of the others, which they call atypicality, and how widely the society's own answers spread across scenes.

The three values below are standardized coefficients from a single regression. The outcome being predicted is the model's error for that society, and the sample is the 85 societies with an available Human Development Index. A positive coefficient means the model missed by more in societies scoring high on that variable; a negative one means the reverse.

Predictor Standardized coefficient, range over four models
Development level−0.17 to +0.15
Atypicality+0.22 to +0.51
Spread of answers−0.59 to −0.75

Development level flips sign in one of the four models (+0.15) and is not significant in another (−0.05). Source: Eriksson et al. (2026), Table S8.

What actually drives the error is spread and atypicality, not development. In the authors' words, the development gradient mostly disappeared after adjustment, and in some cases disappeared entirely. The paper also states plainly that development and spread are entangled at 0.71, so the model cannot cleanly separate the two. That is why this article does not run the summary "AI only knows the norms of rich countries." The summary rests on a single surface correlation, one the authors themselves dismantled. The section 2 finding that the default is America is far sturdier, having had two objections blocked in advance.

The same caution applies to quoting results by region. The paper divides the world into the West, Latin America, Asia and Africa and reports two separate measures, whose orderings differ. On deviation recovery, which asks how much of a society's departure from each scene's global mean the model reproduces, Asia comes out highest in all four models. On within-society accuracy, which asks how well the scenes are ordered inside a single society, the West is highest and Africa lowest. Merge the two into "AI knows Asian norms better" and you are guaranteed to be wrong. There is no way to write the sentence without naming which measure it refers to.

7

How the Paper Handled Its Own Data

The supplement carries records of two incidents. They rarely make it into coverage elsewhere, and they sit closer to practice than the result figures do, so they get their own section here.

The first is a missing space in a prompt template. When the sentence was assembled, the gap between "the appropriateness of" and the scene text was dropped, so throughout the main collection what reached the models was the string "the appropriateness ofargue at a party". All four models received the same template, so comparisons between models still hold. The team nonetheless re-collected 26 societies and 120 scenes twice per cell with the space restored, and confirmed the conclusions did not change.

The second is heavier. After the confirmatory study had been preregistered, a parsing error turned up in the stored file from the exploratory run. Of 67,340 calls, only 15.6% had the rating landing correctly in its column. The grid for change over time was worse still, at 20.0%. The numbers the hypotheses had been built on were standing on bad values.

Recovery was possible for one reason. The full raw response returned by the model had been preserved separately for every call. Of the 10,503 calls where both records existed, all 10,503 agreed. Fixing the parsing rule allowed the whole set to be re-read from the raw responses. Because the preregistered hypotheses had not specified particular GPT-5 values, the decision rules stayed in force as written.

How far the before and after diverged is recorded in the supplement across three measures. Every figure this article has quoted so far is a post-recovery figure.

Measure With the parsing error After recovery
Cross-cultural correlation0.2090.27
Compression0.580.49
Vulgarity partial correlation0.4640.52

Exploratory study, GPT-5. All three measures moved, and compression moved downward, in the direction that strengthens the conclusion. Source: Eriksson et al. (2026), supplementary material.

The remaining noise in the collection is written out in numbers too. Refusals numbered 195 calls, 0.24% of the 82,935 total, and all came from a single model. For the 84 calls that returned empty, the team disclosed that the three retries specified in the preregistration were not performed. One model registered to run without reasoning returned reasoning tokens on every call despite its budget being set to the minimum, and that too is logged as a deviation.

The point of this section is not to fault the team. It is that derived columns can be wrong, and that nothing was irreversible because the raw responses were never discarded. A column can sit quietly empty and the averages will still compute and the charts will still render. A correlation of 0.209 from a column that was 15.6% full and a correlation of 0.27 after full recovery are both plausible-looking numbers, so no amount of staring at the value will tell you which one is the accident. What caught the incident was a check, and what made it fixable was retention.

8

Why This Matters to Pebblous

Our interest should be declared first. Pebblous is a company that diagnoses data quality. A story about training data failing to hold the world evenly leans our way. So the four passages below are not a product pitch. They are about the shapes the structure shown in the preceding sections takes when it reappears in our own work.

8.1Flattening Arrives Wearing the Face of a Normal Output

The failure this paper caught is not "I don't know." The models produced a number every time, and each of those numbers looked reasonable on its own. What had happened is that the gaps between societies had shrunk to about half their real size. In data quality work this is the most awkward failure shape there is. Missing values are conspicuous and outliers set off alarms, but flattening arrives wearing the face of a normal output. On top of that, inside a single society the models widen their range by nearly half again, so from the output alone the model looks more confident rather than less.

Catching a confidently flattened value means looking at variance by axis rather than at the result. Reduce what this paper did to one line and it is not that the models were tested harder. The yardstick was changed. Spread was measured in place of hit rate, and that spread was split into within-society and between-society. That the two came out in opposite directions from the same outputs is something a single score can never reveal.

8.2A Model Knows the Writing About the World, Not the World

The vulgarity gradient in section 4 is a story about a corpus, not about a model. If, as the authors suggest, differences bound up with vulgarity leave a more explicit trace in text, then the order in which a model knows things follows the order in which they were written down, not the order in which they matter. Our own claim that AI-Ready Data is a question of which axis you selected on rather than of volume sits at exactly that point. Just as DataClinic diagnoses the class distribution and the outliers of a dataset, the object of diagnosis here is the distribution along the axis called society.

We add the language distribution of the web only as background. In the most recent public crawl statistics, English accounts for 41.86% and Korean for 0.84%. Korean speakers are also somewhere around 1% of world population, though, so reading that figure as under-representation relative to population does not hold up. What to read here is the absolute quantity rather than the share. When the total volume of text written in a language is small, the slice of it given over to discussing norms is smaller still. Wikipedia article counts point the same way, at roughly 7.25 million in English against roughly 770,000 in Korean, a factor of 9.5.

8.3Foreign Models Already Occupy Those Seats

Korea is a useful worked example here, and the shape generalizes to any market that imports its models. Foreign models already decide "is this behavior appropriate" for Korean users in plenty of places: content moderation, the tone of a support reply, the enforcement of community rules, the action a robot or a vehicle selects. How deep that dependence runs varies with what is being counted, so each figure is written with the name of its measure attached. Counted by users, a Ministry of Science and ICT survey of 2,500 adults in the fourth quarter of 2025 found 81.9% using one of two foreign services. Counted by what services are built on, reporting puts 97% of domestic AI services on foreign models. The two numbers have different denominators and different methods and must not be merged into one.

There is also a figure for how far Korea sits from the default. The baseline survey team grafted the norm-tightness scale from Gelfand (2011) onto their society-level dataset, and the Korean row reads 4.32. The tightness measure the survey itself collected sits in the same dataset at 4.6, on a different scale, so the two cannot be mixed. The provenance grade goes in as well: this value was read not from the body of a paper but from a file in the view-only repository the authors released for reproduction, and that repository is marked for public release upon acceptance.

8.4The Defect Only Shows Against a Korean Baseline

It is not that culture-related evaluation does not exist. Lay out four of the instruments in common use today and some include Korea, while one was built for Korea specifically.

Instrument Size and coverage Task format
NormAdEtiquette situations from 75 countriesAppropriate / inappropriate judgment
CulturalBench1,696 questions, 45 regions, 17 topicsMultiple choice
BLEnD52.6k question-answer pairs, 16 countries and regions, 13 languages (both South and North Korea included)Short answer + multiple choice
KorNATKorea-specific. 4K social value items + 6K common knowledge items, keyed to a survey of 6,174 KoreansMultiple choice

Sizes and task formats as stated in each benchmark's own paper. Sources: Rao et al. (2025), Chiu et al. (2025), BLEnD (NeurIPS 2024 D&B), Lee et al. (2024).

What follows is interpretation. The task formats in the table are facts stated in each paper; the conclusion drawn from them is ours. All four grade a question that has a right answer as either hit or miss. The failure this paper caught, however, is a question of width rather than of right answers. A model that puts the societies in roughly the right order while halving the intervals between them loses no points under any of these four instruments. If anything, crowding answers toward the mean improves a multiple-choice score. The tools in use for cultural evaluation cannot, as a matter of construction, see this defect.

A Korea-specific instrument existing does not fill the gap either. KorNAT is an alignment benchmark keyed to the responses of 6,174 Koreans, and that alone makes it a rare asset. What it measures, though, is the hit rate on value options, not the distribution of appropriateness across situations. Different grid, different task. So the empty cell is neither the model nor the presence of Korean-language data. It is a baseline recording how our own population's answers are spread, and a procedure that measures that spread axis by axis.

How to build one has already been laid out by this paper as a blueprint. Assemble a shared item grid from the situations a service actually has to judge, attach a baseline of human responses, measure the between-society and within-society variance ratios separately, and keep a condition with the society name removed as a control. None of the four steps requires building a new model, and the last of the four is remarkably cheap. Asking once more with the name left out costs almost nothing, and it shows immediately where your service's default is standing.

What we could not check is worth recording too. The paper reports no individual figures for Korea anywhere. Two figures state in their captions that societies with large residuals are labeled, but those labels live inside the images and we could not read them, so whether Korea is among them is unknown to us. We did not open the system cards of the other two vendors, and we did not investigate this time which clauses of international standards and norm documents require evaluation of cultural or regional bias. Finally, training-data imbalance is the explanation the authors put forward, not a conclusion the paper established; the text says outright that their analysis cannot distinguish an absence of information from a prompt failing to elicit it. Thank you for reading this far.

R

References

The list comes in three groups. Items 1 through 3 are the three papers this article is built on, all read in full. Items 4 through 13 are prior work on cultural evaluation; among them, items 8 through 12 were verified bibliographically only, with their conclusions left unread, and no conclusion of theirs is cited in the body. From item 14 onward come primary documents and statistics, together with our own earlier articles that this one continues.

Academic

  • 1.Eriksson, K., Vartanova, I., & Strimling, P. (2026). "Large language models underestimate and partly misrepresent cultural variation in everyday norms." arXiv:2609.30896v1 (2026-09-25), CC BY 4.0. Primary source for this article. Read in full. arxiv.org/abs/2609.30896
  • 2.Eriksson, K., Karlsson, N., Vartanova, I., & Strimling, P. (2026). "Large language models outperform humans at estimating society's everyday norms." Communications AI & Computing 1(1):15. doi 10.1038/s44488-026-00018-8. The contrast case in sections 2 and 3. Read in full.
  • 3.Eriksson, K., Strimling, P., Vartanova, I., Simpson, B., Persson, C., et al. (2025). "Everyday norms have become more permissive over time and vary across cultures." Communications Psychology 3(1):66. doi 10.1038/s44271-025-00324-4. The original paper for the Global Study of Everyday Norms, which serves as the answer key here. Preregistration osf.io/qz82x. Read in full.
  • 4.Rao, A., Yerukola, A., Shah, V., Reinecke, K., & Sap, M. (2025). "NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models." NAACL 2025, 2373–2403.
  • 5.Chiu, Y. Y., et al. (2025). "CulturalBench: A Robust, Diverse, and Challenging Benchmark for Measuring LLMs' Cultural Knowledge." 1,696 questions, 45 regions.
  • 6."BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages." NeurIPS 2024 Datasets & Benchmarks Track. 16 countries and regions, 13 languages, 52.6k question-answer pairs.
  • 7.Lee, J., Kim, M., Kim, S., Kim, J., Won, S., Lee, H., & Choi, E. (2024). "KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge." Findings of ACL 2024, 11177–11213. arXiv:2402.13605.
  • 8.Ramezani, A., & Xu, Y. (2023). "Knowledge of cultural moral norms in large language models." ACL 2023, 428–446. [Bibliographic check only — conclusions not reviewed]
  • 9.Tao, Y., Viberg, O., Baker, R. S., & Kizilcec, R. F. (2024). "Cultural bias and cultural alignment of large language models." PNAS Nexus 3(9), pgae346. [Bibliographic check only]
  • 10.Zhao, W., et al. (2024). "WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models." LREC-COLING 2024, 17696–17706. [Bibliographic check only]
  • 11.Bulté, B., & Rigouts Terryn, A. (2026). "LLMs and cultural values: the role of prompt language and explicit cultural framing." Computational Linguistics 52(2), 407–494. [Bibliographic check only]
  • 12.Adilazuarda, M. F., et al. (2024). "Towards Measuring and Modeling 'Culture' in LLMs: A Survey." EMNLP 2024, 15763–15784. The survey cited by reference 1 as grounds for the vulgarity explanation. [Bibliographic check only]
  • 13.Gelfand, M. J., et al. (2011). "Differences between tight and loose cultures: a 33-nation study." Science 332(6033), 1100–1104. Source of the twelve behaviors, the ten situations, and the early-2000s comparison data. The original tables sit behind a paywall and could not be opened.

Primary Documents and Statistics

  • 14.OpenAI (2025). GPT-5 System Card, §3.11 Multilingual Performance and Table 11. Also distributed as arXiv:2601.03267v2. The Korean scores and the absence of the word "culture" in section 5 were confirmed by extracting the full text of this document as plain text.
  • 15.Common Crawl Foundation. Crawl language statistics, CC-MAIN-2026-39 (CLD2 language detection). commoncrawl.github.io
  • 16.Wikimedia Foundation. List of Wikipedias. Article counts for the English and Korean editions. meta.wikimedia.org
  • 17.Authors' reproduction repository, osf.io/jzhqu (view-only link, marked for public release upon acceptance). The norm-tightness value in section 8 was read from the GSEN society-level aggregate file deposited there. Two preregistrations: osf.io/d7km9 (2025-09-11, GPT-5 exploratory) and osf.io/w3fxd (2026-04-13, three-model confirmatory).
  • 18.Korean Ministry of Science and ICT, Q4 2025 survey (2,500 respondents aged 19–69), and reporting on the domestic AI industry's reliance on foreign models (Ajou Economy, 2026-09-19). Sources for the two figures in section 8.3, which differ in denominator and in method.

Related Pebblous Articles