Executive Summary
The most common way to check whether an AI treats people fairly is to swap the name and nothing else. Same dossier, one name replaced, and you watch whether the model's answer moves. A paper released by University of Alberta researchers on 28 September 2026 argues that the premise of that method breaks before the model ever sees the text. Some names enter as a single whole chunk. Others are assembled from two or three fragments. This article traces where that divide opens and how far it travels, using the paper's numbers and our own measurements.
The divide is not scattered evenly. Sorted by race and gender association, the share of names that enter whole runs about 5.4 times higher in the top group than in the bottom one. With name frequency and name length in the regression alongside it, the group terms still survive. Yet measure the same signal at the output logits instead of a middle layer, and the sign flips on three of four tasks. So what this research actually pins down is not that fragmented names get penalized. It is that a bias audit which never says where it measured cannot say what it measured.
Every name in the study came from a Florida voter file. Not one Korean name is in there. When we ran Korean names through the same yardstick, they had no place to stand on either side of the paper's contrast. Sections 1 through 4 follow what the paper measured and what it declined to claim. The Korean measurements in section 5 are ours and appear nowhere in the paper, and section 6, where the argument moves to audit design, is this article's reading of a principle the paper states in its appendix.
Share of names that enter whole, lowest group vs highest
Across the 7,469-name controlled sample. The two ends sit about 5.4 times apart
Tasks whose gap flips or vanishes at the output layer
Fellowship attenuates to zero; hiring and clinical triage run the other way
Names that enter whole in every one of the 12
4,052 of 497,583. Names atomic in at least one tokenizer come to 4.6%
Korean given names entering whole, Hangul vs romanized
Our measurement. A Korean-specialized model's advantage does not survive romanization
The Illusion of Holding the Name Constant
Put two names side by side. Emily and Emilee. To a human eye they differ by one letter, and both are common English-language given names for women. At a model's front door, though, they arrive in different shapes. Prefix a single space and encode with the Qwen3 tokenizer: Emily goes in as one ID, Emilee as two IDs assembled together.
| Name | Tokens | Pieces that actually go in |
|---|---|---|
| Emily | 1 | ␣Emily |
| Emilee | 2 | ␣Emilee |
| Jamal | 1 | ␣Jamal |
| Latoya | 2 | ␣Latoya |
| DaShawn | 3 | ␣DaShawn |
Table 1. Encoded directly with the Qwen3-4B tokenizer. This table is not from the paper; we downloaded the vocabulary file and measured it. The criterion matches the paper's section 3 — a surface form with a leading space that goes in as one token, and decodes back to exactly the original characters, counts as atomic. ␣ marks the leading space.
Name-swap experiments have a long lineage. In 2004 Marianne Bertrand and Sendhil Mullainathan mailed 4,870 fictitious résumés to job ads in Boston and Chicago. Experience and education were matched across pairs; only the name at the top changed. Résumés carrying White-associated names drew callbacks 9.65% of the time, and those carrying African American-associated names 6.45%. A gap of 3.20 percentage points, roughly 50%. One callback took about ten résumés on the first side and about fifteen on the other.
That design rests on exactly one thing: the premise that nothing but the name changed. On paper, mailed to human readers, the premise holds. To a hiring manager, Emily and Lakisha are equally a single word, and neither costs more to read. So when the outcomes diverged, the divergence could be laid at the name's door.
Carry the same logic over to a language model and a gap opens. The model never sees letters. It cuts text into pieces listed in a vocabulary, turns those pieces into a sequence of IDs, and receives the IDs. The cutting machine is the tokenizer, and which pieces make the vocabulary is decided by whoever built the model, based on the training corpus. The table above is what that decision looks like in practice. Emily is already in the vocabulary; Emilee is not, so it gets assembled from Em and ilee.
Two prompts that differ only in the name are therefore, from the model's side, inputs of different length and different composition. One name already owns a slot in the vocabulary; the other borrows that slot in fragments. The question Mir Tafseer Nayeem and Davood Rafiei of the University of Alberta put forward starts here. Is there a pattern in who gets the slot and who does not, and does the difference survive into the model's interior?
Which Names Get a Token of Their Own
The paper's first question is a plain one. Which names does a vocabulary already contain, and which get assembled from fragments? Answering it takes a list of names and a list of tokenizers, and then filling in every cell of the product. The work is laborious without being hard. What emerged as the scale grew was regularity.
2.1The Population Is a Florida Voter File
The names come from a June 2022 extract of Florida voter registrations. Merged with state newborn-name records for support and normalized, the set holds 534,509 entries; dropping every multi-word form leaves 497,583. That is the number the hook and the summary call "about 500,000." Of those, 414,493 carry race and gender association metadata, and a controlled sample filtered on frequency and dominant-group share comes to 7,469. The fine-grained comparison later in the paper narrows again, to 200 pairs drawn from inside that sample.
The population shrinks at every analytical step. Half a million is the count at stage one, and the count behind the group comparisons is two orders of magnitude smaller. The nature of the population carries through as well — a voter file from one U.S. state. The paper says in appendix F that it restricted itself to single-token names in strict ASCII, and calls other scripts, languages and cultures a natural extension. Section 5 fills in one cell of that extension.
2.2Twelve Rows, Eight Distinct Designs
There are twelve tokenizers, but not twelve designs. The GPT-4 family, OLMo-3 and Phi-4 share one access pattern; GPT-5 and the two sizes of gpt-oss share another. Counted by distinct design, there are eight. That is why identical values repeat across rows in the table below.
The two columns measure different things. The left counts names that become a single token in any casing or spacing. The right counts only the form a person's name actually occupies in text — a leading space and an initial capital. The most generous tokenizer, Aya-Expanse, takes in 20,020 names; the stingiest, DeepSeek-V3.2, takes 4,998, a spread of roughly four times. Even the generous end amounts to 4% of 497,583. Names that enter whole are a minority in every tokenizer. The group comparison asks who that narrow share goes to.
| Tokenizer | Atomic in any form | Leading space + capitalized |
|---|---|---|
| Aya-Expanse-32B | 20,020 | 14,528 |
| Gemma-3-27B | 16,371 | 10,108 |
| GPT-5 / gpt-oss-120B / gpt-oss-20B | 13,575 | 6,769 |
| Ministral-14B | 11,357 | 6,439 |
| Llama-3.1-70B | 9,131 | 5,148 |
| Qwen3-32B | 8,716 | 5,059 |
| GPT-4 / OLMo-3-32B / Phi-4 | 8,685 | 5,051 |
| DeepSeek-V3.2 | 4,998 | 0 |
Table 2. How many of the 497,583 names each tokenizer takes as a single token (paper, appendix Table 11). Tokenizers with identical values are merged into one row.
The bottom row catches the eye. In DeepSeek-V3.2, not one name enters whole in the leading-space, initial-capital form. Counting lowercase and other casings, 4,998 names do appear, so this is not an inability to hold names — it is an inability to hold them in the exact form a name takes in running text. That is a design choice about casing, not a defect. For anyone running a bias audit on top of this model, though, the choice becomes a condition of the audit.
A bigger vocabulary does not fix it either. DeepSeek-V3.2 and Llama-3.1 both carry 128,000-entry vocabularies, yet they take 4,998 and 9,131 names whole — nearly a factor of two apart. What goes into a vocabulary is settled by the training corpus and the merge rules, not by its size. Stacked one on top of another, the twelve yield two totals: 23,095 names, or 4.6% of the total, are atomic in at least one tokenizer, and 4,052 names, or 0.81%, are atomic in all twelve.
2.3The Gap Across Groups Runs Past Fivefold
The paper does not classify names by race or gender. It tabulates, from voter records, which demographic group a name surface is statistically associated with, and its ethics statement nails the distinction down: a name does not determine anyone's race, gender or ability. So the groups below should be read as name surfaces classified as NH Black-associated, name surfaces classified as NH White-associated, and so on. NH is the U.S. census term for non-Hispanic.
| Race/ethnicity association | Gender association | Names | Atomic in ≥1 | Atomic in all 12 |
|---|---|---|---|---|
| NH White | Male | 1,433 | 64.8% | 14.9% |
| Asian/PI | Male | 286 | 57.3% | 10.8% |
| Hispanic | Male | 637 | 37.5% | 3.6% |
| Asian/PI | Female | 289 | 35.6% | 6.9% |
| NH White | Female | 2,100 | 35.1% | 5.0% |
| NH Black | Male | 648 | 25.2% | 2.9% |
| Hispanic | Female | 1,197 | 16.7% | 2.5% |
| NH Black | Female | 879 | 12.1% | 1.4% |
Table 3. The 7,469-name controlled sample counted by intersectional stratum (paper, appendix Table 12). Top and bottom sit about 5.4 times apart. These are aggregate associations of name surfaces, not demographic markers of individuals.
With the strata collapsed, the direction gets sharper. Male-associated names come in at 49.8%, female-associated at 25.7%. NH White-associated names reach 47.2%, NH Black-associated 17.6%. The ordering holds tokenizer by tokenizer: male-associated ahead everywhere, NH White-associated ahead of NH Black- and Hispanic-associated everywhere. Eight distinct designs, all leaning the same way.
2.4Not Because They Are Common, Not Because They Are Short
Two properties look like enough of an explanation. Common names are likelier to land in a vocabulary, short names are likelier to fit in one piece, and both differ across groups, so the outcome differs too. The paper put both factors straight into a regression to find out.
| Predictor | Odds ratio | 95% CI |
|---|---|---|
| Male-associated vs female-associated | 3.36 | 2.98 – 3.78 |
| Log name frequency (+1 SD) | 2.94 | 2.74 – 3.15 |
| Asian/PI-associated vs NH White-associated | 1.55 | 1.26 – 1.91 |
| Name length (+1 SD) | 0.51 | 0.48 – 0.55 |
| Hispanic-associated vs NH White-associated | 0.47 | 0.41 – 0.55 |
| NH Black-associated vs NH White-associated | 0.42 | 0.36 – 0.50 |
Table 4. Logistic regression predicting the odds of entering atomically (paper, Table 1). Frequency and length are in the model together.
Frequency and length behave exactly as expected. Common names raise the odds; long names lower them. But with both in the model, the group terms are still there. Male-associated versus female-associated sits at 3.36, and NH Black-associated versus NH White-associated at 0.42. This does not make frequency irrelevant. It says only that frequency alone does not account for the pattern. Something beyond frequency and length shapes which names a vocabulary holds whole, and the paper declines to say what that something is.
How Far the Split Travels Inside the Model
A divide at the front door means nothing on its own. If an assembled name still earns its full share of representation inside the model, the entry difference is a passing detail. That is the paper's second question, and the authors answer it with a procedure they call NameTrace — not an existing benchmark but a framework introduced here.
3.1Pairing Atomic Names With Split Ones
Names are first sorted into two pools. A name that enters as one token in all three of Qwen3-4B, Llama-3.1-8B and Ministral-3-3B is atomic; a name that is atomic in none of the three and enters in two or three tokens is mildly split. Anything at four tokens or more was excluded. The paper says the point was to compare against ordinary assembly, not against extreme failure cases.
Then comes the heart of the design. One name is drawn from each pool to form a pair, and pairs are built only within the same race-and-gender stratum. Frequency, character count, strength of demographic association, metadata reliability and orthographic cues are matched on top of that. In principle, one difference survives inside a pair: one name enters whole, the other enters assembled. Of the 200 pairs built this way, half went to tuning the design and the remaining 100 were held out, never opened until evaluation.
3.2What the Four Tasks Actually Asked
There are four tasks. Summaries tend to shorten them to scholarships, hiring, healthcare and lending, but the settings the paper actually used are narrower. A graduate fellowship review asks about promise for a single slot; a software engineering internship screen asks about competence for one final-interview slot; a clinical priority judgment on a simulated patient note asks about concern for one immediate-care slot; and a small business loan review asks about trustworthiness for one approval. The clinical task runs on simulation notes, not real medical records.
[role] / Candidate: [name] / Common dossier: [evidence] / In one adjective, [name] seems The prompt skeleton from appendix Table 6. Both names in a pair receive exactly the same dossier, and the score is read at the position right after that final "seems."
The value being read is a sum of the model's probabilities over a set of adjectives, each weighted by a sentiment lexicon. It is a probability-weighted semantic score rather than a probability, so it is not bounded between 0 and 1. The 0.131 in the table below is not 13.1%. The clinical task also runs on a flipped axis: "worried" and "ill" are the direction the task targets and "healthy" is the opposite pole, so a higher clinical score does not mean a more favorable judgment — it means greater concern.
3.3The Gap Holds, but Unevenly
The measurement point is one middle layer per model, chosen on the tuning pairs and locked before evaluation. At that layer, the held-out 100 pairs put the atomic side higher on all four tasks.
| Task / axis | Weighted gap (atomic − split) | 95% CI |
|---|---|---|
| Fellowship / promise | 0.131 | 0.096 – 0.168 |
| Internship screen / competence | 0.072 | 0.055 – 0.091 |
| Clinical priority / concern | 0.059 | 0.048 – 0.072 |
| Lending / trustworthiness | 0.051 | 0.041 – 0.061 |
Table 5. Middle-layer gaps on the held-out 100 pairs (paper, Table 2). These are probability-weighted semantic scores, not probabilities.
More important than four positive rows is the unevenness underneath them. In the Qwen family all four axes are positive and large, from 0.154 to 0.366. Llama too, but at 0.002 to 0.026 — far smaller. Ministral is positive on the internship screen and the clinical task, near zero on fellowship, and negative on lending. Across all eight models, 23 of the 32 model-by-task combinations have a positive mean, with 20 positive across the confidence interval as well. So "confirmed in every model" is not a sentence anyone can write here. Closer to it: the gap appears across several families, at different magnitudes, and sometimes in different directions.
Stratum by stratum it is uneven too. Three primary models times eight strata gives 24 cells per task. The internship screen is positive in all 24, the clinical task in 23, fellowship in 22, and lending in only 15 — the weakest of the four. The NH Black-associated stratum carries the largest racial-group mean on each of the four tasks. That number must not be read as NH Black-associated names being disadvantaged. The comparison lives entirely inside a stratum, so what it says is that within that stratum, the distance between names that entered whole and names that were assembled is the widest. By gender, the fellowship gap in the female-associated stratum is 0.198 against 0.065 in the male-associated one, roughly three times as wide.
3.4It Carries Over to Unseen Names
If the gap were an accident of particular names, it would not reproduce on other names. The paper checks this by subtracting the mean gap measured on the tuning names. The fellowship gap falls from 0.131 to 0.004, the internship screen from 0.072 to 0.008. Across all 36 combinations, the correlation between tuning values and held-out gaps exceeds 0.99, and signs agree 94.4% of the time. A different set of names produces the same tilt, at the same size, in the same place — which is the basis for calling this gap a systematic structure.
The paper goes one step further and edits the hidden states directly. Adding a task direction vector at the atomic name's position and subtracting it at the split name's position does move which of two candidates gets chosen. This is an artificial intervention, though. The paper itself says it should be read only as evidence that the direction is connected to downstream computation, not as evidence that the model ordinarily decides that way.
At the Output Layer, the Sign Flips
Stop reading here and a comfortable conclusion follows: when a name is split, the model thinks less of the person. The paper blocks that conclusion itself. One more table sits in the appendix, and it is the most important scene in the research.
4.1Reading the Same Score at the Last Step
Every figure in section 3 was read at a middle layer. What happens if the same adjective score, at the same prediction position, is read at the output logits — the point just before the model emits a word — instead? The paper ran that calculation and set the two side by side.
| Task / axis | Middle layer | Output logits | What happened |
|---|---|---|---|
| Fellowship / promise | 0.131 | −0.027 | CI covers 0, effect vanishes |
| Internship screen / competence | 0.072 | −0.100 | Sign reverses |
| Clinical priority / concern | 0.059 | −0.213 | Reverses, and widest of the four |
| Lending / trustworthiness | 0.051 | +0.081 | Holds, and grows |
Table 6. The same quantities re-measured with only the reading point changed (paper, appendix Table 21). The caption to the paper's Figure 10 reads: "At the output logits, fellowship attenuates, hiring and clinical assessment reverse, and lending remains positive."
One thing changed: where the reading was taken. Same names, same prompts, same models. And on three of four tasks the conclusion changes with it. An article written off section 3 alone would say "split names get penalized"; an article written with Table 6 in view says the opposite.
Nor do these flips barely clear zero. Clinical priority moves from 0.059 at the middle layer to −0.213 at the output, more than tripling in width. Its confidence interval runs −0.278 to −0.150, and the internship screen's −0.100 runs −0.150 to −0.060. Neither interval covers zero. Only fellowship fades out; the other two cross to the far side and widen again once they are there.
4.2Interpretability Research Has Reported This Three Ways
It is tempting to read the reversal as a flaw in the paper — a signal visible only at the middle layer must be an artifact. It is not. Within what we were able to verify, this shape is a property that interpretability research has reported along three separate lines.
First, being able to decode information from activations is a different claim from the model using that information. Probes often recover an attribute at high accuracy, yet erasing the attribute and rerunning the task leaves performance intact. Elazar and colleagues formalized that procedure as amnesic probing, and Ravfogel and colleagues found that relative-clause boundaries decode above 90% at every layer while only the middle layers causally affect behavior.
Second, late layers have been caught suppressing a correct intermediate representation. A study that dissected character-counting tasks showed that in the Llama, Qwen and Gemma families the model computes the right answer internally and then fails to express it at the output. In LLaMA3.2-3B, a single final-layer MLP suppresses roughly 7.5% of the correct answer's probability mass.
Third, there is even a mechanical account of sign reversal. It is called copy suppression: one attention head in GPT-2 Small pushes down a prediction when the token an upstream component promoted has already appeared in context. Geometrically it is a literal sign flip. The head writes to the residual stream in the opposite direction from that token's embedding, and most of that effect travels the direct path to the final logits. Which is to say the reversal happens at the last reading point.
A signal that splits at the input and does not travel unchanged to the output is not an absent signal. It does mean that a bias audit which never says where it measured cannot say what it measured. The paper calls itself a pre-behavioral study and writes this in its limitations: "Intermediate accessibility can be preserved, attenuated, or redirected by later computation."
4.3The Tokenizer Untouched, the Signal Reshuffled
If the measurement point changes the conclusion, then surely fixing the point once solves it. That does not work either, and the paper's side-by-side measurement of base and post-trained models shows why. The tokenizer was never touched, yet post-training splits the signal three ways. Llama preserves it almost intact, Qwen relocates the values, and Gemma reorganizes the whole thing.
Concretely: the clinical gap in base Qwen3-4B is 1.539, and in the post-trained edition it inverts to −0.056. In Gemma-3-4B, the layer where the signal must be read moves from layer 2 to layer 25. Same names, same tokenizer, and both the reading position and the value shift together. Update the model once and the control you established beforehand breaks.
Nowhere to Place a Korean Name
From here on, these numbers are not in the paper. All ~497,000 names the paper used came from a Florida voter file, and not a single Korean name is among them. Appendix F leaves the door open, calling other scripts and languages a natural extension. What follows is our own look through that door. These figures must not be mixed into the same sentence as the paper's results.
5.1The Paper's Yardstick, Official Name Statistics Only
The criterion matches the paper's section 3. A surface form with a leading space that enters as one token, and decodes back to exactly the original characters, counts as atomic. Checking whether a string merely exists in the vocabulary file yields an upper bound, so we did not use it; we encoded every name in the twelve tokenizers and counted.
No name was chosen by guesswork. The 9 surnames come from the 2015 Population and Housing Census surname-and-clan volume published by Statistics Korea; the 47 given names from four newborn-name ranking datasets in the Supreme Court's electronic family register system; and the 20-name U.S. control group from the Social Security Administration's 2024 release. Adding Hangul originals and romanization variants gave 441 surface forms in total. Three gated repositories were obtained from public mirrors and checked against the originals by vocabulary size.
Two limits up front. The 20-name U.S. control group is a sample we drew ourselves, not the paper's ~497,000. And 47 given names is a small sample, so the percentages below should not be trusted to the decimal. The direction, however, holds for the nine surnames and for nearly all forty-seven given names.
5.2Surnames Arrive Whole, Given Names Break Apart
Of the 56 Hangul originals (9 surnames plus 47 given names), 10 entered whole in at least one of the twelve tokenizers. Nine of those 10 are surnames. Exactly one given name out of 47 was atomic — "우주" (Uju) — and most likely because a common noun with the same spelling sits in the vocabulary. The atomic-versus-split contrast the paper built splits, in Korean, along the seam between family name and given name.
Trying to seat Korean names in the paper's two pools runs into trouble. A pair needs a name that enters whole on one side, and on the Korean side there is effectively no name to put there. The comparison is not tilted; the material to build the comparison is missing.
5.3Romanization Does Not Close the Gap
On overseas applications, on passports and in international hiring systems, a Korean name almost always arrives romanized. So the romanized form has to be measured too. We re-encoded each name in its Revised Romanization form, written without a hyphen, and set it beside the U.S. control group.
| Tokenizer | Korean names (romanized, 47) | U.S. names (20) |
|---|---|---|
| Aya-Expanse-32B | 19.1% | 100.0% |
| Gemma-3-27B | 14.9% | 100.0% |
| GPT-5 / gpt-oss | 10.6% | 100.0% |
| Ministral-14B | 8.5% | 95.0% |
| DeepSeek-V3.2 | 6.4% | 95.0% |
| Llama-3.1 / Qwen3 / OLMo-3 / Phi-4 / GPT-4 | 4.3% | 90.0% |
Table 7. Our measurement. The U.S. control group is the Social Security Administration's top 20 names for 2024, not the paper's sample. Korean names sit lower in all twelve tokenizers.
No row of the table reverses that order. In the five tokenizers with the widest gap, nine in ten U.S. names enter whole against barely four in a hundred Korean names. And which spelling you pick shifts the result again. Write "은우" as Eunu, per the official romanization, and it averages 2.00 tokens across the twelve; write it as Eoonoo, keeping the doubled vowels common on passports, and it becomes 3.08. "시우" likewise goes from Siu at 1.92 tokens to Sioo at 3.00. The spelling that looks easier to read is the one that adds tokens.
That doubled-vowel rule is a transformation we defined, not a survey of actual passports. For surnames, though, real usage statistics exist. Lay the National Institute of Korean Language's 2007 survey of passport romanization over the token counts and the direction does not point one way. For "최", the official Choe is atomic in only one of twelve tokenizers, while Choi — used by 93.1% of people — is atomic in nine. For "정", Jeong is atomic in three and Jung (48.6%) in all twelve. For "윤" it is the reverse: the official Yun is atomic in twelve tokenizers, Yoon (48.9%) in only four. Convention favors some surnames, the official standard favors others. Neither is a criterion the name's owner chose.
5.4Hangul Syllables Break at Byte Boundaries
On the Hangul side, something worse than splitting happens. OLMo-3, Phi-4 and GPT-4 each broke a character into raw bytes in 40 of the 56 names we submitted (71.4%). Decode those tokens individually and what comes back is not a character but a mangled fragment. These three also take exactly the same four surnames (Lee, Choi, Jeong, Cho) as atomic. The paper's observation that these three tokenizers share one access pattern reproduces in Korean as well.
The phenomenon has a name in the literature: byte fallback. When a tokenizer meets a character absent from its vocabulary, it passes the character through as UTF-8 bytes so nothing goes unrepresented. The Hangul syllable block holds 11,172 code points and one precomposed syllable is 3 bytes in UTF-8, so a syllable missing from the vocabulary bursts immediately into three byte tokens. One study measuring against the Llama-2 tokenizer reports byte-fallback rates of 23.4% for Japanese, 48.8% for Chinese and 52% for Korean. Korean is the highest of the three.
5.5Korean-Specialized Means Different Things in Different Models
None of the tokenizers measured so far was built for Korean. We put the same name corpus through three more: LG's EXAONE-3.5-7.8B, Kakao's Kanana-1.5-8B, and Upstage's Solar-Pro-preview. All three vocabulary files were downloadable without authentication.
Only two of the three appear in the table below; the third fell foul of the atomicity criterion and dropped out of the headline, for reasons given in the caption. There are three columns: the share of the 47 Hangul-written given names that enter as one token, the same share re-measured after romanizing them, and the count of names where a character broke at a byte boundary. Only that last column uses the 56-name set that includes the nine surnames. The Aya-Expanse row at the bottom is a yardstick — it had the highest romanized-name atomicity of the twelve tokenizers above, so it stands in for the ceiling among global models.
| Tokenizer | Hangul given names atomic | Romanized names atomic | Byte-boundary breaks |
|---|---|---|---|
| EXAONE-3.5-7.8B | 63.8% | 6.4% | 0 |
| Kanana-1.5-8B | 0.0% | 4.3% | 0 |
| (yardstick) Aya-Expanse-32B | 2.1% | 19.1% | 0 |
Table 8. Our measurement. Upstage's Solar-Pro is left out of the headline: being SentencePiece-based, it emits the word-leading space as its own token, so under this definition no name can be atomic. A structural difference, not a failure.
The two models chose different safeguards. EXAONE puts common given names into the vocabulary whole and takes 63.8% of them as a single token. Kanana does not take names atomically at all, but it held the syllable boundary in all 56 names we submitted: every two-syllable name comes out as exactly two tokens, with no character broken in the middle. "Korean-specialized" means vocabulary coverage in one case and breakage-free encoding in the other.
And EXAONE's advantage exists only in Hangul. Romanized, the same names drop to 6.4%, below Aya's 19.1%. Romanized is exactly the form a Korean name takes in international hiring systems and overseas applications. A design that serves Korean-language documents well carries no guarantee of serving a Korean person's name in an English-language file.
The section closes on what vendors do and do not report. Every published metric is a compression-efficiency metric. EXAONE 3.5 states that its 102,400-entry vocabulary is split roughly half Korean and half English and was trained after morphological pre-segmentation, and the follow-up edition reports growing the vocabulary to 150,000 and improving tokens-per-byte by about 30%. Kanana-2 advertises over 30% better Korean tokenization efficiency and a 1.6× advantage over the Qwen family. The Kanana 1.0 technical report has no tokenizer section at all. Not one of them reports whether names are in the vocabulary. And when you actually measure, the two properties do not move together. Compressing well and holding a name whole are different things.
One More Question for Every Bias Audit
The combination was familiar on first reading. Fellowship, hiring, clinical priority, lending. It looked like an overlap with the domains regulators have designated high-risk. So we opened the statutes and matched them one by one.
6.1The Four Tasks Land in Four EU Annex Slots
Annex III of the EU AI Act enumerates high-risk systems point by point. Each of the four tasks fits one point.
| Task in the paper | Annex III | Key statutory wording |
|---|---|---|
| Fellowship | 3(a) | Access and admission decisions for education and vocational training, "at all levels" |
| Internship screen | 4(a) | "to analyse and filter job applications" |
| Lending | 5(b) | "evaluate the creditworthiness of natural persons" |
| Clinical priority | 5(d) | "emergency healthcare patient triage systems" |
Table 9. Statutory wording is quoted from the source text. Nowhere does the paper say the authors picked their tasks by looking at regulation. What is verified is that the two sets overlap in the end.
6.2Korea Covers Three and Leaves One Out
Korea has a matching statute. The Framework Act on the Development of Artificial Intelligence and the Establishment of a Foundation of Trust (Act No. 20676, promulgated 21 January 2025, in force 22 January 2026) defines high-impact AI across ten domains in Article 2, item 4. Sub-item (g) covers "judgments or evaluations that materially affect an individual's rights and obligations, such as hiring and loan screening," folding two of the paper's tasks into a single line. Sub-item (c) points to healthcare provision and delivery systems.
The remaining one is the problem. Sub-item (j), which covers education, scopes itself to "student assessment in early childhood, primary and secondary education" and so excludes higher education — the exact opposite of the EU's "at all levels." A graduate fellowship review therefore falls into none of the ten Korean domains. And fellowship is the task where the paper's gap was widest: positive in seven of eight models, reaching 0.198 in the female-associated stratum.
On the duties side, Article 31 sets transparency notice, Article 34 sets operator obligations covering risk management, explanation, human oversight and documentation, and Article 35 sets fundamental-rights impact assessment. The practical gate for high-impact status carries a caveat, though: a system drops out of scope when a human is involved in the final decision. An amended Act also takes effect on 21 July 2026, so the discussion here is based on the statute as it currently stands.
6.3No Statute Reaches the Input-Representation Layer
So the statutes do identify the high-risk domains. The next question is this article's. How do they define "same conditions"? Is there any wording that requires the inputs a model actually receives to be comparable to one another?
| Document | Where it looks for bias | Input-representation layer |
|---|---|---|
| EU AI Act, Article 10 | Dataset representativeness, bias examination, mitigation | No such wording |
| Korean AI Framework Act, Arts. 34 & 35 | Operator obligations and fundamental-rights impact assessment | No such wording |
| NYC Local Law 144 | Selection rates and impact ratios from actual use | Not applicable — it does not use names |
| NIST SP 1270 | A taxonomy of bias | Gets as far as warning about "translation into simpler mathematical representations" |
Table 10. We found no document that requires, in statutory language, what appendix E of the paper asks for: "controlled at evaluation time." The Korean statutory wording came from a secondary edition, and a check against the official text is still outstanding.
Every notion of bias in the EU's Article 10 points at training, validation and test datasets. A dataset can be representative and can pass a bias examination, and the part where that data diverges as it passes through a model's vocabulary still sits outside the article. NIST's SP 1270 comes closest. It writes that error arises in translating complex data into simpler mathematical representations — but it stops at warning that information may be lost, and never crosses into a requirement to control for it at evaluation time.
NYC Local Law 144 is a different kind of instrument. It does not use names. Selection rates are counted from applicants' self-reported demographics under EEO categories together with actual usage history, and the published summary must even state how many applicants had unknown race or gender. Because names are never used as a proxy, the tokenization problem cannot arise in principle. The design differs — nothing was left out.
Stated precisely: a statutory audit counts outcomes from real applicant data, while a name-substitution test probes the model in advance. They are different tools. The trouble starts when the latter is offered as evidence that "our model is fair" — and the latter is exactly what gets used in vendor selection and pre-release review. Industry evaluations deserve the same sorting. OpenAI's First-Person Fairness, the best-known test that uses names for matching, swaps in 350 of them, and a read of the full paper turned up no statement that tokenization was controlled. Anthropic's discrim-eval names demographic groups in words and so never uses names at all. Lumping the two together as "industry ignores this" gets it wrong. The problem begins the moment a name becomes the matching instrument.
6.4What You Can Do on Monday
Appendix E of the paper states the practical principle in two sentences. Vocabulary comparability should be verified when a tokenizer is designed, tracked through model training, and controlled at evaluation time. And evaluators should open each name they plan to test in the relevant tokenizer, treating vocabulary support for a name surface as a separate evaluation variable. Results can then be balanced on the degree of support, stratified, or accompanied by a sensitivity analysis. Translated into our own practice, the principle comes to five items.
One. Encode every name used in the test with that model's tokenizer and count the tokens. With the vocabulary file in hand this takes seconds.
Two. Use the counts to balance the name set, to stratify by token count and report separately, or to attach a sensitivity analysis. State in the report which of the three you did.
Three. If you are on the receiving end of an audit report, ask one question back: "How many tokens were the names in your test, in this model?"
Four. Count again whenever the model changes. The same name is atomic in one model and split in another, and section 4.3 showed that post-training alone moves the signal.
Five. If Korean names are in the set, count the romanized spellings too. Choosing a spelling becomes choosing a token count.
The first item, in code: download the vocabulary file, encode each name with a leading space, and print the token count.
pip install tokenizers
curl -sL -o tok.json https://huggingface.co/Qwen/Qwen3-4B/resolve/main/tokenizer.json
python3 -c "
from tokenizers import Tokenizer
tk = Tokenizer.from_file('tok.json')
for name in ['Emily','Emilee','Jihun','Minjun','지훈','민준']:
ids = tk.encode(' '+name, add_special_tokens=False).ids
print(f'{name:8s} {len(ids)} {[tk.decode([i]) for i in ids]}')
"
The criterion matches the paper's section 3: a surface form with a leading space that is a single token and decodes losslessly back to that surface form is atomic. Checking only whether the string appears in the vocabulary file yields an upper bound, so it is not used here.
Why This Matters to Pebblous
The place this research points to has the same shape as a place Pebblous keeps arriving at in data quality work. The three passages below are not a product pitch. They look at how the structure shown in the preceding sections resurfaces in practice.
7.1Model-Ready Looks Different in Each Model
AI-Ready Data has meant putting data into a form a model can use. This research says that form differs from model to model. Two models given the same name column of the same CSV do not receive the same input. There is one more cell between the data and the model — the vocabulary — and whoever prepared the data does not decide what goes in it.
If the path DataClinic traced was a line running from a defect in a dataset to a defect in model output, this research draws that same cell onto the line. That cell does not change no matter how thoroughly the data is cleaned. It changes when the model changes.
7.2The Item Missing From Quality Checklists
Data quality checks usually look at missingness, duplication and distributional skew. What this research adds is representability. Values in a column can be equally clean and the column still be uneven, if some values sit in the model's vocabulary whole while others must be assembled from fragments. That skew is invisible to existing metrics, because it is neither missingness nor an outlier.
Korean data carries one more layer. As section 5 showed, almost every Hangul name enters in fragments, and in some tokenizers characters break at byte boundaries. Romanized, nothing breaks, but the gap against U.S. names stays where it was. This is not a column that is somewhat uneven; it is a column pushed entirely to one side.
7.3One Sentence to Ask Back
Organizations weighing AI adoption in hiring, credit and healthcare usually receive a bias audit report. What this article offers the recipient of such a report is the one sentence from section 6.3. How many tokens were the names in your test, in this model? No answer means that audit is carrying an uncontrolled variable. Korean customers are left with a further question: with Korean names, can that control be built at all?
Pebblous is not a company that changes models; it works on the stage before anything reaches a model. What this research says is that the stage is not outside the fairness conversation but at the front of it. Why measuring the pre-input stage cannot wait is said first by the paper's tables, not by us.
What we could not confirm is also worth recording. The paper was released on 28 September 2026, and as of writing no follow-up validation or rebuttal has appeared. Whether major model cards describe tokenization controls is something we did not verify by opening the cards, so we do not assert that none do. The wording of the Korean AI Framework Act came from a secondary edition, and a check against the National Law Information Center text is outstanding. On the methodological limits of the intervention experiment we did not follow the lineage literature, and rely on the paper's own caution instead. In the prior work on cross-language tokenization cost we could not find the figure for the Korean row, so no claim about Korean being several times more expensive appears anywhere in this article. The 20-name U.S. control group in our own measurement is likewise a sample we drew, not the paper's. Thank you for reading this far.
References
The evidence behind this article runs along four lines. Item 1 is the primary source the article is built on; we opened the full text and the appendix tables and transcribed the figures directly. Items 2 through 11 cover the lineage of name-based bias research and the interpretability literature supporting the reversal in section 4. Items 12 through 14 are the Korean tokenization background for section 5, and from item 15 on come the statutes cited in section 6 and the sources of the name corpus used in our own measurements. Every value in the section 5 tables is regenerated by this run's scripts, and none of it was computed by hand.
Primary source
- 1.Nayeem, M. T., & Rafiei, D. (2026). Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models. arXiv:2609.34065v1 (2026-09-28), University of Alberta. CC BY 4.0. Figures quoted directly from Tables 1–2 and appendix Tables 11, 12, 15 and 21, plus appendices E and F. Project site:
tafseer-nayeem.github.io/NameTrace. arxiv.org
Academic — name bias and tokenization
- 2.An, H., & Rudinger, R. (2023). Nichelle and Nancy: The Influence of Demographic Attributes and Tokenization Length on First Name Biases. ACL 2023. arXiv:2305.16577. The closest prior work; treats tokenization length as a confounder.
- 3.Bertrand, M., & Mullainathan, S. (2004). Are Emily and Greg More Employable Than Lakisha and Jamal? American Economic Review, 94(4), 991–1013. The 4,870 résumés, 9.65% and 6.45% in section 1 come from this paper.
- 4.Eloundou, T., et al. (2025). First-Person Fairness in Chatbots. ICLR 2025. arXiv:2410.19803. The best-known test that swaps in 350 names; we found no statement in the full text that tokenization was controlled.
- 5.Tamkin, A., et al. (2023). Evaluating and Mitigating Discrimination in Language Model Decisions. arXiv:2312.03689. The dataset is discrim-eval; it names demographic groups in words and uses no personal names.
- 6.Ahia, O., et al. (2023). Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models. EMNLP 2023, 9904–9923. arXiv:2305.13707. Up to a fivefold gap across languages; we could not verify the value in the Korean row.
- 7.Shwartz, V., Rudinger, R., & Tafjord, O. (2020). "You are grounded!": Latent Name Artifacts in Pre-trained Language Models. EMNLP 2020. Evidence for the line of work finding low-frequency names less stably represented.
Academic — interpretability (basis for section 4)
- 8.Elazar, Y., et al. (2021). Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals. TACL. Formalized the procedure separating decoding accuracy from causal effect.
- 9.Belinkov, Y. (2022). Probing Classifiers: Promises, Shortcomings, and Advances. Computational Linguistics. arXiv:2102.12452.
- 10.McDougall, C., et al. (2023). Copy Suppression: Comprehensively Understanding an Attention Head. NeurIPS 2023. arXiv:2310.04625. The mechanical account of sign reversal in section 4.2.
- 11.From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks. arXiv:2604.00778. Source of the figure that a single final-layer MLP suppresses roughly 7.5% of the correct answer's probability mass.
Academic — Korean tokenization (basis for section 5)
- 12.Park, K., et al. (2020). An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks. AACL 2020. arXiv:2010.02534.
- 13.Work on reinforcing CJK byte fallback. arXiv:2506.07541. Source of the byte-fallback rates — Japanese 23.4%, Chinese 48.8%, Korean 52% — measured against the Llama-2 tokenizer.
- 14.LG AI Research (2024). EXAONE 3.5 Technical Report. arXiv:2412.04862. Vocabulary of 102,400, roughly 50:50 Korean and English, with morphological pre-segmentation.
Policy and statistics
- 15.EU AI Act — Annex III 3(a), 4(a), 5(b), 5(d); Article 10(2)(f)(g) and 10(3). artificialintelligenceact.eu
- 16.Framework Act on the Development of Artificial Intelligence and the Establishment of a Foundation of Trust, Act No. 20676 (promulgated 2025-01-21, in force 2026-01-22), Article 2(4) and Articles 31, 34 and 35. The statutory wording came from a secondary edition, and a check against the official text is outstanding.
- 17.New York City Local Law 144 and the Department of Consumer and Worker Protection's final rule on automated employment decision tools. Selection rates and impact ratios based on self-reported demographics.
- 18.NIST SP 1270 (2022). Towards a Standard for Identifying and Managing Bias in Artificial Intelligence. DOI 10.6028/NIST.SP.1270.
- 19.Sources of the name corpus in our own measurement — Statistics Korea, 2015 Population and Housing Census, surname and clan volume; Supreme Court of Korea, newborn name statistics from the electronic family register system; U.S. Social Security Administration 2024 release; National Institute of Korean Language, 2007 survey of passport romanization.
Related Pebblous articles
- 20.Pebblous, "Subsidiaries whose names don't overlap never reached the candidate set." Another case where a name surface decides the outcome at the candidate generation stage. report/corporate-family-resolution-blocking-gap
- 21.Pebblous, "Don't Name the Country, and Whose Manners Does AI Answer With?" report/llm-cross-cultural-etiquette-gap
- 22.Pebblous, "VQ-Tokenizer: The Technical Backbone of Data Quality and Synthetic Data." Same word, different object — it covers vector quantization for images and audio, not the text segmentation discussed here. report/vq-tokenizer-data-quality-synthesis