Executive Summary
Multilingual AI safety benchmarks report coverage at the level of the collection. "Evaluated in N languages" is usually a true sentence. The trouble is that the way it is true and the way a reader takes it are two different things. Three authors, working independently under Black in AI Safety & Ethics (BASE), built an audit instrument that puts one language and one dataset in a single row, then opened all 25 language slices one at a time. Languages that sat side by side in a collection's summary turned out, once the slice was open, to be resources that differed in size, in provenance, in how much verification they had received, and in whether anyone could get them at all.
The firmest evidence came from inside a single pipeline. In a 2026 resource that ran Hausa and Swahili through the same English seed, the same model, the same machine-translation step, the same native-speaker verification protocol and the same acceptance threshold, only the Hausa policy documents fell below the bar the authors themselves had set. The verification effort was nominally identical: one verifier per language, reviewing a sample of twenty pairs. Equal treatment produced unequal reliability. Open the harm categories language by language and some cells stop being a matter of degree. Neither Hausa nor Swahili has native-authored data for self-harm or sexual content, and for Hausa self-harm there is no translation, no transcreation and no synthetic data either.
The most portable sentence the authors leave behind is about overlap. Narrow coverage, thin verification, restricted access and marginal quality are each survivable on their own. The risk lives in the four of them landing on the same slice, and no field on a data card is designed to surface that. Korean is not where the two African languages are, but it sits on the same axis. It appears in five of eight major multilingual safety resources; of those five, exactly one was newly authored in the language, and exactly one prints a slice-level agreement figure. A column name is not a value, and in safety data that holds in precisely the same way.
2 / 6
Harm categories with native Hausa data
French 4, Swahili 4 (self-harm: French only, of the three)
66.37 / 93.30
Policy-corpus translation quality that split inside one pipeline
Hausa / Swahili; the original authors' acceptance bar was 70
6 / 25
Slices reporting inter-annotator agreement at slice level
The lowest figure came from the most conscientiously documented resource
300 / 986
Hausa items actually released against the count the paper describes
TukaBench: a component described in the paper was never uploaded
What "Evaluated in N Languages" Actually Counts
Model cards and technical reports state language coverage in a single line. The audit's introduction reproduces that structure exactly: "A model card or technical report states that the system was 'evaluated in N languages, including low-resource language X,' and a reader naturally takes that evaluation to carry something like the rigour of the model's English safety evaluation." The authors argue that the inference often has nothing under it. They are also careful to explain why the statement is not a lie. The sentence is true. It is just true at a different unit than the one the reader is reading at.
Provider documentation supplies a real example without much searching. OpenAI's GPT-4o system card, published in 2024, reports that red-team participants "spoke a total of 45 languages and represented geographic backgrounds in 29 countries." There is a 45, and there is no per-language breakdown. A later section of the same document, on underrepresented languages, describes evaluations built with external researchers for Amharic, Hausa, Northern Sotho, Swahili and Yoruba. Those evaluations are three: a translated ARC-Easy for elementary science questions, a translated TruthfulQA for misleading questions, and Uhura-Eval, a newly built reading-comprehension test. They measure capability and truthfulness, two of the three are translations, and none of them measures a harm category such as hate speech or self-harm. The same section notes that considerable work remains to raise the quality and coverage of these evaluations. This comparison is possible only because OpenAI is among the few that put African-language evaluation results into a document at all. Documentation that says nothing offers nothing to check against.
What the audit team did was move the unit of counting. A multilingual dataset is not one artefact of uniform quality. In the authors' words, it is "a collection of per-language subsets that may differ in how they were constructed, how many items they contain, who reviewed them, and whether agreement was ever measured." So they set the unit of audit at the language slice. One row of the instrument records one language inside one dataset. A dataset covering both Hausa and Swahili becomes two rows, and the two rows are assessed independently and may legitimately reach different conclusions.
Adopting that unit forces you to write down counting rules, because a row count is not a dataset count. The team states three. A dataset is counted once. A study that re-translates an existing prompt set and re-evaluates it counts as evidence about that dataset, not as a new one. A resource that merges existing corpora and re-labels them is counted once at the merged layer, with the lineage recorded so the originals are not counted twice. These rules change the numbers in practice. Applying the third alone: one Swahili resource presents 101,014 labelled items, but it is built from three corpora already inside the audit, so counting those items as new coverage would inflate the apparent Swahili corpus roughly fourfold.
| Figure | What it counts |
|---|---|
| 25 | Rows in the audit instrument, i.e. language slices. Hausa 7, Swahili 10, French 8. Not a dataset count |
| 21 | Distinctly named resources. Writing "21 datasets" gets it wrong |
| 20 | Datasets after the counting rules are applied (one of the 21 is a derivative replication). Of those, 19 are in scope and 1 is a boundary case |
| 26th slice | Ubisoft ToxBuster (game chat). Reviewed and excluded because the data was never released. It appears in no count |
| 101,014 | Labelled items reported by one Swahili resource. It merges three corpora already in the audit, so counting them as new coverage inflates the Swahili corpus roughly fourfold |
Source: Onuoha, Sunu and Sikiru (2026), arXiv:2608.13695, §4.6 and §5.1
The verification procedure was not a formality either. For every field that could conflict with the primary sources, the team checked three places in order: the paper's own methods section and per-language tables, the linked repository or data card, and, where the repository was live, the actual contents of the release. Where the three disagreed, claims about how something was built deferred to the paper's per-language tables, and claims about whether it can be used today deferred to the release. That procedure changed the record in four ways.
- Language attribution corrected. A dataset first recorded as covering Hausa, on the strength of a plausible secondary description, turned out not to cover that language in the repository's own language list. In the other direction, a language added in a later release had initially been recorded as absent. Neither was catchable from an abstract; both required reading the repository's version history.
- Provenance corrected. Two datasets whose secondary descriptions implied native authoring stated in their own methods sections that the content was written in English and machine-translated. One post-edited the output; the other reports no post-editing.
- Release status corrected. One dataset's actual release contained far fewer items in the target language than the paper describes, because a component named in the paper was never uploaded. Another was announced and never shipped.
- Authorship corrected. A widely cited "expert-written" claim in fact applies only to the English seed material, which the paper itself states was not its own contribution.
The authors record these four not as an erratum but as methodology: a sign that the audit procedure is doing work that reading abstracts cannot do. The same posture carries into how they published. The team released the whole instrument, all 31 fields applied identically across the 25 slices, together with the evidence note written for each judgement. In their words, it is "the same standard of disclosure we ask of dataset authors." Open the file and there is a grain of judgement in it that the final tables do not show: of 26 candidate rows, eight carry a hold flag and two an exclusion flag. That this comparison is possible at all is a consequence of publishing the audit tool.
The first observation from that count runs the unexpected way. Swahili, a mid-resource language, is represented by more slices than French, a high-resource one: 10 to 8. The authors immediately rule out the reading that Swahili is therefore better served. Its slice count is inflated by a chain of re-annotations over the same tweet collection and by studies re-evaluating the same prompt set, while the French slices are largely independent resources. Slice count is a poor proxy for coverage.
The sentence that carries this section sits in the paper's introduction. "A dataset reporting '14 languages, 20,000 items' may resolve, for a given low-resource language, into a few hundred machine-translated prompts checked by a single verifier, with inter-annotator agreement never measured, distributed under a licence or access restriction that a downstream developer is unlikely to notice before citing the dataset as evidence of coverage." A collection-level statement does not show you that resolution happening.
The Same Pipeline Produced Different Reliability
The standard objection to any claim about per-language gaps is confounding. If one language's data looks worse, that may be nothing to do with the language: a different team built it, or a different method, or a different budget. The audit's firmest evidence comes from a place where that objection does not apply. UbuntuGuard, released in 2026, ran several African languages through one pipeline. Same English seed, same generation model, same machine-translation step, same native-speaker verification protocol, same acceptance threshold. The amount of verification effort was nominally identical too: one native verifier per language reviewing a sample of twenty pairs, applied the same way across eleven languages.
Inside that pipeline, Hausa and Swahili came apart. The place they came apart was translation quality on the policy corpus. Hausa policy documents scored 66.37, below the acceptance bar of 70 the authors had set for themselves; Swahili scored 93.30, comfortably above the same bar. Yet Hausa conversation transcripts scored 93.31 and cleared it. The shortfall, in other words, is not a general failure of Hausa translation.
| Slice | Policy corpus, translation quality | Conversation transcripts | Final slice size |
|---|---|---|---|
| Hausa (low-resource) | 66.37 (below the threshold of 70) | 93.31 | 1,656 train / 278 test |
| Swahili (mid-resource) | 93.30 | 96.99 | 1,899 train / 435 test |
Source: GEMBA-SQM translation-quality scores from UbuntuGuard (arXiv:2601.12696) as reported in §6.2 of the audit, with slice sizes from its Table 3. The threshold of 70 was set by UbuntuGuard's own authors
Four numbers, and four things about the columns that have to travel with them. First, these are not scores assigned by an outside evaluator; they are translation-quality figures the dataset's own authors published about their own material. As the audit puts it, "we report the numbers that paper publishes about itself." Second, the threshold of 70 is also theirs, which means the verdict is not an external yardstick brought in after the fact. Third, what fell short was one document type, the policy corpus, while the structurally simpler conversation transcripts cleared the bar. Fourth, the slice that fell short was released anyway, and the collection-level summary presents ten African languages as similarly covered.
One misreading has to be closed off here. 66.37 is neither the only sub-threshold score nor the lowest. The same UbuntuGuard table ran eleven African languages through the same pipeline against the same threshold, and three languages score lower than Hausa on the policy corpus: Igbo at 42.62, Nyanja at 48.61 and Luganda at 62.08. The audit does not claim Hausa is worst either. The two languages were singled out because fixing the pipeline is what removes the non-linguistic confounders. This correction does not shrink the paper's claim; it enlarges it. While a collection-level summary spoke of even coverage across ten languages, the sub-threshold slices inside it were not one but several.
The sentence the authors draw from this case is the section's conclusion. "Equal treatment produced unequal reliability, and this is precisely the mechanism by which a single resource can be procedurally fair and substantively inequitable at the same time." The rule of one verifier per language was applied evenhandedly to every language. An evenhanded rule produced different outcomes by language, and that difference is invisible from inside the rule.
Open It Language by Language and Two Cells Come Up Empty
The gap in the previous section was a difference in scores. This one is a difference between something and nothing. The audit works with six harm categories: hate speech, harassment, self-harm, extremism, sexual content and misinformation. The taxonomy is borrowed from trust-and-safety practice, and it was chosen because most of the corpus maps onto it. Lay the three languages across those six cells and you see at once which cells are filled with what, and which are empty.
| Harm category | Hausa (low-resource) | Swahili (mid-resource) | French (high-resource) |
|---|---|---|---|
| Hate speech | Native, multiple datasets | Native, multiple datasets | Native + translated |
| Harassment | Native | Native | Native + translated |
| Self-harm | No coverage of any form | Transcreation only (native 0) | Native (Aya) + translated |
| Extremism | None | Native (incitement portion) | Translated only |
| Sexual content | Translated only (native 0) | Transcreation only (native 0) | Native (Aya) + translated |
| Misinformation | Synthetic or translated only | Native (PolitiKweli) | Translated only (native 0) |
Source: Table 6 of the audit. Transcreation adapts source material to the target culture as it is carried over, which distinguishes it from material a native speaker wrote in that language from the start
Count the categories with native data and you get French 4, Swahili 4, Hausa 2. High and mid are level; only low-resource trails. There is no clean ranking that descends with resource level. The cell most likely to be misquoted from this table is self-harm, so here is the precise statement. Neither language has self-harm data authored natively. For Hausa there is nothing of any form. The sentence in §6.6 reads: "Hausa self-harm has no coverage of any form. No native, translated, transcreated or synthetic resource in our corpus carries a self-harm label for Hausa." Swahili does have material that arrived by transcreation, so writing that neither language has any self-harm data at all overstates the Swahili case. Sexual content is the same for both in that neither has native material; only the route differs, translation for one and transcreation for the other.
Half the reason the ranking wobbles is the yardstick. The labels low-, mid- and high-resource move depending on what you measure. Among the resources inside the audit corpus, those that operationalise resource level as crawl share classify Swahili as low-resource, because they use a cutoff of 0.1% of the web crawl, and on that cutoff Swahili sits in the same box as Hausa. Other resources, and the authors' own prior classification, treat Swahili as mid-resource: it is one of the most widely spoken languages in Africa and has a comparatively large body of curated NLP material. The authors do not resolve the discrepancy; they expose it as a disagreement within the literature. They add that Swahili's results landing between Hausa's and French's on the dimensions the audit actually measures is independent support for the mid-resource reading, while declining to leave the tension as an unexamined assumption. That tier labels move with the yardstick stays a premise through the next two sections.
The authors' argument turns on which two cells are the empty ones. Self-harm and sexual content are categories where the signal is carried by idiom, euphemism and indirection. Fill them by translation or transcreation alone and what drops out is the way the harm is actually expressed in that language. One Hausa entry in the audit instrument records concretely what drops out: the list of mistranslations left in the "quality concerns" field for the HOC dataset.
| Hausa expression | Meaning | What the translation produced |
|---|---|---|
| dan iska | Thug, hoodlum | "and iska" (treated as a proper noun) |
| yan kutumar uba | Obscene insult | "father's cousins" (word-for-word) |
| wawa jaki | Fool, donkey | "wow what" (mistaken on phonetic similarity) |
| dan jaka | Son of a donkey | "and jackets" |
| sakarai / sakara | Dull-witted (masculine / feminine) | "boy" / "connection" |
Source: the "quality concerns" field of the HOC Hausa row in the published audit instrument (safety-slice-audit). Only the sense of each expression is given here
All six expressions in the table were translated literally or compositionally, and all six lost the idiomatic insult. That is primary evidence for why a harm label reached by translation alone cannot stand in for a native one. Why we can see this list at all matters as much. The HOC resource documents, about itself, that posts under eight words were removed and the text was lowercased and stemmed, so it is not verbatim source material, and that its user-study sample skewed 74.7% male and highly educated, which may have biased the verification judgements themselves. This is visible not because the resource is careless but because it wrote things down.
Evidence that scarcity alone cannot explain the table sits inside the table. On the misinformation row, native coverage exists for mid-resource Swahili and does not exist for high-resource French. The Swahili resource is PolitiKweli, a code-switching corpus built for election response. The authors' sentence in §6.9 explains the reversal: "Coverage and annotation practice follow the priorities and capacity of the community that studies the language, not the language's share of the web crawl." Resource level sets a ceiling on how much well-verified data a community can independently produce, but what gets built under that ceiling is not chosen by resource level.
The taxonomy travels along with the data. The four dominant classification frames in the corpus are a ten-way scheme of jailbreak behaviours, a twelve-category risk hierarchy, a six-category macro safety scheme, and the operational definitions of an automated toxicity API. All four were built in English-speaking contexts, and translated resources inherit them intact. Cases where a local conception of harm was established locally are a minority. Three of the four exceptions are African-language resources. As the authors put it, "the problem of taxonomy transfer is being addressed within these communities, and not by the collections that make the largest claims about multilingual coverage."
Every Claim True, the Overlap False
The most operationally useful conclusion in the audit is not about any particular dataset but about how you check. In §7.2 the authors write that a paper stating "this dataset covers Hausa hate speech" is not lying. The real risk arises when narrow coverage, thin verification, restricted access and marginal quality appear together on the same slice, and "no field on a data card is designed to surface the conjunction." That makes the following sentence a methodological conclusion. Checking claims one at a time "systematically underestimates risk in exactly the cases where risk is highest, because the failure mode is not that any single claim is false but that the correlation structure among the claims is invisible from where the dataset consumer stands."
The audit's table of gap categories by tier puts that correlation structure on one screen. Severity runs on four steps: none, mild, moderate, severe.
The table only yields its meaning if you read a column top to bottom. Read it that way and the severe marks land in different places for each language. Hausa is severe on three rows, provenance, access and licensing, and harm-category coverage; Swahili on two, annotation and agreement, and reuse. French has no severe marks at all, but its quantitative quality stays unmeasurable. Of the six categories, only two, provenance and access/licensing, degrade monotonically with resource level; in the other four the ordering wobbles or inverts. What the table shows is not a ranking but where things overlap.
| Gap category | Hausa | Swahili | French |
|---|---|---|---|
| Provenance | Severe: multiple harm resources entirely translated or synthetic | Moderate: native material exists but single-domain | Mild: native red-team material exists; breadth comes from translation |
| Annotation and agreement | Mixed: strong where reported (κ 0.66–0.75), absent for half | Severe: lowest agreement in the corpus (κ 0.13–0.55) | Moderate: printed for very few slices (α≈0.47) |
| Access and licensing | Severe: unreleased, request-only, no licence, password-gated | Moderate: mostly public, one with no repository | Mild: mostly public, one industry resource unreleased |
| Harm-category coverage | Severe: native material confined to hate and harassment | Moderate: adds misinformation and incitement natively | Mild: all six categories, mostly via inherited taxonomies |
| Reuse and double counting | Mild: limited overlap | Severe: a four-deep re-annotation chain | Mild: shared English seed |
| Quantitative quality | Documented shortfall: policy corpus at 66.37 | Documented advantage: 93.30 / 96.99 | Unmeasurable: French is absent from the only within-pipeline comparison |
Source: Table 5 of the audit. Severity runs none, mild, moderate, severe. Five of the six categories are ordinal judgements; only quantitative quality rests on measurement
The cell most likely to be quoted is the second row. Inter-annotator agreement is reported at slice level for six of the 25 slices. The rule for counting it is strict: it counts only where the original paper prints a numeric agreement figure for that language's own labels. Three things therefore do not count as agreement. One is a single collection-level value never broken down by language. One is a metric that measures something else, such as the rate at which model judgements match human ones. One is a single verifier's calibration standing in place of agreement. Count those three and the number of slices "reporting agreement" inflates.
The direction of those six values also defies expectation. The worst agreement came not from the low-resource language but from the mid-resource one; where Hausa reports agreement, it is on the strong side. The corpus minimum, κ=0.13, came from the Kenya slice of XTREMESPEECH, a resource that hired local fact-checkers, engaged the disagreement between annotators and communities, printed a full datasheet in its appendix, and reported and discussed the low figure honestly. The most conscientiously documented resource ended up holding the lowest number. In §6.4 the authors read the value this way: "We read low agreement in this situation as evidence about the taxonomy, not as evidence about the annotators." Consistent with that, the highest agreement came from a taxonomy the language community designed itself.
Nor is a missing agreement figure necessarily negligence. Prolonged exposure to abusive material is itself a harm, and some resources reduced duplicate annotation for that reason; the authors decline to code this as laziness. The structural consequence stands anyway. The languages where annotator-welfare constraints bind hardest are the languages with the smallest annotator pools, so the least-resourced slices stay the least verified. Locally ethical decisions produce a globally inequitable distribution.
Access and documentation repeat the same overlap. TukaBench's Hausa slice is described in the paper as 986 items; 300 were released, because a component named in the paper was never uploaded, and within this corpus that failure appears only for low-resource languages. In the authors' words: "The citation resolves, the repository exists, the data card renders, and the artefact is not what the paper describes." One resource met every quality criterion and was never released. One has a size that appears in none of the paper's three versions, a documentation finding the audit established by comparing all three rather than something it missed. Add the re-annotation chain from the previous section: if four differently named datasets re-annotate the same underlying text, counting four still leaves you one piece of independent evidence.
Take the items in this section one at a time and none is fatal. A small slice, no native review, no agreement figure, an unclear licence, a released volume that differs from the paper: each on its own is a condition you can accept and work with. The trouble is the five of them sitting on one slice, and that fact is written in no field. Overlap is visible only in the table, never in the fields.
Translation Attacks Closed, Multi-Turn Did Not
A data story carries operational weight when it connects to model behaviour. Here the audit writes carefully, and the caution has to be carried over with the finding or the meaning collapses. The starting point is a 2023 result. MultiJail reported an unsafe-response rate of roughly 83% for Swahili jailbreak scenarios against the models of that moment. Precisely: ChatGPT was at 83.49% in the condition where a jailbreak template was supplied alongside the prompt, and at 7.94% when a translated harmful prompt was entered on its own. Blend the two values and the meaning falls apart.
A 2026 replication reports that this attack surface has largely closed. Translation-based single-turn attacks are blocked. What remains is multi-turn conversational attack, and those figures are the ones most often misquoted. The audit's introduction writes "roughly 42% to 71% depending on model and language," but that range is the across-model range for Kiswahili alone. Precisely, it is 41.8% to 70.9%, and in the same table English runs 52.7% to 83.6% and Afrikaans 60.0% to 78.2%. Hausa was not part of that replication. §7.3 of the same paper states the relation exactly: multi-turn attacks remain "effective at rates comparable to English."
The conclusion these numbers support is therefore not "it breaks because the language is low-resource." Multi-turn is an attack surface that breaks regardless of language, and the gap between languages lies in whether native material exists to close it with. Omit English's multi-turn unsafe-response rate and this becomes a story about scarcity, which is the conclusion the audit explicitly rejects.
The data-side facts have the same shape as that gap. For Hausa and Swahili, the only multi-turn safety resource is a single synthetic and machine-translated pipeline, the UbuntuGuard seen in the previous section. Neither language has any natively authored multi-turn material anywhere in the corpus. The two other multi-turn entries in the corpus are a derivative replication and an out-of-scope boundary case. The authors do not claim this correspondence is causal. Their wording: "We state this as a structural correspondence, not a demonstrated causal chain." They add that there is no way to know what data any commercial provider used, and that the direction could run the other way, with providers not investing in that data because the attack surface had not yet been publicly demonstrated.
Nor is the gap the absence of large resources. PolyGuardMix spans 17 languages at 1.91 million items and includes French, but not Hausa or Swahili. PolygloToxicityPrompts holds an even 25,000 items per language and carries no human harm labels at all, only automated toxicity API scores. M-ALERT puts 15,000 items per language into five languages, all high-resource, its authors stating that they chose depth over breadth. Scale is not coverage, and being high-resource, native and large still guarantees no human verification.
What the audit establishes, then, is not causation but procurability. The authors' claim is this: if a provider set out today to close this particular gap, the raw material for doing it with any confidence, African-language safety data that is natively authored, multi-turn, diverse across harm categories and agreement-verified, "does not yet exist in the published literature." That is a claim about the existence of materials rather than about model vulnerability, and it is the kind of fact an audit alone can settle.
A terminology problem attaches here. "Safety" carries at least three senses in this literature: content safety; jailbreak safety, which asks whether a model can be made to emit what it was trained to refuse; and agentic safety, which asks whether a tool-using model can be steered into fraud or privacy leakage. The audit covers the first two and excludes the third, but rather than quietly dropping a resource it encountered, it kept it as a boundary case. The resource is the Kiswahili deep-dive testing in the third joint exercise of the International Network of AI Safety Institutes, which included Kenya and Korea. It was natively reviewed across 156 tasks per language and is methodologically careful, but the harm categories it covers are fraud and privacy leakage, which do not intersect the audit's six. The authors' sentence marks the implication: "A downstream reader who encounters this resource and concludes that 'Kiswahili safety has been evaluated' is missing that distinction."
One cell runs the other way. Four slices model code-switching explicitly: three Swahili-English and one Hausa-English. It is one of the rare dimensions where African-language resources are ahead of the French ones. The problem of getting real mixed-language speech into data was not solved first by the better-resourced side.
What the Korean Slice Looks Like
Korean is not among the audit's three languages. The method, though, is indifferent to which language you point it at, which makes any mid-to-high-resource language a usable test of whether the finding generalises. Ask the same questions of Korean, opening the major multilingual safety resources one at a time to see how each Korean slice was built and how much verification is actually printed, and answers come out. The table below is the result of doing that for eight resources. Cells that could not be confirmed are marked unconfirmed rather than guessed.
| Resource | Korean | How the Korean slice was built | Slice-level agreement |
|---|---|---|---|
| RTP-LX | Present | Not newly transcreated; an existing Korean hate-speech corpus was reused (the same exception applies to Hebrew, Danish and Brazilian Portuguese) | Unconfirmed |
| PolyGuardMix | Present (of 17 languages) | Naturally occurring conversations plus machine translation of English-only material, with sampled human verification. 1.91M items in total; no standalone Korean table | Yes: 50 items per language, 3 raters, translation quality 81.55 |
| MultiJail | Present (mid-resource tier) | Human translation by native speakers. 315 items (3,150 across 10 languages, the same structure as the paper's Swahili 315) | Unconfirmed (only an overall pass rate across 9 languages is reported) |
| PolygloToxicityPrompts | Present (of 17 languages) | Naturally occurring web text plus automated toxicity scores. 25,000 items per language | None (the design carries no human annotation at all) |
| CultureGuard v3 | Added in v3 | English material culturally adapted, machine-translated, then quality-filtered. Not newly authored in the language. 42,393 Korean items | None (only automated cross-lingual consistency filtering) |
| M-ALERT | Absent | Five languages only: English, French, German, Italian, Spanish. The authors state they chose depth over breadth | Not applicable |
| XSafety | Absent | Korean is not among its 10 languages | Not applicable |
| Aya red-teaming dataset | Absent | Korean is not among its 8 languages. The associated language models support Korean; the red-teaming safety dataset does not include it | Not applicable |
Compiled by reading each resource's paper, appendices and public repository directly (August 2026). The Korean item count for CultureGuard was obtained by counting the released files
Three things read straight off the table. First, Korean appears in five of eight and is absent from three. That is not the zero the African languages face, but the structure is the same: whether a language is included depends on how the resource was designed. A design that goes deep on five high-resource languages has no Korean; a design that spreads across seventeen does. Second, of the five that include it, exactly one, MultiJail, was newly produced by native speakers. The rest are a reused existing corpus, machine translation with sampled verification, naturally occurring web text, and synthesis plus machine translation. That is precisely the native-versus-translated axis seen with Hausa and Swahili. Third, slice-level agreement is confirmed as a printed number in exactly one case, and even there the sample is 50 items per language. The same order of magnitude as six of 25 on the African side.
Korean's resource position points the same way on both yardsticks. In the July 2026 crawl statistics, Korean pages are 0.8429% of the total, 17th among detected languages, and in the resource-tier classification widely used in the field it is Class 4. In the same crawl statistics French is 4.8032%, Swahili 0.0117% and Hausa 0.0026%. Korean is roughly 72 times Swahili; French is roughly 1,850 times Hausa. The precise answer: on these two indicators Korean sits in the upper-middle band, and below French.
Resource position is not the same thing as safety performance, though. The MultiJail table that produced the 83% quoted in the previous section has a Korean row. In the condition where a jailbreak template was supplied, Swahili was at 83.49% and Korean at 80.00%; with a translated harmful prompt entered on its own, Korean's 9.84% was higher than Swahili's 7.94%. English was 72.06% in the first condition and 0.63% in the second. These are results against 2023-era models and do not transfer to today's. Still, the fact that two languages separated by a factor of 72 in web share were similarly open to the same attack repeats, on the Korean side, the previous section's point that resource indicators do not predict safety indicators.
Korean also has an axis that no translated prompt set can hold. A result presented on 7 July 2026 at the second AI Safety Seoul Forum, by Kim Bo-ryeong, a doctoral researcher at Seoul National University, shows it. In the hate-speech category, asking in the plain speech level produced unsafe responses 10.1% of the time; asking in the polite speech level, 3.0%. A factor of 3.3. In the same presentation English came in at 23.2%, higher than Korean plain speech. The presenter's conclusion was that a model can be safe in one language and unsafe in another for the same content. Which models were tested does not appear in the published coverage. Where this meets the audit is clear enough. Korean grammaticalises social distance: the speech level a sentence is cast in is part of how aggression is carried, and a set produced by translating English prompts cannot hold that axis at all. It is the Korean edition of the argument that categories where idiom and euphemism carry the signal are not reachable by translation.
Evidence that the audit's recommendations are not an impossible ask also comes from the Korean side. XL-SafetyBench, released in May 2026, was built by AIM Intelligence with Microsoft, the Korea AI Safety Institute, KT and Seoul National University among others, and holds 5,500 test cases across ten country-language pairs. Two native annotators per country worked independently, and the appendix prints the agreement. Nine binary verification filters average between 92.7% and 98.1% agreement, and the ordinal criteria come to quadratic-weighted kappa of 0.50 for country sensitivity and 0.49 for country specificity. The authors also state that per-country and per-task agreement breakdowns are released with the benchmark. Two of the audit's recommendations, making per-language reporting the default and reporting slice-level agreement, are carried out here exactly as written.
The finding in those results that most resembles this report's subject concerns an illusion in the safety metric. Across ten frontier models, US-grounded prompts were the safest, and jailbreak vulnerability was highest for the United Arab Emirates and Korea. The authors read this as English-centric alignment disproportionately benefiting US-centric contexts. Among country-local models, though, attack success rate and nonsensical-response rate form a nearly linear trade-off, correlating at -0.81. The safety of a local model with a low attack success rate turned out to rest not on principled refusal but on failing to understand the question, which the authors call a safety illusion produced by comprehension failure. Looking safe on a single metric while actually measuring the wrong thing is the same structure this report has followed from the start. That said, this benchmark's ten country-language pairs include no African language, and its authors note in their limitations that the country selection skews towards Western Europe. The contrast, that Korean now has a resource like this and the two African languages do not, is not a matter of one being better than the other; it is another confirmation of the earlier point that who was able to build determines what gets built.
There is also a recent record of the same structure repeating across regions and domains. A Bengali study published in August 2026 examines educational AI infrastructure rather than safety benchmarks, but the shape of the deficit overlaps. Bengali holds under 0.5% of web content against roughly 4% of the world's speakers, and the ratio of English to Bengali training tokens is 67 to 1. Gaps in safety data and gaps in infrastructure are produced in the same place.
Why This Matters to Pebblous
In data quality diagnostics, the first job is not choosing a metric but fixing the unit. Measure completeness per table and the answer is "loaded"; measure it per column and per segment and specific ranges turn out to be empty. This audit moving its unit from the collection to the language slice is the same design decision, reached independently in the different domain of safety data. Our reason for decomposing completeness, accuracy and consistency to the slice level when we report on customer data gets third-party evidence here. The problem of what the catalogue says differing from what a measurement finds showed up in the same shape in our audit of Korea's public AI training data.
1Format and provenance set the model's ceiling
The structural correspondence in section 5 shows the route from a data deficit to a behavioural one, and it shows it per format. The attack surface that had single-turn data behind it closed; the one with no native multi-turn data stayed open. That the format and provenance of data set a ceiling on how robust a model can become is not an abstract claim here. That synthetic and translated pipelines add volume but do not substitute for native collection at the frontier of coverage is written directly into the audit team's recommendations, and the 66.37 from inside one pipeline demonstrates it against those authors' own bar. Our earlier piece on the share of low-resource languages in pretraining data covers the training-data side of the same problem.
2Six questions to put back to a data supplier
Translate the audit team's recommendations that a data buyer can use as-is into questions, and you get these six. They share one property: none of them can be answered from collection-level documentation.
| Question | What happens if you don't ask |
|---|---|
| How large is our language's slice | You read the total item count as the count for that language |
| Is that slice natively authored or machine-translated, and was there native review | You treat a label reached by translation as a native label |
| Is slice-level agreement reported, and if not, is the reason documented | You read a single collection-level value as your language's value |
| Does the item count described in the paper match the count actually released | You plan around data you never received |
| Do labels exist in that language for the harm categories we care about | You read the existence of a category list as the existence of labels |
| Are two differently named datasets re-annotations of the same source | You count one piece of independent evidence several times |
Recommendations R1, R3, R4 and R7 from §8 of the audit, restated as checks to run at procurement and acceptance
These six overlap substantially with what regulation asks for as evidence. Article 10 of the EU AI Act enumerates data preparation as six operations, annotation, labelling, cleaning, updating, enrichment and aggregation; requires training, validation and testing data to be relevant, sufficiently representative and, to the extent possible, free of errors and complete; and states that examination and mitigation of bias, and the identification of data gaps, must be documented. Application has been deferred to 2 December 2027. That the audit found agreement figures for six of 25 slices is a preview of which boxes stay empty when material of this kind has to be submitted as evidence. We covered the problem of turning labelling work into audit evidence in a separate report.
3On how to read this audit
Without the limitations the authors wrote down themselves, this piece reads as a takedown of African-language resources. Five of the six severity categories are ordinal judgements; one rests on measurement. Severity was assigned per slice by a single annotator, and the audit team's own reliability was not measured, which the authors note is the same tension they criticise in others. Corpus discovery was citation chasing and targeted search rather than a preregistered systematic review, and they state that the existence of undiscovered Hausa and Swahili safety resources is "not merely possible but likely." The resources audited in most detail are also not the bad ones. They are the ones whose gaps were visible because they documented their own limitations unusually transparently, and there is no evidence anywhere that they are worse than less transparent alternatives. The misuse the authors worry about most is stated explicitly: these findings being used to justify disinvestment. What the controlled comparison showed, they insist, is not that collecting Hausa safety data is intrinsically harder, but that this is a gap existing methods can close given sufficient native verification effort.
Editor's Note. Pebblous works on data quality diagnostics and lineage design, so we have an interest in this subject. This report is not written to recommend a product; it is an attempt to put on record that the fact that the unit you count coverage in decides whether a safety claim holds has been confirmed independently, outside our own domain. That completeness is measured at the slice and not at the collection is a claim we have made repeatedly in customer diagnostics, and an academic audit arriving at the same conclusion in safety benchmarks suggests the method is closer to a general principle of data quality measurement than a property of any particular product. Which unit to count in remains for the people who build the data and the people who use it to decide for themselves.
References
Primary sources
- 1.Onuoha, C. P.-M., Sunu, B. E., Sikiru, R. (2026). Language-Specific Gaps in AI Safety Training Datasets. arXiv:2608.13695 [cs.CY], 13 August 2026, CC BY 4.0. (Conducted independently by the authors, affiliated with Black in AI Safety & Ethics. Most figures and quotations in this report come from §4–§8, Tables 3, 5 and 6, the Limitations and the Ethics Statement)
- 2.Onuoha, C. (2026). safety-slice-audit (dataset). HuggingFace, CC BY 4.0. (The full audit instrument, 26 rows × 31 fields. The mistranslation record for the HOC Hausa row was verified in this file)
Sources cited by the audit
- 3.Abdullahi et al. (2026). UbuntuGuard. arXiv:2601.12696. (Translation-quality table for 11 African languages run through one pipeline. Hausa 66.37, Swahili 93.30, Igbo 42.62, Nyanja 48.61, Luganda 62.08, threshold 70)
- 4.Akinode et al. (2026). TukaBench. arXiv:2606.01322. (986 Hausa items described against 300 released)
- 5.Deng, Y. et al. (2023). Multilingual Jailbreak Challenges in Large Language Models. arXiv:2310.06474. (MultiJail. Per-language unsafe-response rates in Table 1: Swahili 83.49%, Korean 80.00%, English 72.06%; 315 items per language)
- 6.Marx, D., Dunaiski, M. (2026). 2026 replication. arXiv:2605.18239. (Multi-turn unsafe-response rates: Kiswahili 41.8–70.9%, English 52.7–83.6%, Afrikaans 60.0–78.2%)
- 7.PolyGuard (2025). PolyGuard: A Multilingual Safety Moderation Tool. arXiv:2504.04377. (17 languages, 1.91M items. Per-language translation quality and safety-label agreement in Table 10: Korean 81.55, French 82.12)
- 8.Lai, V. D. et al. (2023). ChatGPT Beyond English. arXiv:2304.05613. (Source of the operational definition that treats a CommonCrawl share below 0.1% as low-resource)
- 9.Joshi, P. et al. (2020). The State and Fate of Linguistic Diversity and Inclusion in the NLP World. arXiv:2004.09095. (Resource-tier classification: Korean Class 4, French Class 5, Swahili Class 2)
- 10.Roy, A., Roy, P. (2026). Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages. arXiv:2608.12278, 12 August 2026. (Bengali. Under 0.5% of web content, a 67-to-1 training-token ratio, 36.5% rural against 71.4% urban)
Korean, industry and policy
- 11.Choi, D. et al. (2026). XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity. arXiv:2605.05662, 7 May 2026, CC BY 4.0. (AIM Intelligence, Microsoft, the Korea AI Safety Institute, KT and Seoul National University among others. 10 country-language pairs, 5,500 cases, two native annotators each, agreement in Appendix C.3, attack success rate correlating with nonsensical-response rate at -0.81)
- 12.Seah, E. W. et al. (2026). Improving Methodologies for Agentic Evaluations Across Domains. arXiv:2601.15679, 22 January 2026. (Third joint testing exercise of the International Network of AI Safety Institutes, with Kenya and Korea participating; Kiswahili deep-dive testing at 156 tasks per language)
- 13.OpenAI (2024). GPT-4o System Card. arXiv:2410.21276. (The statement about 45 languages and 29 countries among red-teamers, and the three underrepresented-language evaluations in §5.4)
- 14.European Union. Regulation (EU) 2024/1689 (AI Act), Article 10. (The six data-preparation operations, the representativeness requirement, and documentation of bias and data gaps. Application deferred to 2 December 2027 under the Digital Omnibus)
- 15.DDaily (2026). Coverage of presentations at the second AI Safety Seoul Forum (in Korean). 7 July 2026. (Plain speech 10.1% / polite speech 3.0% / English 23.2%. Cited via a secondary source; the models tested are not named in the coverage)
- 16.Common Crawl. cc-crawl-statistics, languages (CC-MAIN-2026-30). (French 4.8032%, Korean 0.8429%, Swahili 0.0117%, Hausa 0.0026%)
Earlier Pebblous reports
- 17.Pebblous (2026). The Government Built the Training Data. Then Other Agencies Built It Again. 17 August 2026.
- 18.Pebblous (2026). Every Click on the Labeling Screen Becomes Audit Evidence.
- 19.Pebblous (2026). The EU Ordered a 400-Billion-Parameter AI. Maltese Is 0.03% of the Data.