Executive Summary
On 12 August 2026, Korea's Board of Audit and Inspection published the results of an audit into public AI training data. Local governments and public agencies had been building the same datasets the national government had already built. The Ministry of Science and ICT spent ₩7.38 billion (roughly US$5.5 million) on four categories of image data covering wildlife, household waste, concrete cracks and eggs. Five other bodies, including the Seoul Metropolitan Government and the Korea Expressway Corporation, then spent another ₩610 million building something similar. More than half of the household-waste images collected by Seo-gu, a district of the city of Daejeon, resembled images that were already public.
Duplication is a symptom, not a cause. When an agency began a project, it had no practical way to find out what already existed. That absence is what the audit's phrase "no prior review process" actually describes, and it is a catalog and metadata problem before it is a budget problem. Quality follows the same shape. Nine of the 20 largest builders released their data without third-party verification. Quality standards exist in abundance. What did not exist was any party designated to apply one.
What renders data unusable is usually not the data. It is the line that was never attached to it. Some 2,680 CCTV frames collected for autonomous-driving research were unusable. Nothing was wrong with the photographs. The annotation marking where each object sat in the frame was empty. The US Government Accountability Office found something comparable when it reviewed AI inventories at 23 federal agencies and judged only five complete, so this is not a Korean peculiarity. If your company has been pulling public datasets into a training pipeline because they are free and therefore unaudited, this audit is about you.
₩1.63tn
Spent building public AI training data (about US$1.2bn)
Ministry of Science and ICT, 2017–2024
57.8%
Of one district's waste images resembled data already published
5,200 of 9,000 frames, ₩130m project
9 / 20
Agencies that released data with no third-party verification
Among the 20 largest builders by volume
84.2%
Of sampled datasets could not be quality-tested at all
96 of 114 datasets at five agencies had no schema spec
What the audit counted, and what it could not
The Board of Audit and Inspection of Korea, the country's supreme audit institution, ran its audit from September to December 2025 under the title "State of AI Industry Promotion III (AI Training Data)." Its subjects were the Ministry of Science and ICT, the Ministry of the Interior and Safety, the Personal Information Protection Commission and the National Information Society Agency, among others, and the findings were published on 12 August 2026. The audit counted two things. One is the ₩1.63 trillion of budget that went into building AI training data between 2017 and 2024. The other is the 908 datasets that were built with that money and published on AI Hub, the national training-data portal, as of November 2025.
The two figures are measured against different clocks. The first is eight years of cumulative spending; the second is a count at a single point in time. Dividing one by the other to get a cost per dataset would produce a number that means nothing. What they do say together is how this program has been measuring itself. Budget execution and dataset count were the two axes on which performance was reported, and the count kept climbing, from 833 datasets in 2023 to 908 in November 2025.
The stage for this audit, though, is outside those 908. Twenty-six public bodies had published a further 313 datasets on their own portals rather than on AI Hub. These are datasets that provincial governments and state-owned enterprises funded separately, built separately and posted on separate websites. The 313 do not appear alongside the 908 in any single search. Neither list could see what the other contained, and both grew anyway.
| What the audit counted | Figure | Basis |
|---|---|---|
| Build budget | ₩1.6328 trillion | Ministry of Science and ICT, cumulative 2017–2024 |
| Datasets published on AI Hub | 908 | As of November 2025 |
| Datasets published outside AI Hub | 313 | On the portals of 26 individual public bodies |
| Audit period and subjects | Sept–Dec 2025 | Science and ICT, Interior and Safety, PIPC, NIA and others |
Source: Korean press coverage (Kyunghyang Shinmun, Aju Business Daily, Seoul Economic Daily, DigitalDaily, SBS) of the Board of Audit and Inspection's "State of AI Industry Promotion III (AI Training Data)" audit, released 12 August 2026.
What went uncounted sits on the other side of that ledger. What already existed, and how much of it was in a condition to be reused, was nobody's tally until this audit drew a sample by hand. The Board says plainly that the audit rests on samples rather than a census. That is why the sample size shifts from finding to finding: the 20 largest builders here, 114 datasets at five agencies there, six datasets from Gyeonggi Province and the Korea Expressway Corporation somewhere else. A sample result does not settle the whole population. But the fact that nobody knew until a sample was drawn does apply to the whole population.
Read this audit as a story about data nobody uses and you get half of it. AI Hub datasets had been downloaded or otherwise used more than 1.51 million times cumulatively as of July 2026. Because the infrastructure is heavily used, the duplication and the gaps mixed into it propagate quietly downstream. The Board's own summing-up points at demand as well as waste: AI companies keep saying there is no usable data, so the volume and the quality have to improve together.
This closing paragraph is Pebblous reading rather than audit finding. This program had a metric for how much it had built and no metric for how much could be reused. The duplication and the poor quality that follow are two faces of that single fact.
Why the same data got built twice
The duplication was reported at two levels. The first is the level of programs. Across four categories — wildlife, household waste, concrete cracks and eggs — the Ministry of Science and ICT spent ₩7.38 billion building datasets, after which five other bodies spent a further ₩610 million building similar ones. In the pill, oral-cavity and autonomous-driving families, the ministry spent ₩23.3 billion between 2021 and 2023, while three other bodies including the Ministry of Food and Drug Safety put in a separate ₩840 million. The second level is the individual dataset, and that is where the shape of the duplication becomes legible.
| Dataset | Built first by | Built again by | Overlap |
|---|---|---|---|
| Household waste images | Ministry of Science and ICT, 2020 ₩1.6bn · 128 classes · 150,000 images |
Seo-gu district, Daejeon ₩130m · 9,000 images |
5,200 images (57.8%) resembled published data |
| Oral cavity images | Ministry of Science and ICT, 2023 ₩1.753bn · 145,140 images |
Personal Information Protection Commission, 2023 ₩308m · 1,000 images |
All 1,000 were similar |
| Wildlife images | Ministry of Science and ICT, 2021 11 species · about 330,000 images |
Korea Expressway Corporation, 2022 12 species · 60,000 images |
About 25,000 (roughly 42%) were similar |
| Pill identification images | Ministry of Science and ICT, 2021 | Ministry of Food and Drug Safety, 2021 (built separately the same year) |
Images of 1,411 pill types overlapped |
Source: Korean press coverage of the audit findings (Aju Business Daily, Kyunghyang Shinmun, Seoul Economic Daily, DigitalDaily). The wildlife overlap rate appears as both 41.7% and 42% depending on the outlet; these are the same figure.
Read the table down its columns and one feature stands out. Every second builder is small. Seo-gu collected 9,000 images, the Personal Information Protection Commission 1,000, the Korea Expressway Corporation 60,000. Set against the 150,000, 145,140 and 330,000 images built earlier, these are much smaller quantities bought with much smaller budgets. This does not look like agencies imitating a flagship program. It looks like each one procuring the amount its own work required. Seo-gu presumably needed images for municipal waste collection; the expressway operator needed animals on roads.
The Board's phrase for this is the absence of a prior review process. What that phrase points to, though, is not a missing committee or a missing consultation step. SBS reported the Board's description in more concrete terms: agencies drew up project plans and carried them out individually without checking what had already been built. A sentence saying they did not check presupposes that they could have. In this case that premise does not hold.
As the previous section noted, the 908 datasets on AI Hub and the 313 on 26 separate agency portals are not searchable from one place. Even if they were, a problem would remain. Unless a dataset's description records which classes were captured under which conditions and how many images of each, you cannot tell from a title in a list whether it fits your work. A quantitative study of German open-data portals reached the same conclusion: metadata lacking meaningful descriptions makes cross-portal discovery and uniqueness assessment difficult. A catalog with empty descriptions is a catalog you still cannot search.
Read the duplication as a coordination failure and you miss the cause. What produced it was that neither side could look the other up. Each agency reasoned its way to "it doesn't exist, so we'll build it," and the reasoning was sound. What was missing was not judgment but the material to judge with.
One gap in this audit is worth recording. What criterion the Board used to judge two datasets similar does not appear in the published coverage. Whether it was pixel-level comparison, image hashing or overlap in class composition is unknown, and the full audit report itself sits behind a bulletin board that cannot be retrieved programmatically, so this article could not consult it either. Guessing at the methodology would be worse than leaving it open. The gap does connect to the next section, though. If the similarity criterion is not published, no agency can run the same check on itself before starting a project, and everyone finds out four years later through an audit.
The standards exist. Nobody was assigned to apply them.
The Board looked at the 20 agencies with the largest build volumes and asked how far quality management had actually gone. The answer came back as three different numbers, and because they sit at different layers, mixing them together distorts all three. Four agencies had no quality-management standard of their own. Nine did not commission third-party quality verification before accepting delivery. Two required neither third-party verification nor even a self-check by the contractor that built the data. Those two were named: the Ministry of Food and Drug Safety and Korea East-West Power.
4
No quality-management standard at all
No document inside the agency defines what counts as a pass
9
No third-party verification before delivery
The builder graded its own work and the data went public
2
No contractor self-check required either
Ministry of Food and Drug Safety, Korea East-West Power
Skipped inspections are not the whole of it. Some data could not have been inspected even if someone had tried. Of 114 datasets built by five agencies, 96 had no syntax-rule specification. That specification is the document stating which fields a dataset consists of and what format each field must follow. Without it a machine has no baseline to compare against, so quality testing stops being a task that can be run. The figure of 84.2% is not a failure rate. It is the share of datasets for which no grade could be assigned.
Where testing was possible, the scores still fell short. The Board sent six datasets from Gyeonggi Province and the Korea Expressway Corporation to a specialist institute for sample verification. One could not be tested. The remaining five scored between 77.3% and 99.2% on format accuracy. The threshold this program had set for itself is 99.5%. One malformed record in a hundred is invisible to the eye, but a training pipeline either halts on that record or quietly reads the wrong value.
None of these three gaps exists because the standards have not been written. They exist because no one in the program structure was designated to write the specification, to run the test against it, or to answer for the result. What is missing is not a norm but an owner.
The standards and the tools were already on the shelf
On the standards side the position is unambiguous. The ISO/IEC 5259 series covers data quality for analytics and machine learning: overview and vocabulary in part 1, quality measures in part 2, quality management requirements in part 3 and a quality process framework in part 4, all published in 2024, with a governance framework as part 5 in 2025. Part 3 is the only one in the series written in requirements language, so what has to be done reads as a condition rather than a recommendation. By the second half of 2025, when the audit was running, all of this was already in force.
Finding annotation errors has also moved past relying on human eyes alone. A 2023 method uses a trained detection model to diagnose missing boxes, badly placed boxes and wrong class assignments automatically, and other work detects contamination in object-detection datasets from the distribution of features alone, without training. This is not a claim that the tooling is finished. It is a claim that "quality testing was impossible" describes an organisational condition, not a technical limit.
Governments elsewhere also cannot count what they hold
In December 2023 the US Government Accountability Office reviewed AI use-case inventories at 23 federal agencies. About 1,200 use cases had been reported. Only five agencies had submitted complete information for every case. Inventories at 15 agencies contained incomplete or inaccurate data, and two agencies had listed items that later turned out not to be AI at all. The GAO issued 35 recommendations across 19 agencies.
That comparison should be read in the opposite direction from mockery. Governments failing to count accurately what is in their own hands happens everywhere. This is not a case of unusually careless Korean agencies. The difference is not whether the problem occurs but whether a standing mechanism catches it, and what that mechanism looks like is the subject of section 5.
Unusable data is not one problem but two
Press coverage handled the low-quality findings as a single lump, but from the position of someone actually feeding this data into training, two entirely different things are mixed in there. One is absence: something that should be present is empty. The other is corruption: something is present and wrong. You can throw the first kind away and be done. The second kind, if you fail to catch it, goes into the model.
The case of absence is a road-driving CCTV dataset the Korea Expressway Corporation built in 2025. In 2,680 images the bounding-box coordinates were empty. Coordinates here mean the annotation marking where an object sits inside the frame, not the GPS position of the camera. The photographs are fine, but nothing in the file says there is a car at this spot, so an autonomous-driving model has no signal to learn from. The data exists; the training signal does not.
The case of corruption is a bulky-waste dataset the Seoul Metropolitan Government built in 2020. In 538 of 2,303 images the filename and the actual picture did not match. Open a file labelled as a bag and you find a box or a table. That is 23.4%. Feed this into training as it stands and the model learns the appearance of a box under the name of a bag. Where absence drives the training signal to zero, corruption injects a false one.
The diagram separates the two types, but the boundary blurs if you leave the absent annotations in the training set instead of filtering them out. The standard loss function for training an object detector treats any region without a box as background. Put an image with empty annotations straight into training and you teach the model that the cars and people in it are background. A study proposing that missing boxes be redefined as unlabelled regions rather than background demonstrated exactly this point experimentally: the more labels go missing, the more confusing the conventional training signal becomes. Absence is harmless if you discard it, and behaves like corruption if you do not.
How big a number is 23.4%?
To get a feel for a label error rate you need something to compare it against. Northcutt and colleagues had humans re-examine the test sets of ten widely used benchmarks and reported an average of at least 3.3% label errors, rising to at least 6% in the ImageNet validation set. That much alone reverses model rankings. Score against corrected labels and a six-percentage-point increase in the share of originally mislabelled examples is enough for ResNet-18 to beat ResNet-50. Part of what we call one model being better than another is an artefact of label noise.
Seoul's 23.4% is four to seven times that baseline. The conditions are not identical: benchmarks are datasets researchers have polished over years, while this was a new dataset built for operational use. The direction is clear enough anyway. In a world where a 3% error rate destabilises rankings, a mismatch rate in the twenties makes the training itself untrustworthy, well before any question of which model wins.
One further result matters more. Label errors lower the ceiling on the performance any model can reach. Tschirschwitz and Rodehorst argue that contradictory annotations create a point beyond which no model can climb, and estimated that bound on the LVIS object-detection dataset at between 62.63 and 67.52 mAP. Current models already sit in the upper part of that band. The remaining performance comes from fixing labels, not from making models bigger.
Labels, not models, set the ceiling. Adding more data does not raise it. Treat the quality of public training data purely as a budget-waste question and you lose sight of the fact that you are also lowering the ceiling for every model that will ever be trained on it.
Nobody is left to fix it
The third problem is not a kind of error but the lifespan of one. Between 2021 and 2024 the National Information Society Agency built 721 datasets with 1,270 contractor firms. Sixty-five of those firms went out of business or otherwise disappeared, and the datasets they had delivered number 213, or 29.5% of the total. Eleven datasets against which error complaints had been filed were closed out with no correction and no third-party repair, on the grounds that the contractor no longer existed. Those datasets are still published on AI Hub today. The Board instructed the agency to either correct them directly or stand up a third-party maintenance arrangement.
Software has a convention for this situation. A package whose maintainer disappears gets forked or handed to a new owner, and failing that it is marked archived so anyone about to adopt it sees a warning. Data has no such convention yet. The norm does exist, though. Datasheets for Datasets, the foundational work on dataset documentation, already includes maintenance among its documentation items: who looks after this dataset, where errors should be reported, how updates happen. The standard was on the shelf and went unused.
How much was built, and what can be reused
What follows is Pebblous reading. The summary of the audit ended with the previous four sections. What the Board established by sampling four years of projects by hand comes down to three sentences. This dataset has no schema specification, so it cannot be tested. This agency did not commission third-party verification. This dataset overlaps substantially with one already published. None of those three statements is the kind of fact an audit institution alone can establish through special powers. Each is a judgement a machine returns from reading registered metadata.
Somewhere this already happens. The European Union's open data portal, data.europa.eu, runs metadata quality assurance alongside harvesting. It checks each dataset for conformance with the DCAT-AP specification, flags a dataset as non-conformant if even one mandatory field is missing, stores the results in a standard vocabulary, and publishes a quality report per catalog. It also offers an API so publishers can validate their own metadata before registering it. In the United States, DCAT-US requires ten mandatory fields as a condition of registration, covering title, description, keywords, last update, publisher, contact point, identifier and access level among others.
| Portal | Registration requirement | Who checks, and when |
|---|---|---|
| US data.gov DCAT-US v1.1 |
Ten mandatory fields (title, description, keywords, last update, publisher, contact, identifier, access level and others). Distribution, licence, rights and spatial or temporal coverage are conditionally required | A machine, at registration |
| EU data.europa.eu DCAT-AP + MQA |
Conformance with the DCAT-AP mapping. One missing mandatory field marks the dataset non-conformant. Pre-registration validation API provided | A machine, continuously alongside harvesting. Quality reports published per catalog |
| Korea's AI Hub and others Public AI training data |
A syntax-rule specification exists or does not, project by project (96 of 114 datasets at five agencies had none). No common cross-agency registration requirement was identified | A human, four years later, by sample, in an audit |
Sources: resources.data.gov (DCAT-US v1.1), data.europa.eu metadata quality assurance documentation, press coverage of the audit. The Korean row reflects what this audit examined and does not represent every agency.
One qualification belongs alongside the middle column, in fairness. Even where registration requirements are strict, the fields that describe quality are usually optional. In DCAT-US, the dataQuality field indicating whether a dataset meets the agency's information quality guidelines and the describedBy field pointing at a data dictionary or schema are both non-mandatory. The slot equivalent to the syntax-rule specification that 96 datasets were missing here can therefore be empty over there too. Registration requirements make it possible to count what exists and where; whether that data is in usable condition is the next layer's job.
The EU adds one more instrument. Its high-value datasets implementing regulation has applied since 9 June 2024, requiring six thematic categories of data (geospatial, earth observation and environment, meteorological, statistics, companies and company ownership, and mobility) to be made available free of charge, in machine-readable form, through APIs or bulk download. That regulation is confined to high-value datasets already subject to publication duties, so its scope cannot be laid alongside AI Hub, which covers training data generally. What can be compared is who inspects and when, and that is the table's third column.
Korea counted after the fact, by hand. The other jurisdictions block it beforehand, by machine. This audit substituted a four-year sample study for a judgement that should have been automatic at the moment of registration. That is the difference between a system that counts how much was built and one that counts what can be reused.
Duplication costs more than money
The premise that more is better has already been shaken in the research literature. Abbas and colleagues proposed semantic deduplication and reported that removing half the semantically overlapping examples from large-scale web image-text data cost almost nothing in performance, halved training time and actually improved out-of-distribution results. Lee and colleagues found that deduplicating language-model training data cut memorised regurgitation of training text by a factor of ten and reached equal or better accuracy in fewer training steps. The same work found that over 4% of the validation sets in standard datasets overlapped with their training sets, inflating the evaluations themselves.
These results cannot be transplanted directly onto public training data. Both studies dealt with large corpora scraped from the web, whereas this audit dealt with human-labelled domain datasets, and the two sit at different layers. The direction is the same, though. Duplication is a story about ₩610 million spent twice, and also a story about padded data eating into training efficiency and evaluation trustworthiness at once. Stripping duplicates out is quality work, not thrift.
Reuse takes three layers
The machinery that makes data reusable stacks in three layers. The first is the registration requirement: the layer that decides what, if omitted, blocks registration outright, which is where the DCAT family's mandatory fields sit. The second is continuous inspection: a machine repeatedly checking that what was registered still meets the requirement, and publishing the result. The third is documentation written with machine learning in mind. Datasheets and Data Cards, which record what a dataset was collected for, how, and who maintains it; Croissant, which describes datasets so that different tools can read them the same way; and ISO/IEC 5259, which specifies quality measurement and management processes, all live in this layer.
Each of those three layers is a subject Pebblous has taken up in earlier reports. The argument that the catalog is the first layer of the pipeline appeared in our report on OpenMetadata; the argument that publication and reusability are separate problems appeared in our report on research-data pipelines. Institutions that cannot find their own data is not a Korean condition either, as the same structure shows up in our piece on the findability of Māori research data. Set this article next to our report on Korea's planned manufacturing data library and the contrast sharpens. That one is about intake rules that are still blank; this one is about what eight years of building without them produces.
What to check before you download
With the findings out, the institutions are moving. The Board instructed the Ministry of Science and ICT to consult the Ministry of the Interior and Safety and set up a process for reviewing in advance whether existing data can be used instead, and for sharing and coordinating build plans across agencies. It further instructed the ministry to establish common quality-management standards, self-checks and third-party verification procedures. It told the National Information Society Agency to correct errors in data delivered by now-defunct contractors directly, or to build a third-party maintenance arrangement. The ministry has said it will run a government-wide census to verify the quality of existing data and repair deficient material through specialist data and AI firms, will reorganise AI Hub into an integrated AI training data provision system with a pilot launch this year, and will distribute an AI data build and use guide during the third quarter.
Public training data really does flow into models somewhere, and that belongs in the picture too. A survey of AI adoption in the public sector published by Korea's Software Policy & Research Institute in September 2023 found that of 295 institutions answering a question about where they source training data, 36.3% cited public data. That is the second largest channel after internal data at 56.3%. The survey predates the audit by three years and covers public institutions rather than companies, but the figure confirms that public training data is a material used outside the laboratory.
Meanwhile, nobody downloading data can wait for the reform. So here is a checklist usable today. Each item is one of the audit's findings turned inside out, which means it can be checked without any special tooling.
| Check before you download | If you skip it |
|---|---|
| Identify the originating agency and project year | You take two editions of the same data and pad your training set |
| Confirm whether third-party quality verification was done, and by whom | You inherit the builder's own pass mark as if it were independent |
| Ask for the annotation specification and the mandatory-field completion rate | Images with empty coordinates enter the pipeline disguised as training signal |
| Sample-check filenames and labels against actual content | Wrong labels reach the model and lower its performance ceiling |
| Establish the maintenance owner and warranty period | You may be using data nobody is left to fix when you find an error |
| Search for similar published datasets | You repeat, on the consumption side, the duplication this audit flagged |
The deficiency types the Board identified (missing specification, no third-party verification, missing annotations, label mismatch, absent maintenance owner, duplicate builds) restated as procurement-time checks.
After you download, three things. First, sample-test for label errors. Pulling a few hundred images at random and comparing them by eye is enough to surface a mismatch rate in the twenties almost immediately. Second, deduplicate. Start with identical file hashes, and if you have capacity, clear out semantically overlapping examples in feature space as well. Third, check that your validation set does not overlap your training set. That overlap appears most often when public datasets from several sources are combined, and when it does, the only thing that improves is your evaluation score.
None of this says to refuse data that fails these nine checks. It says to know which of them you have not run. The unchecked items are the first places to look when you later try to explain where your model's performance came from.
Why This Matters to Pebblous
The layer this audit found missing occupies one specific position in the path public data travels from creation to use. It is the step that measures and reports absence, duplication and specification violations before anyone takes the data and trains on it. Because that step was empty, the Board substituted a four-year sample study for it. What Pebblous does in data quality diagnostics belongs in that position.
Absence and corruption call for opposite responses
The distinction drawn in section 4 is the one most often flattened in practice. An image with missing coordinates is data carrying no training signal; an image whose filename does not match its contents is data injecting a false one. Bundle both under the single word "low quality" and the response gets bundled too. The first can simply be filtered out; the second gets learned unless it is found. The claim that gaps in training data migrate into a model's internal representations is not an abstraction. It has already been quantified as the result that label errors set the upper bound on model performance.
Free data still needs inspection
Because public data costs nothing to obtain, the acceptance procedures that any team would run on purchased data routinely get skipped. This audit prices that habit. Data with no one left to repair it, such as the 213 datasets delivered by firms that have since closed, remains published, and AI Hub datasets had been used more than 1.51 million times cumulatively as of July 2026. Inspection at procurement is a precondition rather than an option, and the list in section 6 is the minimum needed to begin. An earlier report on data lineage and safety documentation covers the regulatory backdrop.
Editor's Note. Pebblous works on data quality diagnostics and lineage design, so we have an interest in this subject. This article is not a pitch. It is an attempt to record what was missing, given that both the standards and the tools already existed, such that the same dataset got built twice. Pebblous sits on the side that judges whether existing data is in usable condition, not the side that builds more of it. Who performs that judgement, and when, is for the people building and using the data to decide.
References
Academic
- 1.Abbas, A., Tirumala, K., Simig, D., Ganguli, S., Morcos, A. S. (2023). SemDeDup: Data-efficient learning at web-scale through semantic deduplication. arXiv:2303.09540. (Removing 50% of semantic duplicates with minimal performance loss and half the training time.)
- 2.Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., Carlini, N. (2021). Deduplicating Training Data Makes Language Models Better. arXiv:2107.06499. (Tenfold reduction in memorised output; over 4% of validation sets overlap training sets.)
- 3.Northcutt, C. G., Athalye, A., Mueller, J. (2021). Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. arXiv:2103.14749. (At least 3.3% average label errors across ten benchmarks; at least 6% in the ImageNet validation set.)
- 4.Tschirschwitz, D., Rodehorst, V. (2024). Label Convergence: Defining an Upper Performance Bound in Object Recognition through Contradictory Annotations. arXiv:2409.09412. (Label convergence bound on LVIS estimated at 62.63–67.52 mAP.)
- 5.Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., Crawford, K. (2018/2021). Datasheets for Datasets. arXiv:1803.09010. (Maintenance included among the documentation items.)
- 6.Pushkarna, M., Zaldivar, A., Kjartansson, O. (2022). Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI. arXiv:2204.01075.
- 7.Akhtar, M. et al. (2024). Croissant: A Metadata Format for ML-Ready Datasets. arXiv:2403.19546.
- 8.Tkachenko, U., Thyagarajan, A., Mueller, J. (2023). ObjectLab: Automated Diagnosis of Mislabeled Images in Object Detection Data. arXiv:2309.00832.
- 9.Wenige, L., Stadler, C., Martin, M., Figura, R., Sauter, R., Frank, C. W. (2021). Open Data and the Status Quo. arXiv:2106.09590. (Missing metadata descriptions and findability in German open-data portals.)
- 10.Yang, Y., Liang, K. J., Carin, L. (2020). Object Detection as a Positive-Unlabeled Problem. arXiv:2002.04672. (Missing boxes are learned as background under the standard loss, confusing the training signal.)
Policy, standards and statistics
- 11.Board of Audit and Inspection of Korea (2026). "State of AI Industry Promotion III (AI Training Data)" audit findings, released 12 August 2026. (The full report could not be retrieved; cited via press coverage.)
- 12.U.S. Government Accountability Office (2023). Artificial Intelligence: Agencies Have Begun Implementation but Need to Complete Key Requirements. GAO-24-105980, 12 December 2023. (Five of 23 agencies complete; 15 incomplete or inaccurate; 35 recommendations.)
- 13.Commission Implementing Regulation (EU) 2023/138 on the list of high-value datasets and the arrangements for their publication. Applicable from 9 June 2024.
- 14.DCAT-US v1.1 metadata schema. resources.data.gov. (Ten mandatory fields.)
- 15.data.europa.eu. Metadata Quality Assurance (MQA) methodology and validation service documentation. (Automatic DCAT-AP conformance judgement, per-catalog quality reports, pre-registration validation API.)
- 16.ISO/IEC 5259-1 to 5259-4:2024 and ISO/IEC 5259-5:2025. Artificial intelligence — Data quality for analytics and machine learning (ML). ISO/IEC JTC 1/SC 42.
- 17.Software Policy & Research Institute (SPRi) (2023). Survey on AI Adoption in the Public Sector, 22 September 2023. (400 of 408 institutions responded. Among training-data sources, 36.3% cited public data and 56.3% internal data.)
Press
- 18.The Kyunghyang Shinmun (2026). Coverage of the audit findings on public AI training data, 12 August 2026, and follow-up, "₩1.6 trillion poured into public AI data". (In Korean.)
- 19.Aju Business Daily (2026). Detailed coverage of the AI training data audit, 12 August 2026. (In Korean. Sole source for several dataset-level figures.)
- 20.The Seoul Economic Daily (2026). "₩1.6tn of public AI training data: poor quality and weak management" and "₩7.4bn of AI training data built, then ₩600m spent duplicating it", 12 August 2026. (In Korean. Source for the wildlife case's years, species counts and totals.)
- 21.DigitalDaily (2026). Audit findings and the ministry's follow-up measures, 13 August 2026. (In Korean. Source for the bounding-box coordinates, the integrated provision system and the 1.51 million cumulative uses.)
- 22.SBS (2026). Coverage of the audit findings, 12 August 2026. (In Korean. Source for the description of agencies not checking existing data when drawing up project plans.)
Earlier Pebblous reports
- 23.Pebblous (2026). OpenMetadata Completes the AI Ready Data Stack.
- 24.Pebblous (2026). From Stored Research Data to AI Training Data.
- 25.Pebblous (2026). A Manufacturing Data Library That Hasn't Decided What Counts as Correct, 12 August 2026.
- 26.Pebblous (2026). New Zealand's Universities Cannot Count the Māori Data They Hold.