Executive Summary
Korea sits at the very top of the world rankings for the upper storey of open-data policy. The middle storey — the one that translates that policy into specifications a machine can read — still has no single name. Search the project register of Korea's Telecommunications Technology Association (TTA) on 1 September 2026 for a published, revised or in-progress standard called AI-Ready Data and nothing comes back. What is there instead is three separate families of component standards rolling along on their own tracks: catalog vocabularies, data pipelines, and training-data metadata. This report puts those components on one map, works out what Korea's rules for AI-friendly management of public data actually demand, and sets out the order in which an ontology-based semantic layer should be built.
The emptiest cell on that map is the semantic layer. The word ontology does not appear once in the full text of the national catalog standard guide published by the National Information Society Agency (NIA) in October 2025. In the guidelines issued by the Ministry of the Interior and Safety (MOIS) and the NIA in March 2026 it appears exactly once, as a recommended item on an appendix checklist. Both documents fill the catalog and metadata layers densely with international standards, and to the semantic layer they assign, under that name, one line saying check consistency if one already exists and one line saying apply it if needed. Counting words, though, misses something. The requirement to use the same term for the same concept is in the checklist as a mandatory item, and the namespace table already declares the address of the ontology language itself. The requirements are there, scattered; what is missing is a document telling anyone in what order to satisfy them.
On that order, recent evidence points somewhere fairly clearly. A paired-comparison experiment released in April 2026 reported that adding a 4 KB hand-written document to a condition that had been given only the warehouse schema raised natural-language query accuracy markedly, and did so for all three frontier models tested. Pulling the other way, a benchmark paper published the same year opens by noting how often graph-based retrieval-augmented generation is reported to underperform ordinary retrieval-augmented generation on real tasks. The cheapest layer produces the surest gain; the most expensive layer is conditional. The national catalog standard guide says something similar about itself: apply just three standards, it notes, and the basic framework is complete.
This is a map, not a report card. How many times a word appears in a document is a checkable fact; why it came out that way is not something we can know, and we do not guess at it here. All queries were run on 1 September 2026, a date that fell inside the written session of the TTA's big-data standardization committee for the second half of the year.
0×
Mentions of ontology in the national catalog standard guide
Word count over the full 60-page text. Semantic appears once, in the optional-standards table of an appendix
37.8
Score for AI-friendly, high-value data release
2025 statutory assessment of 684 institutions. Its first score, sitting alongside 89.5 for governance
+17–23pp
Accuracy change from attaching a 4 KB document
Paired comparison, 3 frontier models, 99 of 100 questions. From 45.5–50.5% to 67.7–68.7%
1.00
Korea's OECD score for a data-driven public sector
A perfect sub-dimension score and first place in the 2025 Digital Government Index. Korea is also first overall, at 0.95
One concept, three names
This report began in a search box. The TTA's standardization committees keep their project register open to the public. It is a window onto every standardization project filed since 2006, searchable by name, including those still at the drafting, project-adopted and published stages. Typing AI-Ready Data into it is the shortest route to seeing what Korea's standards landscape actually looks like.
As of the 1 September 2026 query, all five search terms returned no registered data. The search itself is not broken. Put ontology into the same window and 37 projects come back, from robotics, energy, sensor networks and addressing; semantic returns 14, and data catalog returns 6.
| Search term | Register results | Control term | Register results |
|---|---|---|---|
| AI Ready | 0 | Ontology (온톨로지) | 37 |
| AI 레디 (AI Ready, Korean) | 0 | Semantic (시맨틱) | 14 |
| 인공지능 준비 (AI-prepared) | 0 | Data catalog | 6 |
| Knowledge graph (지식그래프) | 0 | Data fabric (데이터 패브릭) | 1 |
| Semantic layer (시맨틱 레이어) | 0 | AI data (AI 데이터) | 8 |
Project-name search of the TTA standardization committee register, all committees, all processing stages since 2006, queried 1 September 2026. Search terms were entered in Korean, the language of the register. The five on the left were checked three times: once during planning, once during synthesis, and once more while writing this report.
All this result supports is that, at the time of the query, no project under those names was in the database. Documents at a pre-registration stage — a standardization proposal, or a new-project submission — do not surface in this window at all. And the day we queried fell in the middle of the big-data standardization committee's 54th written session, which ran from 27 August to 2 September 2026 and was reviewing the draft standards for the second half of the year. So rather than writing that no such standard exists, we write that it was not found, and attach the query date, the register searched and the terms used.
1.1Absent from the standards, present in the dictionary
Another window at the same organization gives a different answer. The TTA's ICT terminology dictionary carries this concept as a headword. It is entered as 인공지능 준비 데이터 / 人工知能準備- / AI-ready data, and a search returns six hits: four dictionary entries and two current-affairs terms. The first line of the entry reads:
The entry goes on to say that raw data carries missing values, errors, format mismatches and duplicates, which is why it has to pass through cleaning, transformation, labelling, feature engineering, quality checks and metadata work before it can be used. Then it nails down the difference from preprocessing: preprocessing is the preparatory process, while AI-ready data is the usable state of the data that results from it. Note what the definition includes alongside quality, structure and labels — accessibility and governance requirements. The requirement is not merely a property of the data but of the management system around it, which lines up with how the MOIS guidelines covered in section 2 are organized.
Search the same dictionary for AI 레디 데이터, the direct Korean transliteration of AI-ready data, and you get zero hits — nothing under the spelling that industry and consulting material use most. The MOIS uses a third term again: its guidelines are titled Guidelines for AI-Friendly Management of Public Data. The three names are not strangers to one another, though. Appendix 2 of those guidelines defines the headword as 인공지능 친화적 (Artificial Intelligence Ready) and glosses it as a state of data whose quality, structure and metadata have been organized to suit AI training and inference. The Korean spellings have diverged, but an official document has already recorded that the English name behind them is the same. One concept, then, with three official names — and the fact that the spelling splits from document to document is the linguistic surface of a standards landscape that is itself scattered. This report uses AI-ready data in its own prose, and follows each document's own wording when quoting policy and standards texts.
1.2No umbrella, but the parts are real
Failing to find an umbrella standard does not mean the landscape is empty. The survey found the opposite, in fact. The component standards that make up AI-ready data do exist, and they have been rather busy over the past few years. They are simply running on three separate tracks: catalog vocabularies, pipelines, and training-data metadata.
The top three cells have filled up over the past three years with documents that carry standard numbers. The bottom cell holds two, and neither is a methodology for designing domain vocabulary: one is a domestic adoption of an international standard, the other a technical report.
In those top three cells the dates bunch together. All six standards published in 2025 are dated 5 December. Reading that as a surge in standardization during 2025 would overstate it, because the big-data standardization committee works on a cycle that publishes a year's worth of standards at once at its second-half general meeting. Still, count the 88 standards that committee has listed to date by year and the range runs from three to twelve a year, with five in 2025. Those five are Data Catalog Vocabulary Version 3, RDF Dataset Canonicalization, DataOps Parts 1 and 4, and Data Fabric Part 1 — every one of them in this family. Even allowing for the artificial concentration of a single meeting, the fact remains that the committee spent that meeting entirely on this family. The sixth item, metadata for AI model training datasets, falls under a different committee and is not counted in the 88.
The bottom cell needs a balancing note too, because writing that the semantic layer is entirely empty would be inaccurate. The ontology for data provenance is a domestic adoption of the W3C's PROV-O as a Korean ICT standard, and its scope is the narrow one of data origin and history. It is not a domain ontology standard Korea developed itself. And the RDF Dataset Canonicalization standard published on the same day in December 2025 lists ontology-based data processing among its own fields of application. The materials for a semantic layer are not wholly absent from the standards landscape; what is absent is a document telling anyone what to build with them, and how.
We have published separately on the general conditions AI-ready data has to meet, and on the international-standards axis — the relationship between FAIR and AI-readiness, ISO/IEC 5259, and Croissant. This report checks those concepts against Korea's own standards documents rather than restating the definitions. See the conditions AI-ready data has to meet and making research data AI-ready: standards, quality and governance.
What the guidelines actually demand
The document currently doing the practical work where standards are scattered is the Guidelines for AI-Friendly Management of Public Data V1.0, issued by the MOIS and the NIA on 31 March 2026 and released publicly on 24 April. It runs to 70 pages and can be downloaded from the resource library of Korea's public data portal in both PDF and Markdown. On scope it says only that it applies to every public institution covered by Article 2(1) of the Public Data Act; no institution count is given.
The document is organized into five areas. The principle stated for each is below, translated from the wording the document itself uses.
| Area | Sub-components | Principle (as written) |
|---|---|---|
| 1. Format and structure | Open formats and machine-understandable structure / partial queries and real-time access | Data shall be provided in an open manner and structured in a machine-understandable format. |
| 2. Information about the data | Metadata / documentation | Comprehensive, machine-understandable metadata describing the structure and meaning of the data shall be provided. |
| 3. Quality indicators | Completeness, consistency, accuracy, timeliness, validity, uniqueness | The quality of public data for artificial-intelligence use shall be continuously checked and improved. |
| 4. Standard codes | Application of standard codes | Standard codes shall be applied consistently to secure data linkage and interoperability across institutions. |
| 5. Management system | Accountable party / ease of access / feedback / access rights / governance / resolution of legal conflicts | The party accountable for data quality, security, ethics and updates, and the management system for them, shall be clearly designated. |
Source: MOIS and NIA, Guidelines for AI-Friendly Management of Public Data V1.0 (March 2026), Table 4. Checked directly against the Markdown original, queried 2026-09-01. English rendering ours.
That table carries only the headline sentence for each area. The principles themselves are set out separately in Appendix 1 as a numbered list running from Principle 01 to Principle 12. The fourth is the one that bears most directly on this report's argument.
The document does not stop at demanding documentation; it lists which documents an institution must hold. A data card, a schema definition, a metadata specification, an API access document, and a quality diagnostic report. The schema definition carries data dictionary as a parenthetical alternative name, and the text asks for the physical and logical names of each variable, its data type, its unit, and the relationships between fields. Why that list matters becomes clear again in section 4: the layer that produced the surest gain in a 2026 controlled experiment was exactly this one — a document recording the definitions of measures, their units, and the rules for telling ambiguous cases apart.
Beneath those five areas and twelve principles sit six quality indicators, and an appendix checklist of 38 items carries the practical work. The checklist splits into six common areas and three data-type-specific ones, with 18 mandatory items and 20 recommended.
| Checklist area | Mandatory | Recommended | Total |
|---|---|---|---|
| [Common] Format and structure | 5 | 2 | 7 |
| [Common] Meaning and standards | 3 | 2 | 5 |
| [Common] Quality and history | 3 | 6 | 9 |
| [Common] Management and security | 2 | 2 | 4 |
| [Common] Data provision | 0 | 4 | 4 |
| [Common] Maintenance | 2 | 2 | 4 |
| [By type] Numeric · image/video · personal-data protection | 3 | 2 | 5 |
| Total | 18 | 20 | 38 |
Appendix 4, management checklist. It is a self-assessment marked in three columns — adequate, insufficient, not applicable — and carries no score. The single item in which ontology appears is on the second row above, a recommended item in the meaning-and-standards area.
What the very first principle — structure the data in a machine-understandable format — actually demands is spelled out in chapter 2. Formats come first. Recommended: ODT, CSV, JSON, XML, Parquet and ORC for text; GeoJSON, GeoPackage and GeoTIFF for geospatial; PNG, JPG, TIFF and SVG for images; IFC, OBJ and GLTF for construction and architecture. Then it draws up a separate list of formats it does not recommend, on the grounds of vendor lock-in.
* Not recommended means unsuitable as the primary delivery format for AI use; it does not restrict use for preserving originals or for internal administrative purposes.
That puts HWP — the Hangul word-processor format that is the de facto default for Korean public documents — and Excel formally outside the recommended set for primary delivery. A reason is attached: proprietary formats that require paid software are hard to process automatically on cloud or Linux-based GPU servers, or incur a separate conversion cost. The document also explains why it prefers columnar formats: during training you can read only the columns you need instead of every row, which cuts I/O.
On structure it sets hierarchical and flat forms side by side and pins down three conditions for the flat case. Each row is one unique observation, each column is one unique attribute, and each cell holds exactly one value. Anyone who has fought a merged-cell spreadsheet will recognize the practical instinct in those three lines. Value formats are fixed to international standards as well: ISO 8601 for dates and times, ISO 3166-1 for country codes, ISO 4217 for currency, ISO 639-1 for language, WGS84 and ISO 6709 for geographic position, and SI for units of measure. Missing values are to be marked consistently so that they read as missing rather than as errors. For access it asks for a RESTful API, and adds that a response should carry not just values but the schema, the field definitions and a description of what they mean, so that structure and sense can be grasped at once.
The six quality indicators are not a bare list either; each carries an existing definition and an extended one, in two layers. The contrast the document draws for itself is that conventional public-data quality centres on structural conformity, while AI-friendly public-data quality centres on semantic completeness. Consistency, for instance, is conventionally about whether the same entity, code and attribute value are applied without contradiction across a dataset; the extended definition asks whether the same concept is used with the same meaning across different contexts, so that no confusion arises in the course of an AI's interpretation. The extended definition of completeness brings in representativeness and balance; that of uniqueness brings in the risk that duplication leads to bias, overfitting and data leakage. Each indicator comes with diagnostic criteria, a diagnostic method and examples of errors, which makes this chapter read, in practice, like a requirements specification for an automated diagnostic tool.
2.1Four bars to clear at registration
The first thing a practitioner runs into is the minimum quality requirement for shared data. The guidelines present four of them as principles, with a parenthetical noting that the thresholds are examples. They are illustrations rather than settled rules, so we carry that qualifier along with the quotation.
- Completeness of content — the share of missing values in the dataset should be under 20%, with no core information omitted.
- Fullness of metadata — all eleven fields recorded: managing institution, production cycle, last update date, data period, record count, column count and description, data grade, whether de-identification was applied, permitted scope of use, relevant legislation, and contact details for the responsible officer.
- Currency — no more than two update cycles should have elapsed since the last update.
- Conformance to technical standards — one of the platform's standard formats, such as CSV, JSON, XML or Parquet.
Data is assigned one of three grades. Grade 1, where there is no personal-information or security concern, is openly shared; grade 2, made safe through de-identification, is shared with restrictions; grade 3, where external sharing remains inappropriate even after de-identification, is internal only. Access scope splits four ways again by type of institution. On licensing, a new AI type, for machine-learning use, has been added to the five existing Korea Open Government License types; it permits commercial and non-commercial use and derivative works without even an attribution condition — a licence type carrying fewer conditions than KOGL Type 1.
2.2Why a voluntary document exerts pressure
This document is what its name says: guidelines. The appendix checklist is an unscored self-assessment, and there are no sanctions anywhere in the text. It does note that data posted to the sharing platform will be quality-checked once a year, that institutions doing well may receive quality certification, and that those falling short may be asked to accept restrictions on data sharing or to take corrective action.
The more direct route runs outside the document. On the same policy line sits the statutory assessment of public-data provision and operation, which scores 684 institutions under Article 9 of the Public Data Act. From the 2025 round, the release of AI-friendly, high-value data was added as a new indicator. The guidelines themselves are advisory, but the state they call for comes back as real pressure through the scorecard of a statutory assessment.
Here is how the first year came out. Average scores by area in the 2025 assessment were 89.5 for governance, 72.5 for quality and 59.2 for release and use. The newly added AI-friendliness indicator scored 37.8.
Governance already scores near the top and quality is in the seventies, while the first indicator to put a number on AI-friendliness comes in markedly lower. Second-hand citations that report 37.8 as the score for the whole release-and-use area are in circulation, so it is safer to state which level the number belongs to.
The overall trend is not bad. The share of institutions rated good or better rose from 36.2% in 2023 to 40.9% in 2024 and 50.9% in 2025, putting 348 institutions in that band. By type, public enterprises and quasi-governmental institutions lead at 92.5 and central government bodies at 90.2, while basic local governments trail at 60.3 and other public institutions at 57.4. The shape of the 2025 report card is a general rise with the AI-friendliness indicator starting low.
2.3The standards the guidelines invoke, and the ones they do not
The guidelines state that they do not write their own metadata specification but defer to international standards. The relevant sentence and its footnote read:
Footnote: W3C, Data Catalog Vocabulary (DCAT) Ver.3 · DCMI, DCMI Metadata Terms · MLCommons, Croissant Specification
The figure describing the development methodology also names the foreign literature consulted. Carried over as the figure writes them: ODI, A Framework for AI-Ready DATA; US Department of Commerce, Generative AI and Data Guidelines; the EU AI Act; the OECD AI Principles; and the national data catalog standard guide. The figure uses shortened names, with the full titles in footnotes and a separate table. The Commerce document's full title is Generative Artificial Intelligence and Open Data: Guidelines and Best Practices (2025), and footnotes throughout the body cite it down to the page — 14, 11, and 23–24. The Open Data Institute paper is cited at pages 8 and 14. The reference list also includes the UK government's AI Opportunities Action Plan and material from the Open Data Policy Lab, which shows these guidelines did not skim the foreign literature but read it page by page.
Across the 92,690-character full text, TTA, Telecommunications Technology Association, and Korean ICT standard each appear zero times. DCAT appears 25 times, Croissant 5, and ISO 9. We counted where those nine ISOs land. Most of them are the standards that fix the format of a value — ISO 8601, 3166-1, 4217, 639-1 and 6709 — plus a passage introducing Dublin Core as an international specification. No international standard for data quality appears among the nine. Korea's own standards body revised DCAT Version 3 into a domestic standard in the same period and already holds a Korean DCAT application profile as a standard, yet the policy document bypasses that route and defers directly to the international specifications. Why that choice was made cannot be established from the document alone, so all we record is the word frequency. The pattern of the policy layer and the standards layer not calling on each other turns up again, in the same shape, in section 3.
Not naming an organization is a different matter from not using its vocabulary. The same document maps each field of the public-data registration form prescribed by the enforcement rules of the Public Data Act onto an international vocabulary property. Name of the public data is dct:title, description is dct:description, permitted use is dct:rights, update cycle is dct:accrualPeriodicity, location is dcat:landingPage, and media type is dcat:mediaType. The entries on a statutory form and the properties of a catalog vocabulary are already lined up in a single table. A companion figure treats the existing registration fields as base metadata and extends them with metadata for data type, use, lineage and quality.
The catalog comes first
The other document the MOIS guidelines defer to is the NIA's Data Catalog Standard Guide for National Data Integration and Linkage v1.0. It appeared in October 2025, issued by the national data infrastructure team of the NIA's AI data division. Its colophon notes that the guide is an output of an ICT and broadcasting R&D project funded by the Broadcasting and Communications Development Fund, and that anyone presenting its contents must state that they are research results of a Ministry of Science and ICT (MSIT) project. Secondary sources that introduce this document as an MSIT publication therefore have some basis for doing so. The MOIS guidelines themselves attribute the guide to the MSIT where they list deferred-to documents, and render its title two different ways — once as a data catalog standard guide for national data interoperability, once as one for national data integration and linkage. The precise formulation writes both facts at once: published by the NIA, and an output of an MSIT R&D project.
The document has a single purpose: to standardize which fields, under which names, scattered data platforms should fill in when they push a dataset listing to One-Window, the national integrated data platform. Its scope statement specifies that this covers not only system-to-system linkage but access by AI agents.
3.1The international vocabulary is kept whole; only the domestic items are new
The structure has two layers: five classes of common items taken straight from international standards, and three classes of extended items newly defined for the practicalities of domestic data distribution. The extended items get their own namespace, ndi:.
| Group | Class | Property count as written | Example mandatory properties |
|---|---|---|---|
| Common (defers to DCAT-AP 2.1) | Catalog | 19 | dct:title dct:description dct:publisher |
| Dataset | 35 | dct:title dct:description | |
| Data Service | 7 | dct:title dcat:endpointURL | |
| Distribution | 23 | dcat:accessURL | |
| Catalog Record | 9 | foaf:primaryTopic dct:modified | |
| Extended (newly defined as ndi:) | Data Product | 10 | ndi:viewCnt ndi:downloadCnt ndi:srvcType ndi:pltfrmAccessToken |
| Data Curation | 10 | None (all optional) | |
| Data Quality | 4 | None (all optional) |
Source: Data Catalog Standard Guide for National Data Integration and Linkage v1.0, Table 5. The property counts for the common items are the numbers the original prints next to each class name. They sum to 93, but that total is ours, not one the document calculates.
The passage on the practicalities of linkage has the smell of the field about it. The principle is to convert existing XML metadata into an RDF catalog by script and to strengthen semantic expression by attaching curation and profiling information — and then an exception arrives on the very next line.
The standard is RDF, and there is a second door at the entrance marked Excel. This need not be read as a flaw. Linkage has to actually run, and that door is what makes it run; what remains a separate question is how far up the stack the values arriving through it are machine-readable. It is one more reason the layer distinction in section 4 is worth drawing.
3.2Two DCAT generations, five months apart
The left-hand column of that table records what the guide defers to: DCAT-AP 2.1. The same designation appears in the body text describing the basic structure and in the caption of a reference figure. The RDF example in Appendix 3 embeds the URL of the 2.1.0 release change log in the European SEMIC repository. Yet the term definitions in Appendix 1 of the same document give DCAT's source as DCAT V3. The designation splits even within one document.
And the MOIS guidelines that followed five months later state in a footnote that they defer to the W3C's DCAT Version 3. In between, the TTA published a revision adopting Data Catalog Vocabulary Version 3 as a domestic standard. Laid out in time order, it looks like this.
The two orange points are the Korean policy documents. Five months apart, they ended up deferring to different generations of DCAT, with the TTA's domestic revision falling in between.
Reading this timeline as a story about one document being out of date would be wrong, because the two sit at different levels. DCAT-AP is an application profile in which the EU adds constraints for a European context; DCAT is the base vocabulary the W3C maintains. It is not a relationship in which 2.1 is a lower version than 3. What remains is an arrangement of facts: two Korean policy documents, five months apart, ended up deferring to different generations of the DCAT family, with a domestic standard revision in between.
The generational difference is not merely nominal, though. The DCAT-AP 3.0 specification records that several property addresses used in 2.x were deprecated and moved into the DCAT namespace — dct:hasVersion became dcat:hasVersion, dct:isVersionOf became dcat:isVersionOf, and so on. Since the table above lists dct:hasVersion as an optional property of Dataset, the generational split shows up at the level of actual property names.
3.3Two Korean DCAT extensions exist side by side
The extension side is a little more tangled, because a DCAT extension for Korean data portals already exists as a TTA standard: the DCAT Application Profile for Data Portals of the Republic of Korea (DCAT-AP-KR), standard number TTAK.OT-10.1406, published on 7 December 2022, 41 pages. Its standard summary reads:
Open the public specification page and the newly defined vocabulary turns out to be ten terms under the dcatkr: prefix. The ndi: extension the national data catalog standard guide defined three years later contains items of the same kind. View count and download count appear in both vocabularies, under different names.
| Item | TTA standard DCAT-AP-KR | National data catalog standard guide |
|---|---|---|
| Status | Korean ICT standard TTAK.OT-10.1406 (2022-12-07) | Guide v1.0 (2025-10) |
| Extension prefix | dcatkr: | ndi: |
| Base | DCAT-AP 2.1.0 | DCAT-AP 2.1 |
| View count | dcatkr:numberOfView | ndi:viewCnt |
| Download count | dcatkr:numberOfDownload | ndi:downloadCnt |
| Request count | dcatkr:numberOfRequest | Not applicable |
| Other extensions | maintainer fee legalBasis numberOfRow derivedSystem nextRegistrationDate numberOfRequestLimit | 24 properties across three classes: product, curation, quality |
Source: TTA standard summary page (TTAK.OT-10.1406), the public DCAT-AP-KR specification at vocab.datahub.kr, and the national data catalog standard guide, Table 5. Queried 2026-09-01.
The two vocabularies are not strangers. The national data catalog standard guide's bibliography runs to eight lines, and one of them is the URL of that TTA standard's summary page. Yet nowhere in the 60 pages of the body does the name DCAT-AP-KR or the dcatkr: prefix appear. Whether the relationship between the two vocabularies is succession or citation, and whether a mapping table exists between them, could not be established from public documents. So what can be said here stops at this: two Korean DCAT extensions, both built for interoperability, exist in parallel. Calling that a conflict or duplicated investment would be premature before the mapping question is settled.
The impression that Korea came late to catalog vocabularies is not accurate. The TTA's data catalog vocabulary standard was established on 13 December 2017 and has been revised twice, to version 2 in 2020 and version 3 in 2025. The application profile for domestic portals became a standard in 2022. The materials have been around for a long time. What is missing is a single name that gathers them, and the semantic layer to sit on top.
The prov:wasGeneratedBy that the national data catalog standard guide lists as an optional property of Dataset is the same vocabulary the TTA adopted domestically on 10 December 2020 as Ontology for Data Provenance / PROV-O: The PROV Ontology. It is one of the rare points where a policy document and a domestic standard actually mesh on the same vocabulary. How controlled vocabularies differ structurally from free-text entry is something we covered separately in an earlier report on metadata catalogs.
The order of building a semantic layer
That is the map of what exists. What is left is the empty cell. Counting, by the same yardstick, how often words belonging to the semantic layer appear in the two policy documents brings the terrain into view at a glance. For the catalog standard guide we counted plain string occurrences across the full 60-page PDF; for the AI-friendly management guidelines, across the 92,690-character Markdown original.
Above the dotted line the space is packed with the names of international standards; below it you can count on your fingers. Because the two documents differ in length, read each column against itself, top to bottom, rather than comparing the left figures with the right.
What matters more than the counts is where each of the three lines below the dotted line actually sits. In the catalog standard guide, the single occurrence of semantic is in the table of standards applied optionally, in Appendix 4.
In the AI-friendly management guidelines, the single occurrence of ontology is a checklist item in Appendix 4, in the meaning-and-standards area, graded as recommended.
That is a sentence about checking consistency with one that already exists, not about building one, and no method for building one appears in the document. Rather than speculate about why, we record only where the two sentences sit. One row in a table of optional standards and one recommended checklist item: that is the page space assigned under that name.
Counting words has its limits, though, and this report thinks it more honest to state them, because requirements that belong to the semantic layer have entered these documents under other names. In three places. First, the namespace reference table both documents carry includes owl, described as defining relationships between classes and properties, with prov alongside it. The address of the ontology language is already declared in a table. Second, the meaning-and-standards area of the checklist holds a mandatory item beyond the recommended one quoted above: it asks whether the same concept is consistently used with the same label and term. Third, the extended definitions of the quality indicators seen in the previous section demand semantic consistency and semantic completeness. The word semantic appears ten times in that document.
So the accurate statement is this. The requirements of the semantic layer are already present in Korea's policy documents, scattered across several places, down to the declared address of the vocabulary. What is missing is a document saying in what order and by what method to satisfy them. Nowhere does any chapter tell the officer who has just been handed a mandatory item about using the same term for the same concept how to actually do it. It is the same shape as the standards landscape from section 1: the materials are there, and the assembly instructions are not.
4.1Separate the layers and the order appears
The phrase semantic layer confuses because it points at several layers of wildly different cost at once. A 4 KB explanatory document sitting next to a spreadsheet is a semantic layer; so is a system with a query-rewriting engine on top of a domain ontology. The cost of building the two differs by orders of magnitude, and so does the certainty of the payoff. The experiment in the next section sorts this split into two names: the runtime form, where definitions are written as code and queries compile deterministically, and the context form, where the same content is written as a document for the model to read. Once the layers are separated, where to start becomes much clearer.
The two orange layers are the ones Korea's standards and policy documents already specify densely; the three below them are barely addressed. Evidence strength is marked according to the experimental results discussed in the next section.
4.2The difference a 4 KB document made
The most recent attempt to measure whether a semantic layer actually works, by paired comparison, went up on arXiv on 28 April 2026. The authors put a retail dataset on ClickHouse in front of three frontier models and asked 100 natural-language questions. The models were Claude Opus 4.7, Claude Sonnet 4.6 and GPT-5.4. Each was evaluated twice: once given only the warehouse schema, and once given the schema plus a 4 KB hand-authored Markdown document describing the dataset's measures, its conventions, and its rules for resolving ambiguity. The design hands over the CREATE statements for all 25 tables and asks for a single-shot answer, with no tools, no retries and no execution feedback.
One disclosure belongs before the citation. The authors' affiliation is Cube, a company that sells a semantic layer as a commercial product. The paper itself, reviewing related work, notes that vendor-published benchmarks should be read as directional rather than peer-reviewed — a standard that applies to this paper too. We cite the experiment because its design is controlled and because the code, the questions and the judging records are public and can be re-checked; and we use it for the direction and the ordering it implies, not for the magnitude of its result.
The result is that adding the document raised accuracy by 17 to 23 percentage points in all three models. With the document, the three models fall between 67.7% and 68.7% and are statistically indistinguishable; without it they fall between 45.5% and 50.5% and are again indistinguishable. Every cross-cluster comparison was significant at p < 0.01. The authors read the result this way:
A single 4 KB hand-written document, laid over the same schema, lifted accuracy by roughly 20 points, and the size of that lift was nearly identical across all three models — the spread the authors' "changed what the model was asked to do" reading is built on.
One corollary from the paper's practical-implications section belongs here, because it is the question every team asks first. Does the choice of model matter? Within either condition, on this benchmark, it did not: two Anthropic models spanning a wide range of size and cost, plus one OpenAI frontier model, all landed within a five-point spread whether or not the semantic layer was present. The authors' reading is narrow and worth quoting in substance — with a well-authored semantic layer, model choice can be optimized for cost and latency without sacrificing accuracy, and without one, moving to a stronger or more expensive model is not a substitute for writing one. Note also what this is not: a claim that a semantic layer makes a model better. Turning the result into that causal statement is precisely the reading the authors rule out.
Look at the kinds of error the document prevented and it becomes clear why this overlaps with public data. The first type the paper lists is confusing a snapshot with a flow. The inventory table is a daily snapshot, one row per date for each product-and-store pair, and models in the no-document condition would sum on-hand quantity across every snapshot date, producing values roughly 1,000 times the correct magnitude — absurd numbers that are not visibly ill-formed. The second is the reference point in time: the data ends on 31 December 2009, but the model reads last quarter relative to today and returns an empty result. The third is values whose convention is undocumented. Promotion key 1 is a placeholder meaning no discount and must be excluded from discount totals, and the working-day flag is stored as a string rather than a boolean.
The three error types share one property. The generated query is syntactically valid, references real tables and columns, executes without error, and returns rows. The paper calls this a silent hallucination and argues it is more dangerous than a visible error, because the value flows into a business decision with no signal anywhere that something went wrong. The MOIS guidelines also put hallucination in their Appendix 2 glossary, defining it as a response that looks plausible but is wrong. The guidelines' demand that update cycles, data periods, missing-value notation and code systems be recorded as metadata is, read this way, a list of the materials that prevent exactly this kind of quiet failure.
Three qualifications travel with that result. First, roughly 30% of errors remain even with the document attached. Something in the high sixties cannot be called an accuracy you would trust unattended. What this experiment establishes is that the cheapest instrument available closed the stretch from about 45% to about 68%. The residual errors cluster on calculations the document never defined — percentiles, correlations, standard deviations, threshold-based rankings, multi-level CTEs — and the authors treat that as a problem to be fixed by extending the document rather than as a limit of the approach. Second, the experiment measured only the form in which the document goes into the prompt. The paper divides semantic layers into a runtime form, where definitions are written as code and queries compile deterministically, and a context form, where the same content is written as a document for the model to read; it states that what it measured is the latter, and that this represents a lower bound on the former. In the context form the model receives the rules but is not compelled to follow them. Third, the judge belonged to the same model family as two of the systems under test, and the analyst who wrote the document had already seen the dataset. The authors record both as limitations, and add that in a side experiment in the repository, an agentic approach that used tools to go looking for the information it needed did not beat the schema-only baseline.
A second piece of evidence pointing the same way comes from the chief data office at AT&T, in a study on automatic metadata extraction. Given no oracle information at all — no schema or value hints drawn from the reference answers — and working from profiling information alone, their system scored 67.41% on the BIRD benchmark, against 57.13% for the next-best submission under the same condition. Oracle information does not exist in a live service, the authors argue, which is why this condition better represents how such a system actually behaves. Quote 67.41% without the condition attached and you lose what it is being compared against, so we carry both. The same paper also lists the leaders under the hint-using condition: 77.14 in first place, then Google at 76.02, Contextual AI and Alibaba at 75.63.
This study touches practice through its diagnosis more than through its scores. The authors state flatly that the hardest part of query development is understanding what is in the database, and that once that is done, writing the query is comparatively simple. Then they put a number on how little of that understanding survives in documentation: an analysis of query logs found that 25% of join paths were undocumented. They describe joining two well-documented tables on a device identifier and getting an empty result, because one side stored 13 digits and the other 14, with a leading 1 on the longer one. It is hard to show more briefly why the MOIS guidelines' demand that the schema definition record units and inter-field relationships is not a paperwork requirement.
A gap of the same size shows up in people. According to the original BIRD benchmark report, as cited in the paper above, attaching a single sentence of external knowledge to each question lifted GPT-4's accuracy from 34.88% to 54.89% — and lifted human experts on the same questions from 72.37% to 92.96%. A gap of roughly 20 percentage points, almost identical for machine and human. Floundering in front of a context-free schema is not a peculiarly model-shaped problem, which also means that filling in the data dictionary is less an AI-readiness task than something that always needed doing.
4.3The most expensive layer is conditional
Evidence pointing the other way arrived the same year. The GraphRAG-Bench paper at ICLR 2026 opens its problem statement like this: despite being conceptually appealing, recent studies report that graph-based retrieval-augmented generation underperforms ordinary retrieval-augmented generation on a range of real tasks. So the paper's question shifts from whether graphs work to under what conditions graph structure delivers a measurable gain.
Its answer turns on the nature of the task. Working across two corpora, medical guidelines and a novel, and escalating difficulty from fact retrieval through complex reasoning and contextual summarization to creative generation, the authors gather their findings into nine observations. The first two answer the question of order. On simple fact retrieval, ordinary retrieval-augmented generation matches or beats the graph approaches; a baseline method with reranking achieved the highest evidence recall on fact retrieval over the novel corpus, at 83.2%. In this band, the authors explain, graphs pull in logically related but unnecessary information and degrade answer quality. On complex reasoning and summarization, where information scattered across several passages has to be joined, the graph side is clearly ahead: evidence recall on the same corpus's Level 2 and Level 3 questions rises to between 87.9% and 90.9%.
The same paper measured cost. Prompt tokens per query run to around 900 for the ordinary approach, while graph approaches range from the low thousands into the tens of thousands depending on the implementation. For an implementation that runs community summarization, the text records growth from 7,800 to 40,000 tokens as task difficulty rises. A table in one of the paper's figures gives larger values for the same implementation, at odds with the body text by an order of magnitude, so we go no further than saying the orders of magnitude differ and do not compute a multiple. Nor is extracting more relations by itself the answer: the authors recommend building a densely connected graph rather than a large one. The best-performing implementation created many edges while also achieving the highest neighbourhood density, and kept its prompt at around 1,000 tokens.
Set the two bodies of evidence side by side and the layer distinction becomes the conclusion. The gain from a 4 KB document and the gain from a full knowledge graph are not the same story. The cheapest layer produces the surest gain; the most expensive layer is conditional. And that conclusion runs in exactly the direction the national catalog standard guide points in its own appendix.
A national standard guide and an academic experiment point to the same order. One could read the policy documents' near-silence on the semantic layer as an omission — but on the question of order, at least, the advice those documents wrote down and the 2026 experimental results do not disagree.
None of which means layer 3 and above are useless. Cases where ontology-based data access worked in industry have been in the literature for years. The best known is Norway's Statoil, now Equinor, where a system built in an EU project brought seven exploration data sources under a single ontology. One repository alone held roughly 3,000 tables and 37,000 columns. The result the follow-up literature keeps citing is a description: for data that could not be queried at all without an IT specialist, the time for a geologist to ask a question in their own terms and get an answer fell from weeks to minutes. That is a case description rather than a quantitative benchmark, so we quote it as written instead of reworking it into a multiple or a percentage.
There is a measured strand too. Among the related work catalogued by the paired-comparison paper is a 2023 study on an insurance database of 199 tables. Checking the original abstract: when GPT-4 was asked to write SQL directly, accuracy was 16%; when the same data was expressed as a knowledge graph through an ontology and mappings and questions were put on top of that, it was 54%. A follow-up the next year added a layer that verifies queries against the ontology and feeds back explanations of the errors for repair, raising overall accuracy to 72%. And accuracy was not the only figure that follow-up recorded. A further 8% of answers were returned as I don't know, and the overall error rate fell to 20%.
What changed was not the size of the errors but their kind. As the two studies observed, a model bound to an ontology does not invent classes or properties that do not exist; when it is wrong, it is wrong by taking the wrong property path. Set against the silent hallucination of the previous section, the contrast is sharp. Failure in the schema-only condition takes the form of a plausible number delivered with confidence; failure in the semantically bound condition takes the form of an admission of ignorance or a rejection at the verification layer. Measure the benefit of layer 3 only in percentage points of accuracy and that difference disappears. Where the value goes straight out to the public, making the wrongness visible may matter more than reducing the number of wrong answers. The context form measured in the previous section does not have this property: a document tells the model the rules, but it cannot stop the model breaking them.
4.4Why starting with the ontology stalls
A diagnosis that recurs in the 2026 literature bears directly on this question of order. It comes out of the argument that FAIR principles and AI-readiness are not the same thing, and what gets flagged inside that argument is the scalability of the methods used to achieve FAIR. A governance officer hand-crafting an ontology does not scale to the volume of data products an organization now manages. Datasheets for Datasets, the archetype of dataset documentation standards, was itself designed as a hand-written document for capturing what only the data's creators know, not as an automation tool. A large-scale analysis of Hugging Face dataset cards demonstrated the gap empirically: the tools have been proposed, and adoption remains low.
So the movement through 2025 and 2026 has been toward automatically filling in the fields people used to fill by hand. Tools that generate Croissant metadata automatically have appeared, and the automatic metadata extraction study mentioned above belongs to the same family. In Korean academia there is follow-up work combining DCAT-AP-KR with knowledge-graph-based data map search. Research, in other words, is already walking the order in which the knowledge graph comes after the catalog.
Translated into an order of work, the evidence so far comes out as follows.
- Put the catalog in order first. Without a list of what is where, anything you stack on top has nothing to land on. Korea already has two sets of specifications for this — the TTA standards and the national catalog standard guide — so it is not a design job.
- Next, fill in the data dictionary and the documents. That means the definitions of measures, their units, whether a figure is a snapshot or a cumulative total, conventions that hold only inside your organization, the rule for telling apart two things that share a name, and the date at which the data stops. This is the layer where 4 KB of Markdown produced a measurable effect, and the document the MOIS guidelines ask for under the name schema definition.
- Then build the domain vocabulary. This step is expensive, so start it after the first two layers have shown you which ambiguities you keep colliding with. One more criterion applies where answers go straight out to the public or into decisions: whether you need the system to say it does not know when it is wrong. That property comes from a verification layer, not from a document.
- Add a knowledge graph when you actually have tasks in which multi-hop reasoning or structural relationships determine the answer. If most of the traffic is simple lookup, the benchmark observed that ordinary retrieval matches or beats graphs; the band where graphs pull ahead was questions requiring information scattered across several documents to be joined. If you do build one, build for neighbourhood density rather than a large count of extracted relations.
- Whatever stage you are at, measure the current state first. If you do not know the state of your missing values, your structure and your lineage, you cannot know which layer you are stuck at either.
We have covered construction methods layer by layer elsewhere on the Pebblous blog. The pipeline for extracting a semantic layer from unstructured documents is in a benchmark on building semantic layers from unstructured data; where quality breaks down when an ontology is built automatically is in our report on automatic ontology construction. The neuro-symbolic angle is gathered in the neuro-symbolic ontology hub.
Korea placed next to everyone else
Everything so far has been about an empty cell, which makes it easy to come away with an impression that Korea is behind. Put the international comparisons on the table and the picture inverts. The OECD published Working Paper No. 90 on 12 February 2026, approved by its Public Governance Committee, carrying the results of the 2025 Digital Government Index and the Open, Useful and Re-usable Data Index. Data collection ran from January 2023 to December 2024.
| Indicator | Korea | Rank | For reference |
|---|---|---|---|
| Digital Government Index, overall | 0.95 | 1st | Australia 0.88, Portugal 0.86 |
| └ Data-driven public sector | 1.00 | 1st | Czechia 0.94, Estonia 0.93. The OECD average in this dimension rose from 0.63 in 2023 to 0.74 in 2025, the largest gain of any dimension |
| OURdata Index, overall | 0.95 | 2nd | France 0.96, Poland 0.82. OECD average 0.53 |
| └ Data availability | 0.90 | Joint highest with France | Korea's lowest score among the three pillars |
| └ Data accessibility | 0.97 | 2nd | France 0.98 |
| └ Support for re-use | 0.98 | 2nd | France 1.00 |
Source: OECD, Digital Government Index and Open, Useful and Re-usable Data Index: 2025 Results and Key Findings, OECD Working Papers on Public Governance No. 90, approved 2026-02-12. Checked directly against the PDF, queried 2026-09-01.
Measured by institutional and policy maturity, Korea is at the very top. The relatively lower score within that is data availability, the item that looks at how far the volume and range of actual data has widened. It means 0.90 alongside 0.97 and 0.98 on the other two pillars, not a low score. But the arrangement rhymes with the shape this report has been describing. The upper storey is densely built, and the middle storey that translates it into concrete machine-readable specifications is comparatively less filled in.
5.1Europe writes even its versioning policy down
The EU's DCAT-AP has its release history stacked up in a public repository: 2.0.1 in June 2020, 2.1.0 in December 2021, 2.1.1 in August 2022, 3.0.0 in July 2023, an extension for high-value datasets in October 2024, and 3.0.1 in June 2026. Two dates attach to that last one. The repository release went public on 4 June 2026, while the document status the specification itself declares is a SEMIC Recommendation dated 27 October 2025. The point plotted on the timeline in the previous section is the latter date. The same specification also records that the adoption of W3C DCAT 3 in 2023 triggered a fresh alignment effort for DCAT-AP. Which edition is current, and which properties of the previous edition were deprecated, is written into the document itself.
Extension works like this: domain extensions ride on top of the base profile — GeoDCAT-AP for geospatial information, a separate release for high-value datasets. Korea has the same shape. A GeoDCAT application profile for geospatial data portals was published as a standard on 6 December 2023. The design pattern of per-domain extension is already inside Korea's standards system.
One international standard appears in neither policy document. The ISO/IEC 5259 series, which addresses data quality for AI and machine-learning systems, returns zero hits in a full-text search of both. Lay that alongside the earlier finding that the nine ISO mentions were standards fixing the notation of dates, codes and units, and the shape is clear. On how to write a value, they invoke international standards; on how to measure quality, they built a new indicator system of their own. The international standard on the data-quality axis and the policy on Korea's catalog axis have yet to meet inside the same document. Pebblous has published a case study translating that standard document itself into an OWL ontology, ontologizing ISO/IEC 5259.
5.2Every not-ready survey is a global one
Several surveys describe the situation on the corporate side. The ones below are all global surveys, not Korea-specific data. Reporting them as percentages of Korean companies would be inventing a number that does not exist.
| Survey | Sample and timing | Finding |
|---|---|---|
| Gartner (published 2025-02-26) | 1,203 data management leaders, surveyed July 2024 | 63% said they either lack the right data management practices for AI or are unsure whether they have them |
| Cloudera × Harvard Business Review Analytic Services (2026-03-05) | 230+ respondents, surveyed October 2025 | 7% said their organization's data is fully ready for AI adoption |
| Dun & Bradstreet AI momentum survey (reported 2026-05-20) | Sample not disclosed | 97% are running AI pilots, while 5% say their data is ready |
| Informatica CDO Insights 2025 | 600 global data leaders | Data is the number-one obstacle (43%) to moving generative AI from pilot to production |
Second-hand Korean articles reporting the Gartner sample as 248 are in circulation; the original says 1,203. This report follows the original. Where access to the original is blocked, an archive snapshot can be used to check it.
A few forecast lines from the same firm also circulate, and these are not survey findings but analysts' strategic planning assumptions: that through 2026 organizations will abandon 60% of AI projects unsupported by AI-ready data; that by 2027 organizations prioritizing semantics in AI-ready data will improve generative AI model accuracy by up to 80% while cutting costs by up to 60%; and that by 2030 a universal semantic layer will be treated as core infrastructure alongside data platforms and cybersecurity. It is more accurate to carry them as forecasts than to rewrite them as research findings.
Market size resists a single figure altogether. For metadata management alone, 2025 estimates range from $1.5 billion to $13.28 billion depending on the research firm, because the definitions differ over whether enterprise-wide metadata management is included or the scope narrows to the tools market, and over whether data catalogs are counted in. A spread of nearly ninefold is better read as a signal that the boundaries of this category have not been agreed on yet.
5.3There is already a market in the cell the standards left empty
Korea's consulting market is already selling the layer the standards left blank. As of September 2026, the data consultancy EnCore lists AI Ready Data and Entology as separate items on the products menu of its official site, and Ontology and AI Ready Data as separate items on the consulting menu. The same menus carry entries for a data context map and for MCP extensions. What we checked, though, was only the menu structure. The body copy of each page is rendered in JavaScript and could not be captured, and claims about methodology or about results are a vendor's claims, which we do not carry across as neutral fact.
Training-data infrastructure, meanwhile, is growing quickly regardless of the standards gap. As of 1 September 2026, AI Hub lists 977 open datasets and 51 datasets provided by institutions. On 27 August 2026, 29 outputs from the first-stage evaluation of five teams in Korea's sovereign AI foundation model project — roughly 35.44 million items — were released free of charge. Converted to tokens that is on the order of 1.56 trillion, and the MSIT press release said it is theoretically enough to train a large AI model in the 70-to-80-billion-parameter class. The data is piling up; the vocabulary in which to describe it is the problem this report has been about.
How commercial semantic-layer products handle this layer is outside the scope of this report. The difference between a Palantir-style ontology and a classical one is covered in a separate comparison.
Why this matters to Pebblous
Pebblous diagnoses the state of data and issues the result as a report. That makes the layer map drawn here difficult to read as somebody else's story. The arrangement in which layers 1 and 2 are densely specified by international standards and everything from layer 3 up is empty shows plainly which cell each of our products sits in.
6.1The layers of the map overlap the layers of the products
DataClinic measures, at the first layer, what state a given dataset is in; PebbloScope observes structure; CURK works on the semantic layer. Placed on the ladder from section 4, the three products land on different rungs. And the order the ladder implies is the order in which we ought to recommend them. There are situations where a semantic-layer tool can be sold first, and situations where the upper layer does not hold up until the two below it are in order. This report found evidence on both sides — policy documents and academic experiments — that the second situation is far more common.
6.2The guidelines drew a baseline under our work
The AI-readable state the MOIS guidelines demand comes down to structure, missing values and lineage. Under 20% missing values, six quality indicators from completeness to uniqueness, 38 checklist items — all of them measure the current state of a dataset. The document goes a step further, giving each indicator its diagnostic criteria, diagnostic method and examples of errors, and putting a quality diagnostic report on the list of documents an institution must hold. What we have been doing and calling a report card is now written down as a requirement in a policy document. How the state of training data carries through into a model's internal representations is a subject Pebblous keeps returning to, and this document has drawn a concrete domestic baseline under it.
The experiment in section 4 points the same way. Without changing the model, tidying only the context on the data side changed the answers — and, in the same paper's reading, the choice of model made no statistically detectable difference in either condition. That is one more piece of evidence that diagnosing the data can be cheaper and surer than swapping the model.
6.3The question of what to do first now has an answer
Public institutions and their affiliates are at the stage of having received the guidelines and asking what to do first. This report's practical answer is an order: sort out the catalog first, and the ontology after. That order is supported both by the appendix of the national standard guide and by the 2026 controlled experiment.
Whatever the order, the first step is measuring where things stand. Measuring with a diagnostic tool instead of filling in 38 checklist items by hand is where Pebblous comes in. It is the same on the corporate side: if you have been told to make your internal data usable by agents, build the catalog and the data dictionary before commissioning an ontology design.
6.4Whoever documents the empty cell first sets the standard
That the semantic layer is empty in Korea's standards landscape also means whoever first writes down a methodology for filling that cell ends up proposing the benchmark. Pebblous has already published a case study translating ISO/IEC 5259-2 into an OWL ontology. The grounds for standing between an international standard and domestic policy as an interpreter are in that work.
This report mentions Pebblous products because the findings raised questions we have to answer, not because a gap in Korea's standards proves our products are necessary. Please read the judgement about the standards landscape and our own positioning as two separate things. All queries in this report were run on 1 September 2026, and because the big-data standardization committee's written session for the second half of the year was under way that day, the search results in section 1 may change afterwards.
References
Every standard listing, word count and policy-document quotation in this report was checked by Pebblous against the primary source on 1 September 2026. Where a passage from a Korean-language document is quoted, the English is our translation and the Korean original governs. Survey findings and analyst forecasts are attributed to the body that published them, and where access to an original was blocked, that fact is stated on the line concerned.
Korean standards and terminology
- 1.TTA standardization committees. Standardization project search and the PG1004 published-standards list. committee.tta.or.kr — queried 1 September 2026. All five search terms (AI Ready, AI 레디, 인공지능 준비, knowledge graph, semantic layer) returned no results. The same window returned 37 results for ontology, 14 for semantic and 6 for data catalog.
- 2.TTA. Data Catalog Vocabulary (DCAT) — Version 3, TTAE.OT-10.0427/R2, 2025-12-05, 200 pp. The original TTAE.OT-10.0427 dates to 2017-12-13 and version 2 to 2020-12-10. Standard summary
- 3.TTA. DCAT Application Profile for Data Portals of the Republic of Korea (DCAT-AP-KR), TTAK.OT-10.1406, 2022-12-07, 41 pp. Project number 2021-2197, PG1004. Standard summary · public specification at vocab.datahub.kr
- 4.TTA. Ontology for Data Provenance / PROV-O: The PROV Ontology, TTAE.OT-10.0452, 2020-12-10, 80 pp — a domestic adoption of the W3C's PROV-O.
- 5.TTA. RDF Dataset Canonicalization, TTAE.OT-10.0473, 2025-12-05, 70 pp · Data Fabric — Part 1, TTAK.KO-10.1631-Part1, 2025-12-05 · DataOps — Part 1, TTAK.KO-10.1492-Part1/R1, 2025-12-05.
- 6.TTA ICT terminology dictionary. 인공지능 준비 데이터 / 人工知能準備- / AI-ready data. terms.tta.or.kr — queried 1 September 2026. The same dictionary returns zero hits for the transliterated spelling AI 레디 데이터.
Policy documents and statistics
- 7.National Information Society Agency (NIA), national data infrastructure team, AI data division. Data Catalog Standard Guide for National Data Integration and Linkage v1.0, October 2025, 60 pp per the colophon. Posting — an output of an MSIT ICT and broadcasting R&D project funded by the Broadcasting and Communications Development Fund. Full-text word counts were taken from a text conversion of this PDF.
- 8.Ministry of the Interior and Safety (MOIS) and NIA. Guidelines for AI-Friendly Management of Public Data V1.0, issued 2026-03-31, released 2026-04-24, 70 pp. Public data portal library — word counts were taken from the Markdown original in the same posting (92,690 characters).
- 9.MOIS. 2025 Assessment Results for Public Data Provision and Operation, 2026-03-31. Press release — 684 institutions, 10 indicators across 3 areas, conducted under Article 9 of the Public Data Act.
- 10.OECD. Digital Government Index and Open, Useful and Re-usable Data Index: 2025 Results and Key Findings. OECD Working Papers on Public Governance No. 90, approved 2026-02-12.
- 11.SEMIC / Interoperable Europe. DCAT-AP 3.0.1. The document status declared in the specification is SEMIC Recommendation 2025-10-27; the repository release went public on 2026-06-04. Specification · Release history (retrieved through the GitHub releases API)
- 12.MSIT. Press release on the open release of first-stage evaluation outputs from the sovereign AI foundation model project, 2026-08-27. Policy briefing · AI Hub, aihub.or.kr (queried 2026-09-01: 977 open datasets, 51 provided by institutions)
Academic papers
- 13.Michael Rumiantsau, Ivan Fokeev. "Semantic Layers for Reliable LLM-Powered Data Analytics: A Paired Benchmark of Accuracy and Hallucination Across Three Frontier Models." arXiv:2604.25149, 2026-04-28. arXiv — the +17–23pp figure, the 67.7–68.7% / 45.5–50.5% ranges, the error typology, the residual error rate and the context-form / runtime-form distinction were all checked against the full paper (abstract and sections 4, 5 and 6). The authors' affiliation is Cube, a company that builds a commercial semantic-layer product. Code, questions and judging records are in a public repository. Judging covered 99 of the 100 questions, excluding one whose reference query did not execute.
- 14."When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation" (GraphRAG-Bench). arXiv:2506.05690, ICLR 2026. arXiv — an earlier pass could not open the numeric tables, but the nine observations in section 4, Tables 2 and 3, and the token-cost figure were subsequently checked against the full v3 HTML. The recall figures of 83.2% and 87.9–90.9% quoted here are from Table 3, the novel corpus. Token costs are not reworked into a multiple, because the table in the figure and the body text disagree by an order of magnitude.
- 15.Vladislav Shkapenyuk et al. (AT&T CDO). "Automatic Metadata Extraction for Text-to-SQL" (AskData). arXiv:2505.19988, 2025 — 67.41% on BIRD in the no-oracle condition against 57.13% for the next-best submission under the same condition (a gap of 10.28pp, as of January 2025). Leaders in the hint-using condition: 77.14 (AT&T), 76.02 (Google), 75.63 (Contextual AI, Alibaba), 73.17 (IBM). The 25% undocumented join paths and the device-identifier digit-length case are also from this paper.
- 16.Kaveen Hiniduma, Suren Byna, Jean Luca Bez. "Data Readiness for AI: A 360-Degree Survey." ACM Computing Surveys 57(9):219, 2025. arXiv:2404.05779.
- 17.Jane Greenberg et al. "The metadata ecosystem and AI: Enabling FAIR and AI-ready data." AI Magazine, 2026. doi:10.1002/aaai.70060 · S. Majithia et al. "An actionable framework for AI-ready data." AI Magazine, 2026. doi:10.1002/aaai.70054
- 18.MLCommons Croissant Working Group. "Croissant: A Metadata Format for ML-Ready Datasets." arXiv:2403.19546, NeurIPS 2024 Datasets and Benchmarks Track · Timnit Gebru et al. "Datasheets for Datasets." CACM, 2021.
- 19.Evgeny Kharlamov et al. "Ontology Based Data Access in Statoil." Journal of Web Semantics, 2017 — seven data sources, with one repository alone holding roughly 3,000 tables and 37,000 columns. The weeks-to-minutes description is a case account, not a quantitative benchmark.
- 20.Haram Park and Haklae Kim (Department of Library and Information Science, Chung-Ang University). "DCAT-AP-KR: Application Profile for Interoperability of Data Portals in Korea." Journal of Digital Contents Society 23(11), November 2022, pp. 2249–2258. doi:10.9728/dcs.2022.23.11.2249 — public institutions were running 109 data portals as of 1 February 2022; the paper drops 41 subject-specific portals and 23 that publish no dataset metadata, and surveys the remaining 45 (6 central government bodies, 13 local governments, 26 public institutions) to select a set of common metadata items and define new vocabulary on DCAT-AP design principles. The TTA standard of the same name, reference 3 above, was announced on 7 December that year. What this report says about DCAT-AP-KR rests on that standard's summary and public specification; the paper's own analysis is not cited in the body.
- 21.Juan Sequeda, Dean Allemang, Bryon Jacob. "A Benchmark to Understand the Role of Knowledge Graphs on Large Language Model's Accuracy for Question Answering on Enterprise SQL Databases." arXiv:2311.07509 — 16% to 54%. · Dean Allemang, Juan Sequeda. "Increasing the LLM Accuracy for Question Answering: Ontologies to the Rescue!" arXiv:2405.11706, ISWC 2024 — 72% with ontology-based query verification and repair, a further 8% returned as "I don't know," and an overall error rate of 20%. Both figures were checked directly in the papers' abstracts. Reference 13 carries them as 16.7% and 54.2%, but the original abstracts say 16% and 54%, so we follow the originals. The BIRD evidence gap quoted in the body (34.88%→54.89%, human experts 72.37%→92.96%) is carried over from reference 13's citation of it and was not re-verified against the BIRD paper's own PDF (arXiv:2305.03111).
Research-firm publications (including forecasts)
- 22.Gartner. "Lack of AI-Ready Data Puts AI Projects at Risk" (Roxane Edjlali), 2025-02-26 — 1,203 data management leaders, surveyed July 2024. The 60%-abandonment forecast is from the same release.
- 23.Gartner. "Top Data and Analytics Predictions" (Carlie Idoine, Sydney Summit), 2025-06-17 — the 2027 80%/60% forecast is Carlie Idoine's; secondary coverage has attributed it to other analysts. The 2030 universal-semantic-layer forecast from the 2026-03-11 release was also consulted.
- 24.Cloudera × Harvard Business Review Analytic Services, 2026-03-05 — 230+ respondents, surveyed October 2025 · Dun & Bradstreet AI momentum survey, reported 2026-05-20 (sample not disclosed) · Informatica & Wakefield Research, CDO Insights 2025 (600 respondents).
- 25.EnCore official website, en-core.com, queried 2026-09-01 — only the structure of the product and consulting menus was verified. The body copy of each page is rendered in JavaScript and could not be captured.