Executive Summary

Starting in 2026, Korea's Ministry of Trade and Industry will pull the senses and experience of master welders and assemblers out of the shop floor and into files, and from 2027 it will build a manufacturing data library to hold them. The security design is already tight. Data can be used only inside a clean room cut off from outside networks, taking anything out is prohibited, and even viewing requires a separate review. The vault has been designed. What is empty is the rule for what goes into it.

An intake rule settles two things: what gets written down as the correct answer, and under what conditions the data was captured. Attach sensors and you get current waveforms, molten-pool video, arc sound and torch angles, but decide none of it into a definition of a good weld and what remains is not a craftsman's touch but a log. In the research literature the ground truth for the same molten-pool video splits four ways, and because the unit on the report card changes with it, the results cannot be compared. Nor do the pilot figures the ministry has published say which plant or which conditions produced them.

An intake rule is not the opposite of openness but the condition for it. Korean courts have upheld confidential-management status even after material was provided to a third party, where circumstances supported a duty of confidentiality; depositing data in a clean room does not by itself forfeit trade-secret status. Companies hesitate less over their legal position than over not knowing how much of what leaves the building. Standardize the label definitions and the condition metadata instead of demanding whole raw logs, and a company can keep its parameter recipes in house while contributing only the signal a model can learn from. Korea left those same two blanks in 2020, in what it called the world's first manufacturing data sharing norms. So the question this report follows is not whether the bill passes but whether those two items make it into the implementing rules.

30

shop floors where tacit skill is being turned into data

Industry ministry briefing, second half of 2026

20.5%

of manufacturing workers are 55 or older

Roughly triple the 7.1% of 2010, in 14 years

0.1%

of small and mid-sized manufacturers have adopted manufacturing AI

Smart manufacturing innovation survey

49

standard datasets KAMP gathered in six years

Starting from 12 in December 2020

1

The Government Designed the Vault First

M.AX, the flagship agenda Korea's industry ministry set out for the second half of 2026, stands for the AI transformation of manufacturing (Etoday, 2026-08-10). The first item on the list is moving what has lived in a worker's body into files. From 2026 the ministry will extract the senses and experience of master welders and assemblers as digital data; from 2027 it will build a manufacturing data library to store and use that data safely. Two more milestones follow the library. Between 2027 and 2031 it wants full-stack AI factories in which AI robots judge and run processes on their own, and by 2030 it will designate seven regional industrial complexes as M.AX clusters, expanding private 5G networks and edge AI data centers.

When Plan
From 2026 Extract the senses and experience of master welders, assemblers and others as digital data
From 2027 Build a manufacturing data library to store and use the extracted data
2027–2031 Full-stack AI factories in which AI robots judge and operate processes on their own
By 2030 Seven regional industrial complexes designated as M.AX clusters, with private 5G networks and edge AI data centers

Source: Korea's industry ministry, key policy directions for the second half of 2026 (as reported by Etoday, 2026-08-10). The seven clusters are not new construction but selective upgrades among ten complexes already designated in 2025, and the first pilot site is the Banwol-Sihwa complex.

The first thing that stands out in the plan is how finished the security design is. Data in the library may only be used inside a clean room isolated from external networks, taking data out is prohibited, and even viewing runs through a separate review procedure (Ajunews, 2026-06-05). Data is also piling up before the building exists. Because the library will take time to stand up, since May 2026 a manufacturing-AI solution development center run by the Korea Electronics Technology Institute has served as an interim site, storing data from the AI factory program, with a prototype manufacturing-AI foundation model targeted by year end.

The tacit-knowledge program is not just paper, either. In its August 4 briefing the ministry said it would turn the tacit manufacturing knowledge of skilled workers at 30 shop floors into data and develop AI models infused with the experience of master craftspeople. What the policy documents call a manufacturing master is the person this report calls a craftsman. Which industries and which processes those 30 sites cover, and whether they share one label schema, does not appear in any public material. The pattern echoes our earlier report on robot behavior data, where data also flowed into a government-built facility ahead of the rules. The difference this time is ownership: the legal owner of this data is a private company, not the state.

The hurry comes from how fast the people are leaving

According to the Korea Employment Information Service, 20.5% of manufacturing workers are now 55 or older, roughly triple the 7.1% of 2010, over 14 years. Workers in their fifties are the largest band at 24.7%, and those 60 and over account for 13.2%. Shipbuilding employment stands at 93,038, less than half the 2014 peak of 203,400, and the Korea Offshore & Shipbuilding Association projects an annual shortfall of more than 12,000 workers, reaching about 130,000 by 2027.

Read those numbers as a generic labor shortage, though, and the point blurs. E-7-3 visas issued for shipyard trades including welding rose from 264 in 2021 to 13,297 in 2025, and foreign nationals now make up about a quarter of the shipbuilding workforce. By the ministry's own count, 86% of production workers newly hired in shipbuilding during the first three quarters of last year were foreign. The headcount gap is largely being filled. What is not being filled is the layer of judgment that walks out the door with a worker of thirty years' standing, and that judgment is what the data program aims at.

What the bill governs, and what it does not

The enabling statute already exists. The current Act on Promoting Industrial Digital Transformation and the Use of Artificial Intelligence is itself a renaming of the 2022 Industrial Digital Transformation Promotion Act, so a law covering industrial data and AI use together is not new ground. On top of that, lawmaker Chang Chul-min, a ranking member of the National Assembly's industry committee, introduced a full-scale amendment on July 27, 2026 that would rename the statute the Industrial AI Transformation Promotion Act. That is the bill the government and the press call the M.AX law.

The text of the amendment is so far visible only through secondary reporting. According to those reports, Article 2 adds a definition of an industrial data library, and Article 35 requires industrial data collected and held in the library to be sorted into tiers by its impact on national security and the national economy, with security measures established and enforced per tier. Articles 26 through 28 cover performance certification for industrial AI products, revocation of that certification, and procedures for correcting manufacturing defects. Reports also describe provisions for forming and running a manufacturing AX alliance, for using tacit-knowledge data from the shop floor, for collecting and preserving unstructured data such as the working behavior and know-how of skilled technicians, and for exemptions from preliminary feasibility study requirements.

The schedule pressure is unmistakable. On July 8, 2026 a party-government consultation set the M.AX bill on course for passage during the regular session, and the budget ministry has reportedly been weighing a national manufacturing data library line item in the 2027 budget. If the bill slips, the argument goes, the library construction and AI factory conversion due to ramp up next year slip with it.

Within the provisions that have been reported, two of them touch data directly: security after it arrives (Article 35) and product performance after it leaves (Articles 26 to 28). Between those two, a provision covering what arrives and in what format cannot be found. We could not obtain the bill's original text, so we cannot assert that no such provision exists. But line up what is verifiable and the location of the gap becomes clear. That gap is what this report is about.

What the bill covers, and what it leaves out Shop-floor data Welder/assembler senses Intake rule Ground-truth def.? Condition metadata? Clean-room security Article 35 · tiered security Output certification Art. 26-28 · performance cert. Reported provisions cover only the two ends — security in, performance out What enters, and in what format, is not established Pebblous original diagram — reconstructed from reported provisions (Arts. 2, 26-28, 35) of the M.AX bill
▲ The intake-rule gap (Pebblous original diagram) — no provision covers ground-truth definitions or condition metadata between security and performance certification
2

What Gets Written Down as Correct

The hard part of moving welding skill into data is not the sensors. What to measure has largely been settled. The channels that recur across the welding-monitoring literature converge on arc current and voltage time series, molten-pool vision, thermal imaging, arc sound, and three-dimensional bead profiles, and the layer closest to a craftsman's hands is the set of torch motion variables. Devices already exist in the patent record that mount inertial sensors on an expert's torch to track work and travel angles, and derive travel speed and arc length from helmet-mounted camera footage.

Channel What is measured
Electrical signals Arc current and voltage time series, time-frequency features
Molten-pool vision Surface geometry, reflection patterns, structured-light laser reflections
Thermal imaging Temperature distribution across the molten pool
Acoustics Irregularity in the arc sound
3D profile Bead geometry from a laser line scanner
Motion and posture Work angle, travel angle, contact-tip-to-work distance, travel speed, arc length, weaving amplitude, gaze

Source: synthesis of welding quality monitoring and skill measurement literature (2023–2026)

Same video, four different ground truths

The trouble starts one step later. On a single task, predicting penetration depth, the input is the same molten-pool video while the ground truth is sourced in at least four different ways. One group photographs the bead width on the back side of the workpiece and records three categories: insufficient, adequate, excessive. Another sets a brightness threshold on the back-side image and cuts it into categories that way. A third sections the finished coupon and records the measured penetration depth as a continuous value in millimeters. On the laser-welding side, another team calibrates against keyhole depth measured physically by optical interferometry.

How the ground truth is sourced Form of the label Reported performance
Photographing back-bead width 3-class classification 92.70% accuracy
Brightness threshold on back-side image 4-class classification 98.11% accuracy
Offline measurement after sectioning Continuous regression (mm) MAE 0.0516 mm, R² 0.998
Optical interferometry calibration Physically measured keyhole depth Multi-sensor transfer learning (laser welding, 780DP steel)

Source: papers on penetration prediction. These are different label definitions for the same task, so the three figures are not comparable with one another.

Put 92.70%, 98.11% and R² 0.998 side by side and you cannot say which model is better. The first is the rate of correctly calling penetration adequate, the second the rate of separating four classes, the third an error in millimeters. Change the label definition and the same sensor data yields a different model, measured on a different scale. For a library to let data from many factories be pooled and used, that scale has to be fixed before the pooling starts.

Same video, four ground truths Molten-pool video Same input data (penetration task) Back-bead width photo → 3-class 92.70% accuracy Back-side brightness → 4-class 98.11% accuracy Sectioned measurement → regression (mm) MAE 0.0516mm · R² 0.998 Optical interferometry → keyhole depth Laser welding, multi-sensor transfer Pebblous original diagram — sourcing methods for penetration-depth ground truth (Section 2 table)
▲ Same video, four ground truths (Pebblous original diagram) — different sourcing methods put results on different scales that cannot be compared

Ground truth also carries its own cost. The molten-pool surface is measured in real time by a 3D sensor, but the label itself, the back-bead width, has to be measured offline after the experiment, once the workpiece has been cleaned and machined. The researchers say so themselves. A program that collects sensor logs and a program that collects labeled data have different cost structures. When a budget line lists storage capacity and network bandwidth but no annotation work, the warehouse fills with logs.

Whether a public benchmark dataset exists in this field is also unclear. Within the scope of our search we found no public molten-pool welding dataset; each research group runs its own experiments and writes its own labels. That contrasts with additive manufacturing, which has a tradition of shared datasets. Welding AI has not yet reached the stage of training on someone else's data and comparing results. The reason is less that there is too little data than that nobody has agreed on what to record as correct and which conditions to record alongside it.

What can be measured is not all of the skill

WeldAR, presented at ACM CHI this year, aligns digital twins of the torch and the workbench to compute four parameters in real time — contact-tip-to-work distance, work angle, travel angle, travel speed — and delivers them as augmented-reality guidance. The observations are the interesting part. Learners came to experience the continuous feedback as an interruption over time, and in trials with the guidance display removed they fell back on sound, vibration, and the steadiness of the hand holding the torch. Nobody told them to do that.

If that is so, the procedure for writing down ground truth has to include a round trip through the person. Studies that combined interviews, motion capture and video analysis across experts and novices report that tacit knowledge becomes speakable once the analysis is played back to the practitioner. The knowledge-acquisition literature takes a similar stance: rather than asking an expert to explain what they know, elicit the behavior, measure it, and show the measurement back. A label definition is the product of that round trip, and it comes before the sensors go on.

Who writes the labels is itself a data-quality variable. On a task marking tool-wear regions, expert annotators reached mIoU 0.8153 while novices scored significantly lower. A welding defect-detection study documented how an annotator's habits produced systematic bias: left-hand boundaries were consistently marked less accurately than right-hand ones, and that bias accumulated into degraded performance. Systematic error is more dangerous than random error, because a neural network is unusually good at memorizing a regular mistake. Hence the recursion in any program that moves a craftsman's judgment into data: whoever writes the ground truth has to be skilled too.

The word library already appears in the research literature, with a different referent. There, torch travel speed, arc length, welding angle, current and wire feed rate are extracted from a skilled worker's demonstration to form a skill library, and a new task is decomposed into those elements and matched against it. The unit of that library is a parameterized skill element, not a raw-log archive. Policy and research are using the same word for two different things.

3

Numbers Without Conditions Don't Reproduce

Why are labels and conditions an intake-stage problem? The performance figures the government has already published hold the answer. The original text reads roughly as follows: an AI factory pilot applied to 170 worksites nationwide in the first half of the year demonstrated visible results including a 30% gain in productivity, a 15% drop in defect rates and a 20% cut in costs, with a 90% reduction in semiconductor quality inspection time and battery quality prediction accuracy raised to 87% as headline cases. Those sentences have to be cited as the ministry's claim. Which worksite, which process and which conditions produced each number is not in the source.

How the sample was drawn is disclosed, inadvertently, by the wording of another report: the program supported AI transformation at some 170 worksites, and an inspection of the 42 among them where visible results had begun to appear found productivity up an average of 30.1% and defect rates down an average of 15.5%. The sentence states that only the sites where results had started to show were inspected. What happened at the sites left out of that inspection is not established anywhere in the same material.

The unit of counting and the reference date shift between documents, too. The original article said 100 manufacturing processes would be converted into AI factories by 2026; reporting from the August 4 briefing said 50 more would be added in the second half for a cumulative 220 processes by year end; and the population for the performance check was some 170 worksites. Public materials give no way to tell whether processes or worksites is the larger set, or whether the three numbers count the same things. The defect-rate figure splits as well: 16% in the government's own results summary, 15.5% in press coverage of the same announcement, 15% in the original article.

Attribution wobbles too. In the case list released on August 4, the plant that cut inspection time by 90% is not a semiconductor fab but Asan Sungwoo Hitech, a maker of EV battery components, while the semiconductor cases on that list show improvements of 33% and 66.7%. No case matching 87% battery quality prediction accuracy appears in any public list. There is no basis for calling any of this a misreport. What is observable is that the same headline result attaches to different companies and different industries depending on the citation path, and that is the ordinary fate of a number that travels without its condition metadata.

Why it collapses in the next factory

The distribution of equipment data shifts easily with operating conditions, component quality and environmental noise. Distribution mismatch between training and test data is therefore the default rather than the exception, and the literature repeatedly reports it showing up as degraded performance in real deployment scenarios. Deep learning performance rests on the assumption that source and target share a distribution, and that distribution is sensitive even to installation environment, power supply and load variation. This is why another factory's 30% cannot be copied into your own factory's expected value.

Tools for gauging in advance whether a model will fit your plant do exist. Metrics that predict transferability up front have been proposed and reported to track actual diagnostic-accuracy trends. Two caveats from that same literature land directly on this report's point. One is that generic distribution distances do not capture the whole difference. The other is that label-space shift also has to be measured, and label spaces in manufacturing fault diagnosis are likely to be mismatched. Both calculations take the source data's conditions and label definitions as inputs. A dataset with no conditions recorded does not even supply the input needed to judge whether it can be used. That releasing data and making it usable are two different things is something we already traced through public research data, and the structure of provenance records and the ISO/IEC 5259 family is covered in a separate report.

The research community pushed this problem into institutional form. NeurIPS requires machine-readable metadata in Croissant format for dataset-track submissions. That toolchain carries an assumption the library cannot inherit: metadata generation typically begins by uploading data to a public platform, an approach that related work notes is not executable inside a controlled repository. A library that forbids export breaks the premise of off-the-shelf metadata practice head-on. There is, in other words, a technical reason it has to write its own intake specification, and no evidence yet that it has.

4

The Law Isn't What Makes Companies Hesitate

The obstacle named most often in policy discussion is ownership of the data. An official at a state-funded research institute said that because process data is a company's top-tier trade secret, firms will not readily hand over core data unless rigorous security and a profit-sharing model are in place first. The same source added that careful incentive design has to come first so that a large manufacturer opening up leads to suppliers actually using the data and lifts the competitiveness of the whole ecosystem. The diagnosis is accurate. The obstacle, though, sits in a slightly different place.

The law does not block the intake

Korea's Unfair Competition Prevention and Trade Secret Protection Act sets three requirements for a trade secret: that it is not publicly known, that it has economic value, and that it is managed as confidential. The third has been progressively relaxed. A 2015 amendment changed "substantial effort" to "reasonable effort," and a 2019 amendment softened it again to simply "managed as confidential." The courts go further. Precedents including Supreme Court decision 96Da16605 have upheld confidential-management status even where material was provided to a third party, absent any explicit agreement, so long as circumstances support a duty of confidentiality under good faith or by implication. The weaker bargaining position of a subcontractor is treated as one of the factors to weigh.

Apply that doctrine to the library and the conclusion points the other way. Depositing data in a clean room does not by itself forfeit trade-secret status. Restricted access, confidentiality obligations, and a management posture visible to an outside observer are, if anything, exactly what the doctrine asks for. The flaw in this policy is neither lax security nor a legal barrier.

So what makes a company hesitate? Not knowing how much of what goes out. Layered on top of that is personal exposure for the employee who signs. Seoul High Court decision 2023No999, affirmed by the Supreme Court as 2023Do10280, treated a trade-secret leak as the act of moving material outside a location the holder designated or approved, and did not require, as an element of the offense, that it reached a third party or created a risk of overseas leakage. Absent an internal approval procedure and a documented basis for it, the very act of uploading data to the library can touch the elements of a criminal offense. Few things explain a delayed signature better.

Meanwhile the law already on the books asks for a good deal. Article 9 of the current Act on Promoting Industrial Digital Transformation and the Use of Artificial Intelligence requires a reasonable profit-sharing contract among the parties who jointly generated industrial data, and sets integrity and reliability of that data as a principle. It also empowers the industry minister to issue guidelines on data-use contracts. The legal basis for both incentives and quality has been there since 2022. What is missing is not a new statute but the specification that would sit underneath it. That is the point the debate over legislative speed keeps skipping.

Not moving the data does not solve the condition problem

Federated learning comes up often as a way to join training without shipping raw data out. Empirical assessments are cautious. Non-IID conditions, where data distributions differ sharply across participants, are named as the principal barrier to industrial adoption, and manufacturing supplies the concrete example: plants running nominally identical equipment to produce different product variants diverge widely in their statistical patterns. In that case a global model can converge on a solution that fits no individual site well. The literature also flags local models being overwritten by the global one, taking site-specific information with them, along with the communication cost of repeated parameter exchange and infrastructure gaps between facilities.

There is also a shortage of grounds for choosing among the options. As of 2026, comprehensive studies evaluating federated-learning frameworks as complete systems under realistic industrial conditions remain scarce; existing work leans toward algorithms and synthetic benchmarks, leaving practitioners with little basis for picking a framework. One study on anomaly detection in smart manufacturing reported F1 of 96.1% on synthetic data and a public bearing dataset, falling to 92.5% as differential-privacy strength was increased. That figure shows the exchange rate between privacy and accuracy, but it is not a score on production data.

Federated learning solves the problem of not moving data, and makes the problem of combining data captured under different conditions more acute rather than less. Without condition metadata, even federated learning has no way to know what to weight and how. Synthetic data and simulation run into the same wall: define nothing about what is being reproduced and you cannot say what was synthesized.

The incentive report card

The formal machinery for sharing the gains is already in place. Korea's certified data brokers grew from 52 in the first cohort in 2022 to 162 by the fourth cohort in 2023. Yet no statistics establish how much actual settlement has occurred on manufacturing data. Academic tooling, by contrast, is abundant. Since Data Shapley in 2019 an entire benchmark ecosystem has grown up around pricing data contribution, but we found no case of it applied to a real settlement on Korean manufacturing data.

The inability to put a price on data shows up in the court statistics as well. A July 2026 survey by the Korea Enterprises Federation found overseas leakage of core technology detected in 9 cases in 2021 rising to 33 in 2025, and among 496 first-instance guilty verdicts for technology leakage between 2015 and 2023, not one accepted a calculation of damages. Leaks rise; value goes unpriced. In that institutional setting, telling companies they will share in the profits if they contribute data does not amount to a promise unless a method for pricing comes with it.

This is where a label schema becomes the way out. Demanding whole raw logs is the same as demanding a top-tier trade secret. Standardize label definitions and a minimum set of condition metadata, and a company can keep its parameter recipes in house while contributing only the signal a model can learn from. What counts as a contribution is definable only on top of that specification. The moment contribution is counted in labeled samples and diversity of conditions rather than bytes, the profit-sharing clause becomes a sentence someone can actually compute. Openness holds together only where an intake rule exists.

What goes out: the recipe, or just the signal? Whole raw logs Unfiltered sensor output Top-tier secret exposed Criminal exposure risk Risk Recipe leaves whole → personal criminal risk Label schema + condition metadata Learnable signal only Conditions in, recipe out Open Recipe stays home only signal contributed Pebblous original diagram — exposure comparison: whole raw logs vs. standardized label schema
▲ What a label schema opens up (Pebblous original diagram) — an intake rule is not the opposite of openness, it is the condition for it
5

Korea Stood Here Six Years Ago

Read this policy as a first attempt and the judgment goes wrong. The norms written in 2020 left blank the very items that are blank now. That repetition is the most important fact in this report.

Lesson 1. The norms settled the transaction and left the label blank

On October 29, 2020, Korea's Ministry of SMEs and Startups announced at a strategy committee on AI and manufacturing data that it would establish manufacturing data sharing norms, billing them as the world's first rules for managing manufacturing data. What those norms governed was the definition and scope of manufacturing data, transaction requirements, principles for sharing profits, and the rights among producers, providers and users. Label definitions, quality criteria and condition metadata were not in scope.

Two months later, on December 14, 2020, KAMP opened. The minister at the time, Park Young-sun, described building a "my manufacturing data" system in which participants share and trade manufacturing data under mutually agreed rules and split the gains reasonably, opening what she called the era of the protocol economy, with platform operations promised for the first half of 2022. The standard dataset count rose from 12 to 24 by December 2023 and 49 by October 2025. Yet no public material establishes the cumulative number of companies using it, download counts, or actual transaction counts, and we found no material reporting settlements from profit sharing. A 2025 document says outright that 49 datasets do not cover all manufacturing sectors. The only layer where quality inspection appears at all is a consulting line item in the 2024 program for processing manufacturing data into products.

Lesson 2. A case that records its conditions becomes a reproducible account

A counterexample sits inside the same KAMP. Chosun Refractories, which makes refractory products, applied deep-learning image classification to X-ray-based defect judgment and moved judgment reliability from 90% to 96% while cutting inspection time from 1.5 minutes to 0.5. What separates this figure from the earlier ones is not magnitude but description. What was inspected (X-ray images of refractories) and how the judgment was made are both stated in the sentence. Another company can therefore hold it up against its own process. An average of 30.1% pooled across 42 worksites permits no such comparison.

Lesson 3. The countries that finished the rulebook got stuck on participation

The closest thing to a finished intake rule is Catena-X in the German automotive industry. Compliance with per-use-case guidelines is mandatory, participants who do not accept the data exchange governance cannot trade through a registered connector, and an ODRL profile standardizes permissions, prohibitions and duties into machine-readable contract modules. Version 4 of the product carbon footprint rulebook, released in September 2025, shows the method: where an international standard leaves room for interpretation, insert a binding specification. And yet its own 2026 documentation still describes adoption as in progress, with the number of actively participating companies small relative to the scale the network was designed for. Onboarding infrastructure, credential management and ongoing compliance are a burden for smaller firms.

A bigger program in the same country lands in a similar place. Germany put €150 million into the Manufacturing-X R&D consortium, but a Bitkom survey found high awareness paired with low participation, with only 34% of industrial companies saying value-chain data exchange was decisive for competitiveness (survey conducted 2023–2024). Japan's Ouranos Ecosystem published a reference architecture and demonstrated interoperability with European data spaces, yet describes itself as at an early stage. Korea, meanwhile, already runs several no-export repositories of its own. The national statistics office's microdata service blocks copying and releases only approved result tables; the health and medical data safe zone allows analysis in an environment physically separated from the internet with no downloads; the Korea Data Agency's data safety zone works the same way. Among these we could not find one that publishes annual usage counts or the number of approved projects. Set rule maturity beside the method of drawing participation and the four cases share one trait. The side that finished its rules got stuck on participation, and the side that publishes no numbers leaves no way to tell whether it is stuck at all.

Case Maturity of the rules How participation is drawn in
Catena-X (Germany) Mandatory guideline compliance, ODRL-based usage control, binding rulebooks €23M Data Space Accelerator, June 2026. Up to €30K per company, with proof of actual data exchange as a condition of payment
Manufacturing-X (Germany) €150M into an R&D consortium High awareness, low participation. Only 34% of industrial companies called value-chain data exchange decisive for competitiveness (2023–2024 survey)
Ouranos Ecosystem (Japan) Reference architecture published, interoperability with European data spaces demonstrated Describes itself as early-stage and runs a project certification scheme that solicits exemplary cases
Korean no-export analysis environments No copying, only approved result tables released, physical separation from the internet No instance found that publishes annual usage counts or the number of approved projects

Sources: Catena-X governance documents and PCF Rulebook v4; IDSA Data Space Accelerator call (2026-06); Bitkom survey; METI materials on Ouranos project certification; usage guides for Korea's microdata service and the health and medical data safe zone

Germany's response is the instructive part. The Data Space Accelerator pays up to €30,000 per company for onboarding, certificate management and building a second use case, and it wrote proof of actual data exchange into the payment conditions. The design pays for exchange, not registration. The program calls itself research and measures, through surveys before and after onboarding, whether value appears once a critical mass is exchanging real production data. Japan acknowledged the same problem differently. Judging its own initiative to be at an early stage, it split its call for projects in two: leading projects that already offer data-sharing functions as commercial services, and challenge projects that state goals still ahead of them.

There is also testimony that technology is not the bottleneck. Hartmut Rauen of the German engineering federation VDMA said the barrier in these discussions sits in people's heads rather than in the technology, and that data sharing had failed to work because nobody wants to operate on top of a monopoly. The self-assessment from the head of the same association's platform-economy working group is sharper still: they had never gone to customers and asked what they would be willing to pay for, and how much. It is hard to find a shorter summary of a failed incentive design.

The domestic baseline belongs in the picture too. Among Korean small and mid-sized manufacturers, smart factory adoption stands at 19.5%, manufacturing AI adoption at 0.1%, and only 0.8% have dedicated AI staff or an AI organization. About 80% sit at the basic tier of the four-level smart factory scale; 17.3% still collect data by hand in spreadsheets, and 39.1% by barcode or direct terminal entry. Vice Minister Roh Yong-seok of the SME ministry called AI transformation a precondition for the small firms that make up 99.6% of all manufacturers. In other words, the companies that already hold learnable data worth contributing to a library are a small minority, which is exactly why intake rules cannot be written only for the clean-room conversation among large firms.

6

Seven Things to Settle Before You Hand Over Data

A company weighing whether to contribute data to the library has items to settle internally before it decides. Since the government is writing the implementing rules right now, the items to demand of that specification come from the same list. The seven below include only what the preceding sections established, each paired with what actually happens when it is left blank.

# Settle this first The question to answer What happens if you don't
1 Definition of correct
the label schema
What gets written down as a good weld: bead appearance, three classes of back-bead width, penetration depth in mm, tensile strength, or a skill grade? Report cards can't be compared. 92.70% and R² 0.998 are not scores on the same problem
2 Who writes the labels Are the annotators skilled workers? How many per item? What agreement threshold applies, and is agreement measured in a pilot first? Unskilled judgment gets stamped into the data. Systematic bias is more dangerous than random noise
3 Condition metadata
the minimum set
How far do you record equipment model, base-metal spec, plate thickness, welding position, shielding gas, current range, temperature and humidity, operator grade, measuring instrument? There is no input for judging whether another plant can use it. Transferability metrics require label-space shift as well
4 Provenance records Which line at which point in time, and through which preprocessing steps and which versions? When a result fails to reproduce, the cause cannot be traced
5 Who approves the deposit
and on what documented basis
Who inside the company approves it, under which internal rule, and where is that record kept? It becomes personal criminal exposure for the employee. Case law treats a leak as moving material outside an approved location
6 Unit of output control Is what leaves a model, a result table, or parameters? What are the review criteria for inversion and re-identification? A clean room stops raw data. Know-how can still seep out through the outputs
7 Unit of contribution accounting What counts as a contribution: bytes, labeled samples, diversity of conditions, or measured model improvement? The profit-sharing clause stays blank in the contract. The 2020 norms also set profit sharing, and no settlement figures were ever published

The order matters as much as the list. Measure inter-annotator agreement in a pilot and freeze the label schema. Fix the minimum set of condition metadata and build it into the collection pipeline. Establish the approval line for deposits and the document it rests on. Agree with the counterparty on the criteria for controlling outputs. Then contribute the data. Sensors first is the wrong sequence.

The order to settle before intake — schema before sensors 1 Pilot agreement measurement Freeze label schema 2 Condition metadata Fix minimum set 3 Approval line Documented basis 4 Output control Agree criteria 5 Data intake Pebblous original diagram — execution order for the Section 6 checklist
▲ The order to settle before intake (Pebblous original diagram) — sensors first is the wrong sequence

Nor is there a shortage of material to build with. ISO/IEC 5259-3 carries the requirements for a data quality management system, and 5259-4 covers processes for data labeling, evaluation and lifecycle management. Both were published in 2024. Croissant supplies a machine-readable metadata layer but assumes public upload, so it cannot be lifted into a controlled repository as is. Catena-X's rulebooks demonstrated the method of laying a narrowed specification over the places where an international standard leaves interpretation open. The document the library needs is not a different shape from that.

What to demand on the policy side comes down to two lines: put an intake specification in the implementing rules, and put publication of usage figures in alongside it. If the library inherits the habit of Korea's existing no-export analysis environments and never publishes usage counts, no outside observer will be able to verify whether it succeeded or failed. That, in the end, was the largest gap the 2020 norms left behind as well.

Why Pebblous Cares

Moving tacit skill into data is a labeling-design problem, not a storage problem. A manufacturing log missing its label schema, condition metadata and provenance records is a single case in which three of the defect types data quality diagnosis deals with — inconsistent label definitions, missing metadata, unrecorded conditions — all appear at once. The items Pebblous diagnoses with DataClinic are being repeated verbatim inside a policy document.

What you wrote as correct decides what the model learns

A model trained against bead appearance and a model trained against tensile strength learn different things from the same sensor data. That the definition of the training data determines a model's internal representation is a claim this policy illustrates with public evidence. If ground truth already splits four ways on penetration prediction alone, the goal of pooling data from many factories into one sector-wide model will not be reached without a specification document.

What a company facing the decision actually needs

Manufacturers will soon have to decide whether to contribute data to the library. A security promise is not all that decision requires. It also takes a standard for what to contribute, and in what form, so that the know-how stays in house and only learnable signal leaves the building. The table in section 6 can serve as a first draft of that standard. Because the library's implementing rules are being written now, whatever a company settles internally becomes material it can put directly into the specification debate.

Editor's Note. Pebblous works on data quality diagnosis and provenance design, so we have an interest in this subject. This report is not a pitch for a particular product. It is an attempt to record, at the moment the intake specification is actually being drafted, which omissions would produce the same outcome as six years ago. Which boxes have to be filled before the library opens is a judgment we leave to the companies contributing data and the people writing the rules.

R

References

Policy, law and statistics

  • 1.Seo, B. (2026). Government moves fast on manufacturing AI; the crux is data openness and legislative speed. Etoday, 2026-08-10 (in Korean). Seed report.
  • 2.Act on Promoting Industrial Digital Transformation and the Use of Artificial Intelligence (Act No. 21250), Korea Law Information Center. Article 9: principles for using and protecting industrial data, profit-sharing contracts, integrity and reliability, and the minister's authority to issue guidelines.
  • 3.Unfair Competition Prevention and Trade Secret Protection Act, Article 2, Korea Law Information Center. The three requirements for a trade secret, including management as confidential.
  • 4.Supreme Court of Korea, decision 96Da16605, 1996-12-23. Confidential-management status upheld after provision to a third party where circumstances support a duty of confidentiality under good faith or by implication.
  • 5.Supreme Court of Korea, decision 2023Do10280, 2023-09-27 (lower court: Seoul High Court 2023No999). A leak is the act of moving material outside a designated or approved location.
  • 6.Ministry of SMEs and Startups (2020). World's first manufacturing data sharing norms to be established. Press release, 2020-10-29, Korea Policy Briefing (in Korean).
  • 7.National AI Strategy Committee (2026). Korea's AI Action Plan (AI Basic Plan 2026–2028). 2026-02-25, aikorea.go.kr.
  • 8.Ajunews (2026). Reporting on clean-room operation and the export ban for the manufacturing data library, and KETI's interim site. 2026-06-05 (in Korean). Secondary coverage.
  • 9.Coverage of the industry ministry's briefing for the second half of 2026 (ZDNet Korea and Ajunews, 2026-08-04). Tacit-knowledge data at 30 sites, 220 cumulative processes, results from the 42-worksite inspection, and the published case list. Secondary coverage.
  • 10.Korea Economic Daily (2026). AI transformation: Chang Chul-min introduces the M.AX bill. 2026-07-27; and Goodmorning Economy, "This bill: the Industrial AI Transformation Promotion Act." Basis for the structure of Articles 2, 26–28 and 35 (secondary reporting).
  • 11.Korea Employment Information Service. Analysis of labor market characteristics among older workers in manufacturing. 20.5% aged 55+, 24.7% in their fifties, 13.2% aged 60+.
  • 12.Korea Enterprises Federation (2026). Survey on technology leakage. 2026-07. Detected overseas leakage rose from 9 cases in 2021 to 33 in 2025; among 496 first-instance guilty verdicts, zero accepted a damages calculation.
  • 13.Ministry of SMEs and Startups and the Smart Manufacturing Innovation Agency. KAMP standard dataset status and the smart manufacturing innovation survey. 12 datasets (2020-12) to 49 (2025-10); smart factory adoption 19.5%, manufacturing AI adoption 0.1%.

Research and standards

  • 14.ISO/IEC 5259-3:2024 and ISO/IEC 5259-4:2024. Artificial intelligence — Data quality for analytics and machine learning (ML). ISO/IEC JTC 1/SC 42. Quality management requirements; framework for labeling, evaluation and lifecycle processes.
  • 15.Akhtar, M. et al. (2024). Croissant: A Metadata Format for ML-Ready Datasets. DEEM @ SIGMOD 2024, ACM.
  • 16.Croissant Baker (2026). arXiv:2605.15079. Metadata generation presumes upload to a public platform, making it unexecutable in a controlled repository.
  • 17.WeldAR: Augmenting Live Hands-On Training with In-Situ Guidance for Novice Learners. ACM CHI 2026, arXiv:2603.07959. AR guidance on CTWD, work angle, travel angle and travel speed, and learners shifting to reliance on sound and vibration.
  • 18.Papers on penetration prediction: deep-learning penetration prediction from GTAW molten-pool images (International Journal of Advanced Manufacturing Technology, 2023, doi:10.1007/s00170-023-12855-3); real-time penetration prediction for A-TIG welding of 10 mm 316LN (Scientific Reports, 2025); multi-sensor transfer learning for keyhole depth in laser welding of 780DP steel (PMC12429609).
  • 19.Ghorbani, A. & Zou, J. (2019). Data Shapley: Equitable Valuation of Data for Machine Learning. ICML 2019.
  • 20.Catena-X Automotive Network e.V. PCF Rulebook v4 (2025-09) and Governance Framework. ODRL profile, Data Exchange Governance.
  • 21.IDSA, Catena-X and Cofinity-X (2026). Data Space Accelerator. 2026-06, internationaldataspaces.org. €23M total, up to €30K per company, with proof of actual data exchange required.
  • 22.Ministry of Economy, Trade and Industry, Japan (2025). Project Certification of Ouranos Ecosystem. 2025-05-09, meti.go.jp; plus the Bitkom Manufacturing-X survey (2023–2024) and remarks from VDMA representatives.

Earlier Pebblous reports