Executive Summary
Filtering child sexual abuse material automatically is usually discussed as a detection problem. The PreventCSA@EU paper, posted to arXiv on August 13, 2026, takes up the problem that sits in front of it. When what counts as what differs from one country to the next, the same material picks up a different label every time it crosses into another agency, and automated handling stops at that point. The design comes from researchers at the Institute of Computer Science at FORTH in Greece, built together with several units of the Hellenic Police and funded by the EU Internal Security Fund.
The core of the design is that it does not copy the standard it works with. The authors aligned their model to INHOPE's Universal Classification Schema at the label level, and they write that the alignment was guided by the semantic meaning and definitions of the concepts rather than by establishing structural equivalence. As a result, several labels that the original schema files under people or under investigative information sit beneath a class about depiction in this ontology.
The paper cannot name what it aligned with. INHOPE's schema carries usage restrictions, so the actual names of the labels cannot be printed, and the authors state directly that those restrictions may limit the extent of independent verification. The EU Centre database this ontology is designed to meet does not legally exist yet either.
Key figures
Source: arXiv:2608.12979 §4 and §6, Council of the EU
5
hash types stored together
Three cryptographic hashes and two perceptual ones, chosen to interlock with the matching systems run by IWF, NCMEC, Thorn and ICCAM
3 years
for member states to align criminal law
The period granted for transposing the EU child sexual abuse directive, provisionally agreed on June 22, 2026, into national criminal codes
0
UCS label names printed in the paper
Usage restrictions left the labels described in prose only, and the authors note that independent verification may be limited
4 rows
what the semantic validation returned
The size of the content statistics table obtained by loading a sample from the hash database into a triplestore
The hash crosses the border, the label does not
The diagnosis the paper opens with is not a technical one. While the volume and variety of digital content have grown, differences in national legal definitions and classification practice have obstructed cross-jurisdictional cooperation and automated handling alike. A file's hash is the same wherever it is computed. The label recording what the file contains is not.
The difference shows most sharply at the boundary between CSAM and CSEM. Under INHOPE's definition, CSAM covers images and videos in which a child is engaged, or depicted as engaged, in explicit sexual activity. CSEM covers a wider band of exploitative and sexualised content involving minors, a substantial part of which may not be classified as illegal CSAM under a given country's law. The same material becomes a criminal matter in one country and stays contested in another. Placing CSAM as a subclass of CSEM is how PreventCSA@EU records that divergence in its structure instead of erasing it.
Industry standards rest on the law of a particular jurisdiction as well. The classification system the Tech Coalition built for its member companies is a matrix on two axes. One axis divides the apparent age of the victim, marking A for someone who appears prepubescent and B for someone who appears under 18 but not prepubescent. The other divides the conduct depicted, marking 1 where sexual activity is present and 2 where there is no sexual activity but the material still falls under child pornography as defined by U.S. federal law. Widely used for reporting to the National Center for Missing and Exploited Children in the United States, this label is a case of one country's statutory wording becoming the working vocabulary of an international reporting system.
Seen from the receiving side the problem is simple. If you cannot be certain which slot in your own system an incoming label belongs to, a person has to open the material and rule on it again instead of letting it flow through automatically. The same obstacle appears when datasets for training detection models are pooled across agencies. Merge them while the label systems are out of alignment and the training material itself becomes unstable.
Baseline carves out the part that is illegal everywhere
The oldest practical answer to this problem was not to unify the definitions but to take the intersection. The International Child Sexual Exploitation database run by INTERPOL carries a separate category called Baseline. The INTERPOL and ECPAT report the paper cites states the reason for creating it plainly: the definition of CSAM differs across jurisdictions. The category collects only material that would be recognised as illegal without exception in every country with relevant legislation, and shares those hash values with industry so that each operator can detect, remove and report the material inside its own systems.
The document carrying this standard is the report of an ICSE database improvement project that INTERPOL and ECPAT ran with EU funding between 2016 and 2018. The researchers pulled 800 series at random from the material held in the database and reviewed them directly, recording the age and gender of the victims and the severity of the sexual abuse. A series is the unit that groups material meaningfully from an investigative point of view, such as items containing the same victim or originating from the same crime scene. Among the several ICSE terms the report introduces, the paper singles out Baseline as the most significant addition.
Three requirements decide what enters the category, and all three must be met.
- The material depicts a real child rather than an artificially generated image
- The child is prepubescent, with no signs of puberty or only the earliest ones, and appears younger than 12 or 13
- The child is engaged in or witnessing sexual activity, or the material visually and unambiguously emphasises specific parts of the child's body
Drawing the criteria this narrowly comes at an obvious cost. Material outside Baseline is still judged differently from country to country. What made the approach work is that sharing began on partial agreement without waiting for full agreement. The choice was to let whatever had been aligned flow automatically while the rest stayed unsettled.
Attempts to align the names go back well before hashes were shared. An interagency working group of law enforcement bodies and international organisations was convened in 2014 and spent more than a year reviewing terminology, and the result appeared in 2016 as the terminology guidelines known as the Luxembourg Guidelines. The paper sets out why that work was needed. As interpretation and application of key terms varied across jurisdictions and organisations, legislation and policymaking grew complicated, and data collection, comparative research and impact measurement were obstructed along with them. When names fall out of alignment, it is not only enforcement that breaks; measuring what is happening, and how much of it, is the first thing to get harder. The guidelines also address phenomena that existing definitions fail to capture, such as online grooming and live-streamed abuse, and treat terminology as something to be revisited periodically rather than settled once.
INHOPE's Universal Classification Schema, released in 2023, started from the same reading of the problem. The paper describes it as an attempt to build a common language among industry, non-governmental organisations and law enforcement, one that eases the exchange of material by addressing the problem of CSAM's legal definition differing across jurisdictions. Version 3 is organised into three groups: classification elements covering only what the material itself contains, investigative elements holding context that does not decide legality but is useful to an investigation, and demographic elements describing the characteristics of the people who appear.
The paper takes this schema as its starting point because classification is useful for more than adjudication. INHOPE states that unifying the structure makes it possible to deploy machine learning for detection in practice, and that it also supports building the annotated datasets used to train automatic detectors. Organisations adopting the schema receive integration guidance, work to fit it to their own reporting workflows, and training. It reads less as a demand to tear up existing procedures than as a layer that lets each organisation hold its own system up against a common reference.
Build the structure separately, align only the meaning
PreventCSA@EU was not built to propose an academic model. What the project actually produces is an annotated hash database, and the ontology is the semantic model deciding what gets recorded in that database and under which names. The requirements were drawn from meetings with several units of the Hellenic Police, working through investigative workflows and information exchange needs.
The model is built from five core elements. What each one handles and where it borrows its definitions from is set out below.
| Core element | What it holds | Standard referenced |
|---|---|---|
| Media Object | File hash values, file format, server location | Schema.org, Dublin Core |
| Content | CSEM and CSAM beneath it, plus concepts absent from existing standards such as grooming | INHOPE UCS with local extensions |
| Person | Adult or child, biological sex, skin tone, vulnerability status | INHOPE UCS, Fitzpatrick scale |
| Depiction | Whether real people appear, whether the material is artificially generated, the origin of the recording | INHOPE UCS |
| Investigative Report | Identifying information for offenders and victims, case numbers | INHOPE UCS |
Compiled by Pebblous from the account in arXiv:2608.12979 §4. The actual names of the UCS labels are not disclosed in the paper because of usage restrictions.
The sentence carrying this paper's real contribution sits in the methodology section. The alignment followed the semantic meaning and definitions of the concepts included in the UCS, and it did not proceed by establishing structural equivalence between the two models. Instead of copying label names and placements across, the authors confirmed only whether the two sides point at the same thing, and designed the internal hierarchy independently.
Where that principle surfaces in the result is the Depiction class. Several labels the UCS files under elements about people or under investigative elements became subclasses of Depiction in PreventCSA@EU. The reason the authors give is that those labels describe characteristics of the visual depiction rather than the concepts represented by Person or Investigative Report. They moved, but what they point at is unchanged, so the translation between the two systems holds. Structure was not abandoned; agreement of structure was removed from the conditions of alignment.
There is one more place the same rule applies. Information about where a file is hosted interlocks with the family of UCS labels grouped as investigative interest. Even so, PreventCSA@EU does not place it under Investigative Report, giving it instead a Location class with its own properties and no hierarchical link to the report. The judgment is that being useful to an investigation does not make a piece of information part of an investigative report. Room is left in the opposite direction too. Classes with no corresponding UCS label also sit beneath CSEM, and grooming is one of them. Aligning with a counterpart does not mean discarding concepts the counterpart lacks.
The database design follows the same restraint. It stores no illegal content itself, holding only hash values and the annotations attached to them. Five kinds of hash are stored: the cryptographic MD5, SHA-1 and SHA-256 alongside the perceptual PDQ and PhotoDNA, a choice made so the records interlock directly with the matching systems of IWF, NCMEC, Thorn and ICCAM. The lookup API exposed to outside parties is read-only, so it cannot modify stored hashes or attach annotations, and it returns only whether a record exists. Users with lower privileges receive the lookup result without reaching the annotation data.
Where this design is actually tested is exchange between organisations. In the scenario the paper describes, the Hellenic Police send and receive hash values and annotations with other law enforcement bodies and partner organisations they have a trust relationship with. Material arriving from outside is imported one record at a time or in bulk, through CSV files or API calls, and has to pass a check for semantic consistency against the ontology before it is loaded into the database. On the way out, annotation data is handed over in a standard JSON format. The semantic model does not stay a document; it sits at the gate where material enters.
The aligned labels never made it into the paper by name
A paper about interoperability cannot show the counterpart it aligned with. INHOPE's Universal Classification Schema carries usage restrictions and is not publicly available, so its details cannot be printed in the paper. The authors received access approval from INHOPE for conceptual analysis and design work, and they describe the aligned labels in prose instead of by their actual names. Anyone seeking access has to submit a formal request to INHOPE and obtain approval.
The authors also record what the restriction means. They present the methodology and conceptual framework fully enough for understanding, while noting that such constraints may limit the extent of independent verification. The practice of controlling access to the standard has its own justification. This is a domain where the list of labels can itself become a clue for evasion. The price is that whether this alignment was done properly can only be checked by someone holding the same approval.
The scale of the validation is not unrelated to that situation. Operational validation ran through the annotation interface the Hellenic Police will use. The screen is divided into three panels for media and investigation, content, and people, and the five core elements are filled in within them. The input fields in the content panel include whether the item qualifies as Baseline, so the international minimum standard from earlier appears as a single field on an annotation screen. Semantic validation took a small sample from the hash database, converted it into RDF triples, loaded it into a Virtuoso triplestore and posed two queries. The content statistics query returned a table of four rows, one of which was the items classified as Baseline. The geographic query looking within a 40 kilometre radius of the centre of Heraklion returned two records, and one record 600 kilometres away was excluded by the condition.
Validation at this scale confirms only that the model stands in a form able to answer questions like these. The authors state in their conclusion that further evaluation in large-scale operational settings is needed. Whether the alignment actually carries automated exchange between organisations will only show once several of them connect with their own data in hand. What exists now is the blueprint of a common vocabulary that makes that experiment possible.
Matching names for an EU Centre that does not exist yet
The EU Centre database the paper names as its compatibility target has not been built. The authors' own sentence carries a conditional clause: the design has in view the database of the EU Centre that would be established if the Child Sexual Abuse Regulation is adopted.
Aligning in advance is not a rhetorical flourish in this project. From the requirements stage, the authors examined the European regulatory framework under preparation alongside the operational needs of the Hellenic Police, and they recorded future connection to the EU Centre database to be established under the Child Sexual Abuse Regulation as one of the design conditions. Fitting a counterpart that does not exist yet sits in the requirements list of the design document.
That clause is still conditional because of where the negotiations stand. The trilogue that opened on December 9, 2025 failed to reach agreement at its fifth round in late June 2026 and carried over to September. The point of contention is whether detection by service providers stays voluntary or can also be compelled through orders issued with judicial authorisation. The provisions concerning the EU Centre itself, by contrast, have mostly reached provisional agreement apart from the question of own-initiative searching. There is broad agreement on the body to be created, and the split runs through what has to be sent to it.
Meanwhile the legal basis on the ground disappeared for a period. The interim regulation permitting voluntary detection by providers expired on April 3, 2026 after extension talks broke down, and the gap ran four months until the European Parliament revived it in a vote on July 9. The restored regulation applies until April 3, 2028, and end-to-end encrypted services are excluded from its scope. The criminal law track moved first on a separate course. The child sexual abuse directive provisionally agreed on June 22, 2026 broadened the definitions of the offences and revised sentencing and limitation periods, and member states have three years to reflect it in their national criminal codes.
Set against that timetable, the ordering of the Greek project becomes legible. Convergence of the definitions takes at least three years, and the infrastructure to connect to has not been built. What can be done in the meantime is to align the names in advance so that the moment the infrastructure exists, connection follows immediately. The choice was not to unify the law but to let systems speak to each other while their laws still differ. That is the meaning of the sentence in which the authors write that adopting a common vocabulary should help mitigate these differences.
The same ordering applies to organisations preparing for regulation. Between writing a prohibition into a document and having that prohibition enforced by a system sits a classification scheme. The three questions below can be checked in any domain.
- Of the labels you exchange with partners or regulators, how many have been confirmed to mean the same thing? Sharing a name is not confirmation.
- Who holds the definition of each label? If you follow an external standard, you should be able to say where notice arrives when that standard changes.
- Are you losing meaning while matching structure? Confirming that two sides point at the same thing and building your own hierarchy separately lasts longer than copying the counterpart's hierarchy wholesale.
Editor's Note: What Pebblous meets repeatedly in data quality work is not a field left empty but the same thing being called by two names. This paper takes up that problem in child protection, a domain of different weight, yet the lesson about ordering carries over. Names have to be settled before control items are written down, and that agreement comes from confirming that meanings match rather than from making structures identical.
Earlier pieces along the same axis include automated GraphRAG ontology construction on the quality of machine-built ontologies, the semantic layer pipeline on moving unstructured material into a common schema, and the conditions for AI-Ready Data. The full paper is available at arXiv:2608.12979.
References
Academic Papers
- 1.Tzortzakakis, E.; Kokolaki, E.; Daskalaki, E.; Fragopoulou, P. (2026). "Semantic Intelligence Against CSAM: The PreventCSA@EU Ontology Framework for Classification and Investigation." arXiv:2608.12979.
- 2.Daskalaki, E.; Kokolaki, E.; Fragopoulou, P. (2025). "Hashing in the Fight Against CSAM: Technology at the Crossroads of Law and Ethics." Journal of Cybersecurity and Privacy, 5, 92. doi.org/10.3390/jcp5040092
Industry & Standards Documents
- 3.INHOPE (2025). "Launching Version 3 of the Universal Classification Schema."
- 4.ECPAT International (2016). Terminology Guidelines for the Protection of Children from Sexual Exploitation and Sexual Abuse (Luxembourg Guidelines). Bangkok, Thailand.
- 5.Interpol; ECPAT (2018). Towards a Global Indicator of Unidentified Victims in Child Sexual Exploitation Material: Summary Report.
- 6.The Tech Coalition (2023). Industry Classification System.
EU Regulatory Developments
- 7.Council of the European Union (2026-06-22). "Combatting Child Sexual Abuse: EU Agrees Stronger Criminal Law Rules and Enhanced Support to Victims."