Chapter III · High-risk AI systems
Article 10 — Data and data governance
Article 10 governs the material a high-risk system is built from. It requires the training, validation and testing sets to be documented, examined for bias, and to be relevant, sufficiently representative, as free of errors as possible and complete for the intended purpose — and since the Digital Omnibus, also to satisfy the conditions in the new Article 4a whenever special categories of personal data are processed to find that bias. That cross-reference is the change most providers have not noticed: a data protection condition has become a product conformity criterion, so missing one of the six conditions in Article 4a(1) does not merely breach the GDPR, it means the data sets fail Article 10(1). Two supervisory routes then open on one set of facts, and the Regulation says nothing about how they coordinate. What the article does not do is tell anyone how much is enough. Nine of its predicates carry no measure at all: relevant, sufficiently representative, to the best extent possible, complete, appropriate, strictly necessary. This is not careless drafting. It is a deliberate allocation of judgment to the provider, subject to review afterwards — which is precisely why the judgment has to be recorded at the moment it is made, with the reasoning, the evidence and the version of the law it was made against.
Official text
Each paragraph records where its wording comes from. Only text reproduced unchanged from the Official Journal is authentic; consolidated text is editorial and has no legal value. Paragraphs, subparagraphs and points each have their own link.
High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2, 3 and 4 of this Article and in Article 4a(1) whenever such data sets are used.
Consolidated text — no legal value · amended by Regulation (EU) 2026/1744
Training, validation and testing data sets shall be subject to data governance and management practices appropriate for the intended purpose of the high-risk AI system. Those practices shall concern in particular:
- (a)
the relevant design choices;
- (b)
data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection;
- (c)
relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation;
- (d)
the formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent;
- (e)
an assessment of the availability, quantity and suitability of the data sets that are needed;
- (f)
examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations;
- (g)
appropriate measures to detect, prevent and mitigate possible biases identified according to point (f);
- (h)
the identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed.
Authentic — as published in the Official Journal
Training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used. Those characteristics of the data sets may be met at the level of individual data sets or at the level of a combination thereof.
Authentic — as published in the Official Journal
Data sets shall take into account, to the extent required by the intended purpose, the characteristics or elements that are particular to the specific geographical, contextual, behavioural or functional setting within which the high-risk AI system is intended to be used.
Authentic — as published in the Official Journal
Paragraph 5 — deleted by Regulation (EU) 2026/1744 · no text in force
For the development of high-risk AI systems not using techniques involving the training of AI models, paragraphs 2, 3 and 4 of this Article and Article 4a(1) shall apply only to the testing data sets.
Consolidated text — no legal value · amended by Regulation (EU) 2026/1744
The formal analysis of Article 10
Article 6 is formalised first. The remaining articles follow.
Article 10 is formalised node by node: each rule as a deontic position with its operator, each exception with its rank, each predicate resolved against the definitions in Article 3.
What members get, per article
- the rule logic: every norm as a formal position, with the defeater chain that decides which exception wins
- the competency questions and their answers, every unanswered one marked as a gap and named
- the ontology: predicates bound to AISV (AI Standardisation Vocabulary) and to the definitions they depend on, exportable as JSON-LD and OWL
- the documentation: the Article 6(4) assessment record generated from a fact set, with its derivation and the version of the law it was decided against
Built for providers claiming the 6(3) derogation, for the counsel who has to defend that claim, and for the auditor who reads it afterwards.
Access is invite-only and opening in stages. Article 6 is formalised first; the remaining articles follow.
No mail client? Write to hello@signatu.com
Already a member? Sign in
What the article requiresUnder review
- 01Apply data governance and management practices covering design choices, data origin and, for personal data, the original purpose of collection.
- 02Document data collection processes and the provenance of the data.
- 03Carry out data preparation: annotation, labelling, cleaning, updating, enrichment and aggregation.
- 04State the assumptions the data is meant to represent, including what the data is supposed to measure.
- 05Assess the availability, quantity and suitability of the datasets needed.
- 06Examine the datasets for possible biases likely to affect health, safety or fundamental rights, or to lead to prohibited discrimination, and take measures to detect, prevent and mitigate them.
- 07Identify and address data gaps or shortcomings that prevent compliance.
- 08Ensure datasets are relevant, sufficiently representative, and to the best extent possible free of errors and complete in view of the intended purpose.
- 09Take account of the geographical, contextual, behavioural or functional setting in which the system will be used.
- 10Where strictly necessary for bias detection and correction under Article 10(2), points (f) and (g), providers may exceptionally process special categories of personal data, subject to the conditions and safeguards in Article 4a(1).
In practiceUnder review
Article 10 is where AI compliance and data protection meet. The bias examination is not a one-off statistical exercise; it has to be repeatable and evidenced, and the record has to survive the arrival of a new dataset version two years later.
Standards addressing this articleUnder review
- prEN 18284Quality and governance of datasets in AI systemsSent to the Commission in July 2026 for its pre-enquiry check. Commission comments received in August are under review.
- prEN 18283Concepts, measures and requirements for managing bias in AI systemsSent to the Commission in July 2026 for its pre-enquiry check. Commission comments received in August are under review.