Model-card coverage of the GOPAL declared-input schema
August 29, 2026 · View on GitHub
Generated by scripts/model-card-coverage.sh from docs/model-card-classification.json and the playground manifest. Do not edit by hand.
What this measures, and what it does not
GOPAL v2.0.0 reads 183 distinct declared input fields across the 35 EU and UK policy checks shipped to the validation preview. A declared input is one no tool can measure for you: whether a CE marking was affixed, whether required documentation has been retained for ten years.
This document maps how many of those inputs a fully populated standard Hugging Face model card can supply. It is a measurement of model-card-to-policy-schema overlap. It is explicitly not a percentage of legal compliance, nor of all applicable regulatory obligations, and no figure here should be quoted as one.
Three reasons the denominator is ours rather than the law's:
- The schema decides the granularity. Annex IV technical documentation is collapsed into a small number of completeness fields while CE marking is split across five. Refactoring these policies would move the total without a single word of the law or of a model card changing.
- The library does not cover everything. The EU AI Act matrix lists what is implemented and what is not; the not-implemented rows are absent from this denominator entirely.
- No system is subject to all of it. The fields span provider, deployer, importer, distributor and GPAI duties, which attach to different actors. The union is not any one organisation's obligation set.
Every field is classified below against what a card structurally contains: its YAML metadata (license, datasets, metrics, model-index, base_model, pipeline_tag, library_name) and the sections of the standard template. Answerable means the card itself establishes the fact, without inference or inspecting the repository around it.
Result
| Facts | Share | |
|---|---|---|
| Answerable from a model card | 5 | 3% |
| Partially informed by one | 23 | 13% |
| Untouched by one | 155 | 85% |
| Total | 183 |
A complete, well-written model card supplies 5 of 183 of these inputs outright and gives you material for 23 more. The standard template does not prompt for the remaining 155.
Much of that remainder is organisational, process and supply-chain evidence: who signed the declaration of conformity, whether a notified body was involved, whether the deployer told the affected workers, whether logs are kept for six months. Some of it is not organisational at all and needs technical work a card simply does not carry, such as adversarial testing, cybersecurity controls, dataset quality and fairness measurement.
A card can of course be extended with sections of your own. The finding is that the standard template does not ask for these, so at the point you add them you are building a compliance record that happens to live in a card, not relying on the model-card structure.
What a card does answer
| Declared fact | Where in the card |
|---|---|
accuracy.metrics_declared | model-index / the Evaluation section is exactly a declaration of metrics |
documentation.evaluation_results_documented | the Evaluation section |
documentation.training_process_documented | the Training Details section |
governance.third_party_model_in_use | base_model in the YAML names the upstream model |
transparency.system_purpose_documented | the Uses / Direct Use section |
What it partially informs
These are the ones worth arguing about. In each case the card raises the topic without asserting the fact, and the distinction matters: a Bias, Risks and Limitations section is not the same claim as "a bias assessment was completed".
| Declared fact | Why it is only partial |
|---|---|
accuracy.declared_in_instructions | the card is documentation; whether it is the instructions for use is a legal question |
datasets.complete_to_the_extent_possible | not asserted by any card section |
datasets.contextual_characteristics_considered | Article 10(4) asks about the deployment setting, which a card rarely knows |
datasets.errors_addressed | not asserted by any card section |
datasets.relevant | the datasets field names them; relevance is a claim about them |
datasets.sufficiently_representative | Bias, Risks and Limitations often discusses representativeness without asserting it |
documentation.explainability.completeness | some cards discuss interpretability |
documentation.technical_documentation.completeness | a card is a summary, not the Annex IV file |
documentation.testing_process_documented | Evaluation covers results more often than process |
explainability.method_documented | sometimes present in Technical Specifications |
fairness.bias_assessment_completed | Bias, Risks and Limitations is a discussion, not a completed assessment |
fairness.max_disparity | only if fairness metrics happen to be reported |
fairness.protected_characteristics_tested | named occasionally, and rarely as a tested set |
model.cumulative_training_compute_flops | Training Details and Environmental Impact sometimes give compute, rarely FLOPs |
model.free_and_open_source | license names a licence, but the Article 53(2) open-source treatment also turns on access, use, modification and distribution rights and on parameters being public. A licence string is evidence toward the test, not the test |
model.general_purpose | pipeline_tag and library_name hint at it, but GPAI is a legal classification with a compute threshold |
model.parameters_publicly_available | established by inspecting the repository, not by the card. A card may link weights without asserting their availability |
robustness.performance_thresholds_defined | metrics are given; thresholds are a decision about them |
system.performs_biometric_categorisation | pipeline_tag can indicate the task |
system.sources | the datasets field can indicate provenance of training images |
training_content.public_summary_available | the Training Data section is arguably one, which is untested |
training_content.summary_sufficiently_detailed | sufficiency is judged against Article 53(1)(d), not the card |
transparency.ai_use_disclosed | disclosure is to the affected person, not in the card |
What it does not reach
Grouped by the part of the input they belong to.
| Group | Facts | Examples |
|---|---|---|
system | 30 | affects_individuals, annex_iii_1_a_biometric, annex_iii_category, … |
deployer | 12 | affected_persons_informed, affects_natural_persons, controls_input_data, … |
governance | 11 | accountable_person_named, assumptions_documented, bias_examined, … |
provider | 10 | accessibility_requirements_met, ce_marking_affixed, conformity_assessment_completed, … |
logging | 8 | automatic_recording_enabled, covers_system_lifetime, identifies_risk_situations, … |
importer | 7 | documentation_retained_ten_years, identification_provided, storage_and_transport_conditions_adequate, … |
distributor | 6 | corrective_action_taken, non_conformity_discovered, storage_and_transport_conditions_adequate, … |
oversight | 6 | assigned_persons_competent, automation_bias_addressed, can_disregard_output, … |
risk_management | 6 | lifecycle_monitoring_in_place, mitigation_measures.completeness, monitoring_system.completeness, … |
ce_marking | 5 | affixed, digital_marking_accessible, notified_body_involved, … |
declaration | 5 | annex_v_content_complete, drawn_up, machine_readable, … |
registration | 5 | annex_viii_information_complete, exemption_assessment_registered, provider_registered, … |
assessment | 4 | completed, harmonised_standards_applied, notified_body_involved, … |
decision | 4 | article_9_condition, meaningful_human_involvement, significant, … |
logs | 4 | retained, retention_months, role, … |
redress | 4 | human_review_available, response_timeframe_days, route_communicated, … |
robustness | 4 | failure_handling_documented, fallback_or_fail_safe_documented, feedback_loop_risk_addressed, … |
safeguards | 4 | decision_contestable, human_intervention_available, information_provided, … |
systemic_risk | 4 | cybersecurity_protection_adequate, model_evaluation_with_adversarial_testing, risks_assessed_and_mitigated, … |
copyright | 2 | policy_in_place, respects_article_4_3_reservations |
cybersecurity | 2 | adversarial_attacks_addressed, controls_in_place |
downstream | 2 | annex_xii_information_provided, enables_downstream_compliance |
model | 2 | commission_designated_systemic_risk, systemic_risk |
special_category | 2 | processed_for_bias_correction, safeguards_in_place |
authorisation | 1 | prior_authorisation_obtained |
documentation | 1 | annex_xi_complete |
explainability | 1 | decision_rationale_available |
fairness | 1 | legal_rights_review_completed |
notification | 1 | commission_notified |
security | 1 | security_testing_completed |
Method and limits
- The classification is a judgement, not a measurement, which is why every field is listed with its reason rather than summarised. The
partiallist is where to disagree. - It assumes a complete card following the standard template, which is the best case. Cards that omit sections would map to fewer inputs; no representative sample of published cards has been measured here.
- Applicability is not modelled. Regulation (EU) 2026/1744, in force 27 July 2026, moved Chapter III Sections 1 to 3 to 2 December 2027 for Annex III high-risk systems and 2 August 2028 for Annex I. Prohibitions and GPAI obligations were not deferred. Many inputs counted here therefore relate to duties that are modelled but not yet applicable.
- The UK principles are non-statutory. They are addressed to regulators rather than imposed on firms, as the UK matrix states. UK GDPR Articles 22A to 22D are binding in their own right; the five cross-sector principles are assurance evidence, not a statutory test.
- It counts distinct declared facts, not policies. Fields read by more than one policy are counted once.
- Measured metrics and policy thresholds are excluded: neither is something a person declares.
--checkfails if the classification names a field no policy reads, so this cannot silently drift from the library.