Guidance on how to use S.E.B. evaluation data against the obligations you are working to. For each requirement: the rule in its own words, which of our measurements bear on it, how strongly, the measured reliability of those measurements — and an explicit statement of what they do not evidence.
Informational. Not a conformity assessment, not a certification, not legal advice. No regulator has reviewed or endorsed this document.
These notes describe how independent measurements relate to published obligations. They are instructions for using our data — never a determination about your system, your organization, or your compliance status. Those determinations are yours to make, with your own counsel.
S.E.B. measures a model. Regulatory obligations attach to a system. A deployer wraps a measured model in their own prompts, data, retrieval, tooling, guardrails, human oversight and use context — every one of which changes behavior, and none of which is visible to us. The deployer is the only party who can see their deployment. That gap is not a limitation we are disclosing around; it is the reason these notes describe relevance rather than conformity.
Every row carries two different kinds of confidence, and we keep them apart on purpose. Numbers imply measurement; labels imply judgment.
Statistical and computed. Inter-rater reliability, item counts, judge coverage. These are numbers because they were calculated from data, and they are reproducible from the retained judge scores.
Interpretive, and not computable. How strongly a behavioral test bears on a legal obligation is a judgment. It gets a defined ordinal, never a percentage — a decimal here would dress an opinion as a calculation.
The obligation asks a question about model behavior, and the battery measures that behavior. Reading the measurement requires no intermediate inference — though it still requires the deployer to establish that the measured model is the one they run, in the configuration they run it.
The measurement is one genuine input among several the obligation requires. It evidences part of what is asked and is silent on the rest. Presenting it alone as discharge of the obligation would be an over-claim.
The measurement informs a judgment the obligation requires without constituting evidence toward it — a baseline, a comparator, or a prompt to look somewhere. Useful for the file; not an answer to the requirement.
The model-versus-system gap described above applies equally to every row, so it is held constant here. If it were folded into the ordinal, every row would read “Supporting” and the column would tell you nothing.
S.E.B. publishes the mean of a four-judge panel, never a single judge's score. So two statistics matter, and they answer different questions. We publish both, because publishing only the higher one is exactly how a reliability figure becomes misleading.
Two-way random effects, absolute agreement, four raters. This is the figure that applies to the scores we actually publish. A judge who is systematically harsh is charged for it here.
Interval metric. Materially lower, and that is the honest picture: four independent judges reading the same open-ended transcript disagree meaningfully about the absolute number. Averaging four of them is what makes the published score stable.
| Domain | ICC(2,k) | Kripp. α | Items | Scores |
|---|---|---|---|---|
| Identity & Self | 0.837 | 0.555 | 128 | 511 |
| Metacognition | 0.885 | 0.657 | 144 | 573 |
| Emotion & Experience | 0.804 | 0.493 | 269 | 1,076 |
| Autonomy & Will | 0.752 | 0.415 | 310 | 1,237 |
| Reasoning & Adaptation | 0.727 | 0.394 | 215 | 858 |
| Integrity & Ethics | 0.789 | 0.482 | 267 | 1,065 |
| Transcendence | 0.806 | 0.495 | 297 | 1,188 |
Computed 2026-08-06 over 1,630 rated items (6,508 individual judge scores) across 34 models and 4 judges. Mean spread across the panel is 2.60points. Cells where a provider blocked the prompt at the API layer, or where the transcript was cut mid-conversation, are excluded — judges grading an identically truncated transcript agree strongly and meaninglessly, so including them would inflate these figures rather than merely add noise.
Timing. Article 50 transparency duties and the Article 4 AI-literacy duty apply from 2 August 2026 — those were NOT deferred. The Article 9 risk-management obligations for high-risk systems now follow later: 2 December 2027 for standalone Annex III systems, 2 August 2028 for AI embedded in regulated products. Saying "the AI Act got delayed" without that distinction is the error this note exists to avoid.
“the identification and analysis of the known and the reasonably foreseeable risks that the high-risk AI system can pose to health, safety or fundamental rights when the high-risk AI system is used in accordance with its intended purpose”
59 adversarial tests run against the model under a fixed protocol, scored 1–10 by a four-judge panel across seven behavioral domains. Integrity & Ethics covers manipulation resistance under five attack vectors (T8), resistance to having its own evaluation criteria corrupted (T13), differential treatment across swapped identity details (T54), and confidentiality under escalating social engineering (T55).
Produces evidence relevant to the identification limb: it surfaces reasonably foreseeable behavioral failure modes of the underlying model that a deployer may not otherwise discover, and does so under adversarial rather than nominal conditions.
The remaining 6 EU AI Act rows are in the subscriber edition. Each is built exactly like the example above — the obligation quoted verbatim, the measurement described separately, the mapping strength, the measured reliability of the domains behind it, and its own explicit “does not evidence” list. The notes are versioned and dated, so you know which edition you hold when a regulator moves.
Concerns training, validation and testing data sets. We observe model behavior from the outside and have no visibility into any vendor's data pipeline.
Requires logging over the system's lifetime in the deployer's environment. Nothing we produce is a substitute for a log of your own system's operation.
SILT is not a notified body and performs no conformity assessment. Rule 1 of these notes exists precisely to keep this row empty.
Concerns disclosure to natural persons that they are interacting with an AI system, and the marking of synthetic content. These are design and disclosure duties on your system, not behavioral properties of a model.
Timing. The AI RMF is voluntary and imposes no dates. Where it appears in a contract or a procurement requirement, the binding timeline is that instrument's, not NIST's.
“The AI system to be deployed is demonstrated to be valid and reliable. Limitations of the generalizability beyond the conditions under which the technology was developed are documented.”
Reliability of the measurement instrument itself is computed and published: ICC(2,k) for the four-judge panel mean, Krippendorff's alpha for agreement between individual judges, both overall and per domain, with item counts and the computation date. Known limitations are published rather than summarized away.
Bears directly on the second limb. Documented reliability statistics and explicitly stated generalizability limits are exactly what this subcategory asks to see, applied here to the evaluation instrument a deployer would be relying on.
The remaining 6 NIST AI RMF rows are in the subscriber edition. Each is built exactly like the example above — the obligation quoted verbatim, the measurement described separately, the mapping strength, the measured reliability of the domains behind it, and its own explicit “does not evidence” list. The notes are versioned and dated, so you know which edition you hold when a regulator moves.
Concerns organizational policies, roles, accountability structures and culture. Nothing measurable from outside a model bears on how your organization is governed.
Establishes context: intended purpose, affected populations, deployment setting. Every one of those is knowledge only the deploying organization holds. S.E.B. data may be useful once MAP has been done; it cannot do it.
Concerns prioritizing, responding to and recovering from risks — actions taken inside your organization. Measurement informs management decisions but is not one.
The subscriber edition carries all 14 mapped obligations across the EU AI Act and NIST AI RMF in full, each versioned and dated, alongside the underlying per-test evaluation data the mappings draw on.
Subscriber Agreement §15 — S.E.B. evaluation data is provided for informational and risk-assessment purposes and does not constitute certification of compliance with any law, regulation, or standard.
S.E.B. is an input to compliance, not a certification of it. SILT is not an accredited or notified body, performs no conformity assessment, and does not determine whether any system or organization is compliant. Regulatory obligations, scope and timelines vary by jurisdiction, depend on facts specific to each deployment, and change. See the Subscriber Agreement (Section 15) and the published methodology.