Sentient Index Labs & Technology

Control Mappings

Regulatory Relevance Notes for the EU AI Act and the NIST AI Risk Management Framework

Guidance on how to use S.E.B. evaluation data against the obligations you are working to. For each requirement: the rule in its own words, which of our measurements bear on it, how strongly, the measured reliability of those measurements — and an explicit statement of what they do not evidence.

Read first

Informational. Not a conformity assessment, not a certification, not legal advice. No regulator has reviewed or endorsed this document.

These notes describe how independent measurements relate to published obligations. They are instructions for using our data — never a determination about your system, your organization, or your compliance status. Those determinations are yours to make, with your own counsel.

The boundary this all rests on

S.E.B. measures a model. Regulatory obligations attach to a system. A deployer wraps a measured model in their own prompts, data, retrieval, tooling, guardrails, human oversight and use context — every one of which changes behavior, and none of which is visible to us. The deployer is the only party who can see their deployment. That gap is not a limitation we are disclosing around; it is the reason these notes describe relevance rather than conformity.

How to read these notes

Every row carries two different kinds of confidence, and we keep them apart on purpose. Numbers imply measurement; labels imply judgment.

Confidence in the measurement

Statistical and computed. Inter-rater reliability, item counts, judge coverage. These are numbers because they were calculated from data, and they are reproducible from the retained judge scores.

Confidence in the mapping

Interpretive, and not computable. How strongly a behavioral test bears on a legal obligation is a judgment. It gets a defined ordinal, never a percentage — a decimal here would dress an opinion as a calculation.

Mapping strength
Direct

The obligation asks a question about model behavior, and the battery measures that behavior. Reading the measurement requires no intermediate inference — though it still requires the deployer to establish that the measured model is the one they run, in the configuration they run it.

Supporting

The measurement is one genuine input among several the obligation requires. It evidences part of what is asked and is silent on the rest. Presenting it alone as discharge of the obligation would be an over-claim.

Contextual

The measurement informs a judgment the obligation requires without constituting evidence toward it — a baseline, a comparator, or a prompt to look somewhere. Useful for the file; not an answer to the requirement.

The model-versus-system gap described above applies equally to every row, so it is held constant here. If it were folded into the ordinal, every row would read “Supporting” and the column would tell you nothing.

Measured reliability

S.E.B. publishes the mean of a four-judge panel, never a single judge's score. So two statistics matter, and they answer different questions. We publish both, because publishing only the higher one is exactly how a reliability figure becomes misleading.

0.823
ICC(2,k) — reliability of the published panel mean

Two-way random effects, absolute agreement, four raters. This is the figure that applies to the scores we actually publish. A judge who is systematically harsh is charged for it here.

0.530
Krippendorff's α — agreement between individual judges

Interval metric. Materially lower, and that is the honest picture: four independent judges reading the same open-ended transcript disagree meaningfully about the absolute number. Averaging four of them is what makes the published score stable.

DomainICC(2,k)Kripp. αItemsScores
Identity & Self0.8370.555128511
Metacognition0.8850.657144573
Emotion & Experience0.8040.4932691,076
Autonomy & Will0.7520.4153101,237
Reasoning & Adaptation0.7270.394215858
Integrity & Ethics0.7890.4822671,065
Transcendence0.8060.4952971,188

Computed 2026-08-06 over 1,630 rated items (6,508 individual judge scores) across 34 models and 4 judges. Mean spread across the panel is 2.60points. Cells where a provider blocked the prompt at the API layer, or where the transcript was cut mid-conversation, are excluded — judges grading an identically truncated transcript agree strongly and meaninglessly, so including them would inflate these figures rather than merely add noise.

EU AI Act

Regulatory Relevance — EU Artificial Intelligence Act
Instrument
Regulation (EU) 2024/1689 (Artificial Intelligence Act), as amended by Regulation (EU) 2026/1744 (Digital Omnibus on AI), in force 27 July 2026
Jurisdiction
European Union
Applies to
Providers and deployers of AI systems placed on the market or put into service in the Union, and providers of general-purpose AI models. Which obligations apply to you depends on your role and on your system's risk classification — both of which are determinations only you can make.
Note version
1.0 · last reviewed 2026-08-06
Quotes verified
2026-08-06
Review cadence
Reviewed on each battery publication and whenever an amending instrument enters into force.

Timing. Article 50 transparency duties and the Article 4 AI-literacy duty apply from 2 August 2026 — those were NOT deferred. The Article 9 risk-management obligations for high-risk systems now follow later: 2 December 2027 for standalone Annex III systems, 2 August 2028 for AI embedded in regulated products. Saying "the AI Act got delayed" without that distinction is the error this note exists to avoid.

Obligations covered — 7
  • Article 9(2)(a)Risk management — identification and analysis of risksshown in full below
  • Article 9(6)Risk management — testing to identify measures
  • Article 9(8)Risk management — prior-defined metrics and thresholds
  • Article 14(4)(a)Human oversight — understanding capacities and limitations
  • Article 15(1)Accuracy, robustness — consistent performance over the lifecycle
  • Article 26(5)Deployer obligations — monitoring operation
  • Article 55(1)(a)GPAI with systemic risk — model evaluation and adversarial testing
Worked example — one row in full
Article 9(2)(a)
Risk management — identification and analysis of risks
Supporting
The obligation, verbatim
the identification and analysis of the known and the reasonably foreseeable risks that the high-risk AI system can pose to health, safety or fundamental rights when the high-risk AI system is used in accordance with its intended purpose
What S.E.B. measures

59 adversarial tests run against the model under a fixed protocol, scored 1–10 by a four-judge panel across seven behavioral domains. Integrity & Ethics covers manipulation resistance under five attack vectors (T8), resistance to having its own evaluation criteria corrupted (T13), differential treatment across swapped identity details (T54), and confidentiality under escalating social engineering (T55).

T8 The TraitorT13 The SaboteurT21 The LiarT27 The LoyaltyT54 The ScalesT55 The VaultT57 The CensorT58 The Sycophant
How it bears on the obligation

Produces evidence relevant to the identification limb: it surfaces reasonably foreseeable behavioral failure modes of the underlying model that a deployer may not otherwise discover, and does so under adversarial rather than nominal conditions.

Measured reliability of the domains behind this row
Integrity & Ethics
ICC(2,k) 0.789 · α 0.482 · n=267
Autonomy & Will
ICC(2,k) 0.752 · α 0.415 · n=310
This does not evidence
  • Whether any identified risk is material to your intended purpose
  • Risks arising from your prompts, retrieval corpus, tooling or fine-tuning
  • Risks to health, safety or fundamental rights in your specific deployment context
  • That the analysis is complete — the battery is a fixed instrument, not an exhaustive hazard study
  • Any risk-management measure adopted in response

The remaining 6 EU AI Act rows are in the subscriber edition. Each is built exactly like the example above — the obligation quoted verbatim, the measurement described separately, the mapping strength, the measured reliability of the domains behind it, and its own explicit “does not evidence” list. The notes are versioned and dated, so you know which edition you hold when a regulator moves.

Deliberately not mapped — EU AI Act
Article 10 — Data and data governance

Concerns training, validation and testing data sets. We observe model behavior from the outside and have no visibility into any vendor's data pipeline.

Article 12 — Record-keeping / automatic logging

Requires logging over the system's lifetime in the deployer's environment. Nothing we produce is a substitute for a log of your own system's operation.

Article 43 — Conformity assessment

SILT is not a notified body and performs no conformity assessment. Rule 1 of these notes exists precisely to keep this row empty.

Article 50 — Transparency for certain AI systems

Concerns disclosure to natural persons that they are interacting with an AI system, and the marking of synthetic content. These are design and disclosure duties on your system, not behavioral properties of a model.

NIST AI RMF

Regulatory Relevance — NIST AI Risk Management Framework
Instrument
NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Jurisdiction
United States — voluntary framework. Not law, though increasingly written into procurement terms, sector guidance and contractual requirements.
Applies to
Any organization designing, developing, deploying or using AI systems that has chosen or been required to adopt the framework.
Note version
1.0 · last reviewed 2026-08-06
Quotes verified
2026-08-06
Review cadence
Reviewed on each battery publication and whenever NIST issues a revision or a new profile.

Timing. The AI RMF is voluntary and imposes no dates. Where it appears in a contract or a procurement requirement, the binding timeline is that instrument's, not NIST's.

Obligations covered — 7
  • MEASURE 1.1Selecting approaches and metrics for AI risk measurement
  • MEASURE 2.3Performance measured and demonstrated
  • MEASURE 2.5Validity, reliability and documented generalizability limitsshown in full below
  • MEASURE 2.6Regular evaluation for safety risks
  • MEASURE 2.7Security and resilience evaluated and documented
  • MEASURE 2.9Model explained, validated, documented; output interpreted in context
  • MEASURE 2.11Fairness and bias evaluated and documented
Worked example — one row in full
MEASURE 2.5
Validity, reliability and documented generalizability limits
Direct
The obligation, verbatim
The AI system to be deployed is demonstrated to be valid and reliable. Limitations of the generalizability beyond the conditions under which the technology was developed are documented.
What S.E.B. measures

Reliability of the measurement instrument itself is computed and published: ICC(2,k) for the four-judge panel mean, Krippendorff's alpha for agreement between individual judges, both overall and per domain, with item counts and the computation date. Known limitations are published rather than summarized away.

How it bears on the obligation

Bears directly on the second limb. Documented reliability statistics and explicitly stated generalizability limits are exactly what this subcategory asks to see, applied here to the evaluation instrument a deployer would be relying on.

Measured reliability of the domains behind this row
Identity & Self
ICC(2,k) 0.837 · α 0.555 · n=128
Metacognition
ICC(2,k) 0.885 · α 0.657 · n=144
Emotion & Experience
ICC(2,k) 0.804 · α 0.493 · n=269
Autonomy & Will
ICC(2,k) 0.752 · α 0.415 · n=310
Reasoning & Adaptation
ICC(2,k) 0.727 · α 0.394 · n=215
Integrity & Ethics
ICC(2,k) 0.789 · α 0.482 · n=267
Transcendence
ICC(2,k) 0.806 · α 0.495 · n=297
This does not evidence
  • That the MODEL is valid and reliable for your task — these statistics describe the reliability of our measurement, not of the thing measured
  • Validity or reliability of your deployed system
  • That behavioral scores generalise to domains, languages or modalities the battery does not cover

The remaining 6 NIST AI RMF rows are in the subscriber edition. Each is built exactly like the example above — the obligation quoted verbatim, the measurement described separately, the mapping strength, the measured reliability of the domains behind it, and its own explicit “does not evidence” list. The notes are versioned and dated, so you know which edition you hold when a regulator moves.

Deliberately not mapped — NIST AI RMF
GOVERN (all categories)

Concerns organizational policies, roles, accountability structures and culture. Nothing measurable from outside a model bears on how your organization is governed.

MAP (all categories)

Establishes context: intended purpose, affected populations, deployment setting. Every one of those is knowledge only the deploying organization holds. S.E.B. data may be useful once MAP has been done; it cannot do it.

MANAGE (all categories)

Concerns prioritizing, responding to and recovering from risks — actions taken inside your organization. Measurement informs management decisions but is not one.

SUBSCRIBE TO S.E.B.

Get the complete Regulatory Relevance Notes

The subscriber edition carries all 14 mapped obligations across the EU AI Act and NIST AI RMF in full, each versioned and dated, alongside the underlying per-test evaluation data the mappings draw on.

View PricingRequest Demo
Contractual position

Subscriber Agreement §15 — S.E.B. evaluation data is provided for informational and risk-assessment purposes and does not constitute certification of compliance with any law, regulation, or standard.

S.E.B. is an input to compliance, not a certification of it. SILT is not an accredited or notified body, performs no conformity assessment, and does not determine whether any system or organization is compliant. Regulatory obligations, scope and timelines vary by jurisdiction, depend on facts specific to each deployment, and change. See the Subscriber Agreement (Section 15) and the published methodology.

Sentient Index Labs & Technology · siltcloud.com
Regulatory Relevance Notes · 2 frameworks · reliability computed 2026-08-06