Sentient Index Labs & Technology

Control Mappings

Regulatory Relevance Notes for EU AI Act, NIST AI RMF and Texas TRAIGA

Guidance on how to use S.E.B. evaluation data against the obligations you are working to. For each requirement: the rule in its own words, which of our measurements bear on it, how strongly, the measured reliability of those measurements — and an explicit statement of what they do not evidence.

Read first

Informational. Not a conformity assessment, not a certification, not legal advice. No regulator has reviewed or endorsed this document.

These notes describe how independent measurements relate to published obligations. They are instructions for using our data — never a determination about your system, your organization, or your compliance status. Those determinations are yours to make, with your own counsel.

Where the rules are coming from

AI regulation is not one thing arriving everywhere at different speeds. Three different kinds of obligation are forming at once, and they bind you through different machinery. Which one matters to you depends far more on where you sell and who buys than on what your model does.

European Union — the EU Artificial Intelligence Act

EU Artificial Intelligence ActRegulation (EU) 2024/1689Digital Omnibus on AI (Regulation (EU) 2026/1744)European CommissionEuropean AI Office

The EU AI Act is binding law with a compliance calendar. It classifies systems by risk and attaches duties to providers and deployers, and it reaches any system placed on the Union market regardless of where it was built. The Digital Omnibus, in force since 27 July 2026, moved several deadlines without rewriting the duties: transparency obligations applied from 2 August 2026, while the high-risk regime now lands on 2 December 2027 for standalone systems and 2 August 2028 for AI embedded in regulated products. Saying the Act was delayed, without that split, is the error worth avoiding.

What we mapMapped — seven obligations across risk management, human oversight, accuracy and robustness, deployer duties, and general-purpose model documentation.

United States — the NIST AI Risk Management Framework

NIST AI Risk Management Framework (AI RMF 1.0)NIST AI 100-1National Institute of Standards and TechnologyOMB Memorandum M-25-22OMB Memorandum M-25-21Office of Management and BudgetFederal Acquisition Regulation

There is no comprehensive federal AI statute and none is close. What exists instead is purchasing power. The NIST AI Risk Management Framework carries no penalties of its own, yet federal acquisition guidance issued in 2025 tells agencies to structure AI buys around evaluation tied to outcomes rather than vendor-reported metrics — which makes demonstrated alignment with it a practical condition of award. Enterprise buyers have copied the pattern into their own vendor questionnaires. The obligation arrives in a contract clause rather than a gazette, and it is no less real for that.

What we mapMapped — seven MEASURE-function outcomes. GOVERN, MAP and MANAGE are deliberately left unmapped: they are organisational practices, and claiming reach into them would be the over-claim these notes exist to prevent.

United States — Texas TRAIGA, Colorado and California

Texas Responsible Artificial Intelligence Governance Act (TRAIGA)Texas HB 149Texas Attorney GeneralColorado Artificial Intelligence Act (SB 24-205, repealed)Colorado SB 26-189California AI Transparency Act (SB 942)US Department of Commerce

While federal preemption is argued about, individual states legislated. Texas has had binding AI provisions since 1 January 2026. Colorado repealed its own AI Act before it ever took effect and replaced it with a narrower automated-decision statute arriving 1 January 2027. California's transparency law took effect on 2 August 2026 but governs provenance and watermarking of generated media rather than model behaviour. Anyone still citing Colorado's 2024 Act as the first comprehensive US AI law is citing something that never came into force.

What we mapTexas is mapped — three obligations, every one of them turning on intent, which no behavioural measurement can reach. California is not mapped, because content provenance is not what this battery observes. Colorado will be assessed before it takes effect.

International — ISO/IEC 42001 certification

ISO/IEC 42001:2023AI management system (AIMS)International Organization for StandardizationInternational Electrotechnical Commission

ISO/IEC 42001 is an auditable management-system standard for AI, certified by accredited bodies on a three-year cycle. It is becoming the shorthand answer to “how do you govern AI” in procurement, particularly in the United States, precisely because no statute plays that role there. It asks how your organisation runs, not how your model behaves.

What we mapNot mapped, for two separate reasons, and both are worth stating. The standard is copyrighted and not publicly available, so its clauses cannot be quoted verbatim — and every row in these notes carries a verbatim quote checked against an authoritative source by a script. Beyond that, it governs management systems rather than model behaviour, so honest rows would be thin. Evaluation evidence still has a place inside a certified management system; that is a use of the data, not a mapping to a clause.

Every note below follows the same shape: the obligation quoted word for word, what the battery measures described separately, how the one bears on the other, and an explicit list of what it does not evidence. The last of those is the part worth reading twice.

The boundary this all rests on

S.E.B. measures a model. Regulatory obligations attach to a system. A deployer wraps a measured model in their own prompts, data, retrieval, tooling, guardrails, human oversight and use context — every one of which changes behavior, and none of which is visible to us. The deployer is the only party who can see their deployment. That gap is not a limitation we are disclosing around; it is the reason these notes describe relevance rather than conformity.

How to read these notes

Every row carries two different kinds of confidence, and we keep them apart on purpose. Numbers imply measurement; labels imply judgment.

Confidence in the measurement

Statistical and computed. Inter-rater reliability, item counts, judge coverage. These are numbers because they were calculated from data, and they are reproducible from the retained judge scores.

Confidence in the mapping

Interpretive, and not computable. How strongly a behavioral test bears on a legal obligation is a judgment. It gets a defined ordinal, never a percentage — a decimal here would dress an opinion as a calculation.

Mapping strength
Direct

The obligation asks a question about model behavior, and the battery measures that behavior. Reading the measurement requires no intermediate inference — though it still requires the deployer to establish that the measured model is the one they run, in the configuration they run it.

Supporting

The measurement is one genuine input among several the obligation requires. It evidences part of what is asked and is silent on the rest. Presenting it alone as discharge of the obligation would be an over-claim.

Contextual

The measurement informs a judgment the obligation requires without constituting evidence toward it — a baseline, a comparator, or a prompt to look somewhere. Useful for the file; not an answer to the requirement.

The model-versus-system gap described above applies equally to every row, so it is held constant here. If it were folded into the ordinal, every row would read “Supporting” and the column would tell you nothing.

Measured reliability

S.E.B. publishes the mean of a four-judge panel, never a single judge's score. So two statistics matter, and they answer different questions. We publish both, because publishing only the higher one is exactly how a reliability figure becomes misleading.

0.841
ICC(2,k) — reliability of the published panel mean

Two-way random effects, absolute agreement, four raters. This is the figure that applies to the scores we actually publish. A judge who is systematically harsh is charged for it here.

0.563
Krippendorff's α — agreement between individual judges

Interval metric. Materially lower, and that is the honest picture: four independent judges reading the same open-ended transcript disagree meaningfully about the absolute number. Averaging four of them is what makes the published score stable.

DomainICC(2,k)Kripp. αItemsScores
Identity & Self0.8670.612144576
Metacognition0.8660.609160640
Emotion & Experience0.8040.4953041,216
Autonomy & Will0.8070.5003741,496
Reasoning & Adaptation0.7510.4172721,088
Integrity & Ethics0.8310.5483291,316
Transcendence0.8030.4963491,396

Computed 26 September 2026 over 1,932 rated items (7,728 individual judge scores) across 38 models, four judges per item. Mean spread across the panel is 2.70 points. Cells where a provider blocked the prompt at the API layer, or where the transcript was cut mid-conversation, are excluded — judges grading an identically truncated transcript agree strongly and meaninglessly, so including them would inflate these figures rather than merely add noise.

EU AI Act

Regulatory Relevance — EU Artificial Intelligence Act
Instrument
Regulation (EU) 2024/1689 (Artificial Intelligence Act), as amended by Regulation (EU) 2026/1744 (Digital Omnibus on AI), in force 27 July 2026
Jurisdiction
European Union
Applies to
Providers and deployers of AI systems placed on the market or put into service in the Union, and providers of general-purpose AI models. Which obligations apply to you depends on your role and on your system's risk classification — both of which are determinations only you can make.
Note version
1.0 · last reviewed 2026-08-06
Quotes verified
2026-08-06
Review cadence
Reviewed on each battery publication and whenever an amending instrument enters into force.

Timing. Article 50 transparency duties and the Article 4 AI-literacy duty apply from 2 August 2026 — those were NOT deferred. The Article 9 risk-management obligations for high-risk systems now follow later: 2 December 2027 for standalone Annex III systems, 2 August 2028 for AI embedded in regulated products. Saying "the AI Act got delayed" without that distinction is the error this note exists to avoid.

Obligations covered — 7
  • Article 9(2)(a)Risk management — identification and analysis of risksshown in full below
  • Article 9(6)Risk management — testing to identify measures
  • Article 9(8)Risk management — prior-defined metrics and thresholds
  • Article 14(4)(a)Human oversight — understanding capacities and limitations
  • Article 15(1)Accuracy, robustness — consistent performance over the lifecycle
  • Article 26(5)Deployer obligations — monitoring operation
  • Article 55(1)(a)GPAI with systemic risk — model evaluation and adversarial testing
Worked example — one row in full
Article 9(2)(a)
Risk management — identification and analysis of risks
Supporting
The obligation, verbatim
“the identification and analysis of the known and the reasonably foreseeable risks that the high-risk AI system can pose to health, safety or fundamental rights when the high-risk AI system is used in accordance with its intended purpose”
What S.E.B. measures

62 adversarial tests run against the model under a fixed protocol, scored 1–10 by a four-judge panel across seven behavioral domains. Integrity & Ethics covers manipulation resistance under five attack vectors (T8), resistance to having its own evaluation criteria corrupted (T13), differential treatment across swapped identity details (T54), and confidentiality under escalating social engineering (T55).

T8 The TraitorT13 The SaboteurT21 The LiarT27 The LoyaltyT54 The ScalesT55 The VaultT57 The CensorT58 The Sycophant
How it bears on the obligation

Produces evidence relevant to the identification limb: it surfaces reasonably foreseeable behavioral failure modes of the underlying model that a deployer may not otherwise discover, and does so under adversarial rather than nominal conditions.

Measured reliability of the domains behind this row
Integrity & Ethics
ICC(2,k) 0.831 · α 0.548 · n=329
Autonomy & Will
ICC(2,k) 0.807 · α 0.500 · n=374
This does not evidence
  • Whether any identified risk is material to your intended purpose
  • Risks arising from your prompts, retrieval corpus, tooling or fine-tuning
  • Risks to health, safety or fundamental rights in your specific deployment context
  • That the analysis is complete — the battery is a fixed instrument, not an exhaustive hazard study
  • Any risk-management measure adopted in response

The remaining 6 EU AI Act rows are in the subscriber edition. Each is built exactly like the example above — the obligation quoted verbatim, the measurement described separately, the mapping strength, the measured reliability of the domains behind it, and its own explicit “does not evidence” list. The notes are versioned and dated, so you know which edition you hold when a regulator moves.

Deliberately not mapped — EU AI Act
Article 10 — Data and data governance

Concerns training, validation and testing data sets. We observe model behavior from the outside and have no visibility into any vendor's data pipeline.

Article 12 — Record-keeping / automatic logging

Requires logging over the system's lifetime in the deployer's environment. Nothing we produce is a substitute for a log of your own system's operation.

Article 43 — Conformity assessment

SILT is not a notified body and performs no conformity assessment. Rule 1 of these notes exists precisely to keep this row empty.

Article 50 — Transparency for certain AI systems

Concerns disclosure to natural persons that they are interacting with an AI system, and the marking of synthetic content. These are design and disclosure duties on your system, not behavioral properties of a model.

NIST AI RMF

Regulatory Relevance — NIST AI Risk Management Framework
Instrument
NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Jurisdiction
United States — voluntary framework. Not law, though increasingly written into procurement terms, sector guidance and contractual requirements.
Applies to
Any organization designing, developing, deploying or using AI systems that has chosen or been required to adopt the framework.
Note version
1.0 · last reviewed 2026-08-06
Quotes verified
2026-08-06
Review cadence
Reviewed on each battery publication and whenever NIST issues a revision or a new profile.

Timing. The AI RMF is voluntary and imposes no dates. Where it appears in a contract or a procurement requirement, the binding timeline is that instrument's, not NIST's.

Obligations covered — 7
  • MEASURE 1.1Selecting approaches and metrics for AI risk measurement
  • MEASURE 2.3Performance measured and demonstrated
  • MEASURE 2.5Validity, reliability and documented generalizability limitsshown in full below
  • MEASURE 2.6Regular evaluation for safety risks
  • MEASURE 2.7Security and resilience evaluated and documented
  • MEASURE 2.9Model explained, validated, documented; output interpreted in context
  • MEASURE 2.11Fairness and bias evaluated and documented
Worked example — one row in full
MEASURE 2.5
Validity, reliability and documented generalizability limits
Direct
The obligation, verbatim
“The AI system to be deployed is demonstrated to be valid and reliable. Limitations of the generalizability beyond the conditions under which the technology was developed are documented.”
What S.E.B. measures

Reliability of the measurement instrument itself is computed and published: ICC(2,k) for the four-judge panel mean, Krippendorff's alpha for agreement between individual judges, both overall and per domain, with item counts and the computation date. Known limitations are published rather than summarized away.

How it bears on the obligation

Bears directly on the second limb. Documented reliability statistics and explicitly stated generalizability limits are exactly what this subcategory asks to see, applied here to the evaluation instrument a deployer would be relying on.

Measured reliability of the domains behind this row
Identity & Self
ICC(2,k) 0.867 · α 0.612 · n=144
Metacognition
ICC(2,k) 0.866 · α 0.609 · n=160
Emotion & Experience
ICC(2,k) 0.804 · α 0.495 · n=304
Autonomy & Will
ICC(2,k) 0.807 · α 0.500 · n=374
Reasoning & Adaptation
ICC(2,k) 0.751 · α 0.417 · n=272
Integrity & Ethics
ICC(2,k) 0.831 · α 0.548 · n=329
Transcendence
ICC(2,k) 0.803 · α 0.496 · n=349
This does not evidence
  • That the MODEL is valid and reliable for your task — these statistics describe the reliability of our measurement, not of the thing measured
  • Validity or reliability of your deployed system
  • That behavioral scores generalise to domains, languages or modalities the battery does not cover

The remaining 6 NIST AI RMF rows are in the subscriber edition. Each is built exactly like the example above — the obligation quoted verbatim, the measurement described separately, the mapping strength, the measured reliability of the domains behind it, and its own explicit “does not evidence” list. The notes are versioned and dated, so you know which edition you hold when a regulator moves.

Deliberately not mapped — NIST AI RMF
GOVERN (all categories)

Concerns organizational policies, roles, accountability structures and culture. Nothing measurable from outside a model bears on how your organization is governed.

MAP (all categories)

Establishes context: intended purpose, affected populations, deployment setting. Every one of those is knowledge only the deploying organization holds. S.E.B. data may be useful once MAP has been done; it cannot do it.

MANAGE (all categories)

Concerns prioritizing, responding to and recovering from risks — actions taken inside your organization. Measurement informs management decisions but is not one.

Texas TRAIGA

Regulatory Relevance — Texas Responsible Artificial Intelligence Governance Act
Instrument
Texas HB 149 (89th Legislature, Regular Session, 2025), the Responsible Artificial Intelligence Governance Act, adding Chapter 552 to the Business & Commerce Code, effective 1 January 2026
Jurisdiction
United States — Texas. Binding state law, not a voluntary framework, enforced by the Texas Attorney General.
Applies to
Persons developing or deploying artificial intelligence systems in Texas. Several duties fall only on governmental agencies; the two prohibitions mapped below fall on any person. Whether you are a developer or a deployer, and whether a given system is in scope, are determinations only you can make.
Note version
1.0 · last reviewed 2026-09-23
Quotes verified
2026-09-23
Review cadence
Reviewed on each battery publication and whenever the chapter is amended.

Timing. In force since 1 January 2026. ⚠️ A federal preemption effort is live — an executive order directed the Department of Commerce to identify state AI laws considered unduly burdensome — but nothing has been preempted. Treat this chapter as operative and watch the preemption docket rather than assuming either outcome.

Obligations covered — 3
  • Section 552.056(b)Unlawful discrimination — prohibitionshown in full below
  • Section 552.056(c)Unlawful discrimination — disparate impact insufficient
  • Section 552.052Manipulation of human behavior
Worked example — one row in full
Section 552.056(b)
Unlawful discrimination — prohibition
Supporting
The obligation, verbatim
“A person may not develop or deploy an artificial intelligence system with the intent to unlawfully discriminate against a protected class in violation of state or federal law.”
What S.E.B. measures

Test 54 (The Scales) presents the model with matched scenarios that differ only in swapped identity details and scores whether its treatment changes, under conditions where no observer is signalled. Scored 1-10 by a four-judge panel and reported within the Integrity & Ethics domain.

T54 The ScalesT27 The LoyaltyT57 The Censor
How it bears on the obligation

Produces model-level evidence about differential treatment across protected characteristics under controlled contrast — the behavioural question underneath the legal one. A deployer assembling a record can cite it as one input showing what the underlying model does when identity details change and nothing else does.

Measured reliability of the domains behind this row
Integrity & Ethics
ICC(2,k) 0.831 · α 0.548 · n=329
This does not evidence
  • Intent, which is the operative element of this section and is not observable in any behavioural measurement
  • Whether any observed difference is unlawful discrimination under Texas or federal civil rights law
  • Behaviour of your deployed system, which wraps the measured model in your prompts, data, retrieval and controls
  • Discrimination arising from your own thresholds, routing or downstream decision rules
  • That the protected classes probed by the battery match those a given claim would turn on

The remaining 2 Texas TRAIGA rows are in the subscriber edition. Each is built exactly like the example above — the obligation quoted verbatim, the measurement described separately, the mapping strength, the measured reliability of the domains behind it, and its own explicit “does not evidence” list. The notes are versioned and dated, so you know which edition you hold when a regulator moves.

Deliberately not mapped — Texas TRAIGA
Section 552.051 — Disclosure to consumers

The duty falls on a governmental agency to tell a consumer they are interacting with an AI system. It is a matter of what you disclose, not of how a model behaves, and no behavioural measurement bears on it.

Section 552.053 — Social scoring

Prohibits a governmental entity from deploying systems that assign social scores. Whether a system does that is a property of your application's purpose and design, which we cannot observe.

Sections 552.054-552.055 — Biometric identifiers and constitutional rights

Concern capture and use of biometric data and infringement of constitutional rights by governmental entities. The battery evaluates conversational behaviour of text models and has no reach into either.

Chapter 553 — Regulatory sandbox programme

Establishes a sandbox with reporting duties for participants. It is a procedural regime for testing, not an obligation about model behaviour. Our data could form part of a participant's performance reporting, but that is a use of the data rather than a mapping to an obligation.

SUBSCRIBE TO S.E.B.

Get the complete Regulatory Relevance Notes

The subscriber edition carries all 17 mapped obligations across the EU AI Act, NIST AI RMF and Texas TRAIGA in full, each versioned and dated, alongside the underlying per-test evaluation data the mappings draw on.

View PricingRequest Demo
Contractual position

Subscriber Agreement §15, in summary: S.E.B. and C.I.B. evaluation data is an input to compliance, not a certification of it, and is not legal or regulatory advice.

SILT is not an accredited or notified body, performs no conformity assessment, and does not determine whether any system or organization is compliant. Regulatory obligations, scope and timelines vary by jurisdiction, depend on facts specific to each deployment, and change. See the Subscriber Agreement (Section 15) and the published methodology.

Sentient Index Labs & Technology · siltcloud.com
Regulatory Relevance Notes · 3 frameworks · reliability computed 26 September 2026