What would stop us being bought?

Independence

The short answer is that you should not take it on faith, and we would rather not be asked to be. This company exists because self-reported claims about AI systems are unreliable, and it would be strange to build that instrument and then ask you to accept our own self-report about ourselves.

So below are five claims about how this company and its instrument are structured — not what we value, which costs nothing to say. Each one is paired with what you can do to catch us in it. Where a claim could not be given a check, it is not on this page.

We do not build, deploy or invest in AI models, and we accept no funding, sponsorship or strategic investment from any AI model vendor.

This is not a statement of values. It is the reason a rating here cannot be bought: there is no commercial relationship to jeopardise by publishing a bad one, and no equity position that a downgrade would cost us.

How you check it. A bought rater never downgrades anyone. So do not ask us who funds us — that is our word again. Go to the front page and look for the declines: we publish, by name, which models got worse since we last measured them, on the same screen as the ones that improved. A lab paid by the vendors it rates does not print a named vendor's score falling. If you ever find that we have quietly stopped publishing the downward half of that chart, you have caught us, and you will not need our cooperation to do it.

Every transcript is scored by independent judges who are not told which system produced it, and who cannot see each other's scores.

A judge who knows the vendor can flatter it, and a judge who sees a colleague's score converges on it. Blinding removes both without anyone having to be virtuous.

How you check it. We publish inter-rater reliability rather than summarising it, including the coefficient we originally reported wrongly and corrected. Disagreement between judges is visible in the data. A panel that never disagreed would be evidence of collusion or of a rubric doing the scoring, and ours disagrees.

No judge ever grades a model made by its own company. When the model under test comes from a judge's company, an independent judge from a company with no seat on the panel takes that seat for that result.

Blinding hides which system wrote a transcript, but not its style, and a model can recognise its own. The trimmed mean only removes a self-favouring judge when its score happens to be the most extreme, so relying on it would be relying on the bias being clumsy.

How you check it. Every stored result records which model held each seat, and a stepped-down judge's blind grade is kept beside it, uncounted. When we adopted the rule we re-graded every stored result that had broken it, withdrew from scoring the few that could not be re-graded, and published the correction on the methodology page with both counts. If you find a result graded by its own model, that is a defect in our record, and you found it without our cooperation.

Every model meets the same battery, administered the same way, and results are computed from raw scores with no editorial override.

Comparison means nothing if the instrument moves between subjects. A bespoke evaluation is a testimonial with statistics on it.

How you check it. Every model's coverage is printed beside its result — "10 of 11 tests", not a tidy percentage. A vendor given a shorter or gentler run would show up there as a smaller denominator, which is why the denominator is on the page at all. Compare them to each other. And where a cell could not be scored we report it unscored rather than as a zero, because a zero would flatter the models that refused and punish the ones that tried.

A vendor cannot pay for a favourable rating, for early sight of results, or to be left out of an evaluation.

Exclusion is the quiet version of the same purchase. A register that lets you buy your way off it is an advertisement.

How you check it. Look for absences. If a major vendor is missing from our roster, we state the reason — a provider dropping a model, a retirement, an account we cannot get. If you find a gap we have not explained, that is the question to put to us, and it is the one we would least like to be unable to answer.

What is deliberately not on this page

There is no security claim here. No padlock, no encryption boast, no badge. That is not because we are careless with data — it is because every company says it, so it persuades nobody who was not already convinced, while inviting the one audit that finds your worst day. We have had one. Independence and method are the opposite kind of claim: they are unusual, and they get stronger the harder somebody checks them, which is why they are the two we are willing to stake the page on.

You will also not find a customer logo wall. Where we ever provide complimentary access in exchange for permission to name someone, that arrangement will be disclosed inline wherever the name appears, because an undisclosed one reads as a paid endorsement and would quietly undermine everything above.

And the check we would most like you to run. Do not take this page on trust. Put it into whichever AI system you already use and ask it the unkind question: do the confidence intervals match the sample sizes, and does anything here claim more than it measured? We would rather you asked it than asked us. A page written to survive that reading is a different document from one written to be believed.

If you would rather hand it something structured, everything we claim is published in machine-readable form at silt-seb.com/llms.txt — the same figures, the same denominators, the same dates, and the same list of things we do not claim. It exists so the page can be checked against the data rather than read for tone.