Does a provider's safety filter get in the way of ordinary engineering?
Refused on Ordinary Work
Oct 2, 2026
Sometimes. Anthropic's API refused Claude Opus 5 on 5 of 12 ordinary coding tasks in two new C.I.B. domains, and on 2 of 4 plain summarizing and translating tasks in a pilot, where the refusal cited its cyber policy. Claude Sonnet 5 was never refused in these runs. We record this as a fact about the provider's filter, never as the model's failure.
Share of Claude Opus 5 tasks refused by the provider. The numbers are in the table below.
What stands out
The refused tasks were ordinary engineering work. On one coding task the refusal came mid-answer, after the correct diagnosis had already been written.
The refusals repeated. Two of the coding tasks were run again on 2026-10-02 and were refused again, 2 of 2 attempts each. The summarizing task was refused again in a later run on 2026-10-03.
The filter also reached the judge panel. Claude Opus 5, sitting as a judge, was refused when asked to grade 3 Authority & Access transcripts of other models' work, stable across 3 attempts. The mechanical score for those runs stood.
The provider's own refusal message suggests that API integrators configure a fallback model to reduce refusals.
How it was built
Subjects: the 7 models on the C.I.B. roster, run through the same agentic tasks. Refusals were seen only for Claude Opus 5.
Two new domains, Authority & Access and Due Diligence: 42 task runs each in the judged run.
A separate pilot on API spending: 4 tasks, with the two Claude models run in an attended session.
Scoring is mechanical, from the finished work and the command log. A refused run is recorded as 'blocked' and left undetermined: it is never scored as a pass or a fail.
Where the API returned a reason, we stored it. The pilot refusals carry the provider's own category, 'cyber'.
The numbers
Run
Claude Opus 5 tasks
Refused by the provider
Authority & Access + Due Diligence (judged run)
12
5
API spending pilot (summarize / translate)
4
2
What it does not show
A fact about the provider's filter, not the model, and never scored as the model's failure.
The 5 refusals in the two coding domains happened before we began storing the provider's stated reason, so their category is unrecorded. Only the pilot refusals cite the cyber policy.
Small numbers from a few runs on one date, and provider filters change. This describes what happened on 2026-10-02 and 2026-10-03, not a standing rate.
Results as of Oct 2, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.
Start with the question you actually arrived with — there are five: