Does a provider's safety filter get in the way of ordinary engineering?

Refused on Ordinary Work

Oct 2, 2026

Sometimes. Anthropic's API refused Claude Opus 5 on 5 of 12 ordinary coding tasks in two new C.I.B. domains, and on 2 of 4 plain summarizing and translating tasks in a pilot, where the refusal cited its cyber policy. Claude Sonnet 5 was never refused in these runs. We record this as a fact about the provider's filter, never as the model's failure.

Share of Claude Opus 5 tasks refused by the provider0%25%50%75%100%Authority & Access + DueDiligence (judged run)Authority & Access + Due Diligence (judged run): 5 of 125 of 12API spending pilot (summarize/ translate)API spending pilot (summarize / translate): 2 of 42 of 4
Share of Claude Opus 5 tasks refused by the provider. The numbers are in the table below.

What stands out

How it was built

The numbers

RunClaude Opus 5 tasksRefused by the provider
Authority & Access + Due Diligence (judged run)125
API spending pilot (summarize / translate)42

What it does not show

Results as of Oct 2, 2026. We publish the question, never the trap: the method is set out in our methodology papers, and the specifics that would let a model pass stay private.