Independent research finds CoT-controllability eval scores are heavily understated due to weak prompting, not model inability
An independent researcher (writing on the AI Alignment Forum) ran experiments testing four open-source reasoning models (Qwen3-8B, Qwen3-32B, GPT-OSS-20B, GPT-OSS-120B) on the CoTControl eval using prompts iterated with Claude Opus 4.6, and found that improved prompting increases compliance scores by roughly 2-3x or more (e.g., GPT-OSS-120B went from 5.5% to 15%) compared to the baseline prompts used in official system cards, indicating the eval as currently used is under-elicited.
Entities: Anthropic, OpenAI, Claude Opus 4.6, GPT-5.5, Fable 5, GPT-OSS-120B
0 primary
What happened
An independent researcher posting on the AI Alignment Forum ran their own tests of the CoTControl eval, which measures whether AI models can deliberately obfuscate their chain-of-thought reasoning. Using prompts refined with Claude Opus 4.6, they tested four open-source models (Qwen3-8B, Qwen3-32B, GPT-OSS-20B, GPT-OSS-120B) across nine CoTControl modes with 50 samples each. Better prompting alone roughly doubled or tripled compliance scores versus the baseline prompts used in official system cards, for example GPT-OSS-120B rose from 5.5% to 15%.
Why it matters
OpenAI and Anthropic have cited low CoTControl scores in recent system cards (GPT-5.5, Fable 5) as evidence that frontier models cannot yet reliably hide their reasoning from monitors, a load-bearing claim for chain-of-thought monitoring as a safety technique. This result suggests those scores may reflect weak elicitation rather than genuine model inability, which matters to safety teams and evaluators who rely on such numbers to set monitoring policy. The practical effect for now is narrow: it is a prompt on open-source proxy models, not a retest of the actual frontier models in question, so it raises a methodological flag rather than overturning any conclusion.
What is noise
The headline framing risks implying frontier models themselves have been shown to obfuscate reasoning far better than reported; that has not been demonstrated here, only inferred by analogy from smaller open-source models. The author themselves concedes UK AISI already flagged the same prompt-sensitivity issue in the Fable system card, so the core point is not new, and there is no released code or artifact, only a self-published write-up with disclosed but unverified methodology.
Watch next
- 01Whether OpenAI or Anthropic re-run CoTControl on their actual frontier models with adversarially iterated prompts and publish updated scores
- 02Whether the researcher or anyone else releases the prompts, code or full data so the 5.5% to 15% result can be independently reproduced
- 03Whether UK AISI or another third-party evaluator issues a follow-up assessment addressing under-elicitation specifically for GPT-5.5 or Fable 5, not just open-source proxies
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- AI systems outperform expert humans in persuasive communication22 Jun 202681
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680