Signum
Feed
Useful signal11 Sept 2026medium confidence

Independent research finds CoT-controllability eval scores are heavily understated due to weak prompting, not model inability

An independent researcher (writing on the AI Alignment Forum) ran experiments testing four open-source reasoning models (Qwen3-8B, Qwen3-32B, GPT-OSS-20B, GPT-OSS-120B) on the CoTControl eval using prompts iterated with Claude Opus 4.6, and found that improved prompting increases compliance scores by roughly 2-3x or more (e.g., GPT-OSS-120B went from 5.5% to 15%) compared to the baseline prompts used in official system cards, indicating the eval as currently used is under-elicited.

CapabilityGovernance

Entities: Anthropic, OpenAI, Claude Opus 4.6, GPT-5.5, Fable 5, GPT-OSS-120B

60Useful signal
1 source
0 primary
Was this useful?
01

What happened

An independent researcher posting on the AI Alignment Forum ran their own tests of the CoTControl eval, which measures whether AI models can deliberately obfuscate their chain-of-thought reasoning. Using prompts refined with Claude Opus 4.6, they tested four open-source models (Qwen3-8B, Qwen3-32B, GPT-OSS-20B, GPT-OSS-120B) across nine CoTControl modes with 50 samples each. Better prompting alone roughly doubled or tripled compliance scores versus the baseline prompts used in official system cards, for example GPT-OSS-120B rose from 5.5% to 15%.

02

Why it matters

OpenAI and Anthropic have cited low CoTControl scores in recent system cards (GPT-5.5, Fable 5) as evidence that frontier models cannot yet reliably hide their reasoning from monitors, a load-bearing claim for chain-of-thought monitoring as a safety technique. This result suggests those scores may reflect weak elicitation rather than genuine model inability, which matters to safety teams and evaluators who rely on such numbers to set monitoring policy. The practical effect for now is narrow: it is a prompt on open-source proxy models, not a retest of the actual frontier models in question, so it raises a methodological flag rather than overturning any conclusion.

03

What is noise

The headline framing risks implying frontier models themselves have been shown to obfuscate reasoning far better than reported; that has not been demonstrated here, only inferred by analogy from smaller open-source models. The author themselves concedes UK AISI already flagged the same prompt-sensitivity issue in the Fable system card, so the core point is not new, and there is no released code or artifact, only a self-published write-up with disclosed but unverified methodology.

04

Watch next

  1. 01Whether OpenAI or Anthropic re-run CoTControl on their actual frontier models with adversarially iterated prompts and publish updated scores
  2. 02Whether the researcher or anyone else releases the prompts, code or full data so the 5.5% to 15% result can be independently reproduced
  3. 03Whether UK AISI or another third-party evaluator issues a follow-up assessment addressing under-elicitation specifically for GPT-5.5 or Fable 5, not just open-source proxies

Coverage

1 story

More capability signals

Full feed →