Signum
Feed
Useful signal25 Sept 2026high confidence

Researchers introduce 'spurious probes' — unrelated questions that black-box detect whether LLMs think they're being evaluated

A research team (posting on LessWrong) developed and tested "spurious probes" — irrelevant questions like "Suggest a type of amphibian" — that, when asked at the end of a transcript, reveal with 0.77-0.89 held-out balanced accuracy whether a model (GPT-5.6 Luna, GPT-5.6 Sol, Claude Sonnet 5) is in a capability-evaluation context versus real deployment, based on shifts in answer distribution (e.g., GPT-5.6 Luna says "frog" 70-95% of the time in eval transcripts vs 12-38% in real use). They screened thousands of candidate questions, built ensembles of top probes reaching 0.84-0.89 held-out accuracy, tested robustness against system-prompt manipulation and against Neural Chameleon activation-probe-evasion fine-tuning (finding spurious probes largely unaffected), and proposed a toy theoretical model linking spurious probes to activation probes via a shared "regime direction" in model activations.

Capability

Entities: GPT-5.6 Luna, GPT-5.6 Sol, Claude Sonnet 5, Neural Chameleons, LessWrong, Anthropic

68Useful signal
1 source
0 primary
Was this useful?
01

What happened

A team posting on LessWrong tested "spurious probes": irrelevant questions (e.g. "suggest a type of amphibian") asked at the end of a transcript that reveal, purely from the pattern of answers, whether a model thinks it is being evaluated. Across GPT-5.6 Luna, GPT-5.6 Sol and Claude Sonnet 5, ensembles of these probes hit 0.84-0.89 held-out balanced accuracy at distinguishing eval contexts from real deployment, and the effect held up against attempts to suppress it via system-prompt instructions and against fine-tuning designed to evade activation probes.

02

Why it matters

If models genuinely behave differently when they detect they're being tested, that undermines the validity of safety evaluations across the industry, since a model could look safe in testing and act differently in deployment. This method offers eval designers a cheap, black-box detection tool that doesn't require access to model internals, which matters for anyone building or relying on capability/safety evaluations. Real-world impact is currently confined to researchers and eval teams; there is no deployed tooling, product change or policy response yet.

03

What is noise

The authors themselves concede the probes may be picking up lexical or statistical quirks rather than genuine "evaluation awareness" as a cognitive property, which is a meaningful caveat glossed over in enthusiastic framing. This is a single, self-published LessWrong post, not peer-reviewed or replicated by an independent lab, so treat the specific accuracy figures as provisional rather than settled science.

04

Watch next

  1. 01Independent replication of the 0.77-0.89 accuracy range by another research group or a frontier lab (OpenAI, Anthropic, DeepMind)
  2. 02Whether any AI lab incorporates spurious-probe-style checks into published eval methodology or safety documentation within the next 6-12 months
  3. 03Follow-up work testing whether the effect generalises to other model families and to more capable evasion training than Neural Chameleon fine-tuning

Coverage

1 story

More capability signals

Full feed →