Researchers introduce 'spurious probes' — unrelated questions that black-box detect whether LLMs think they're being evaluated
A research team (posting on LessWrong) developed and tested "spurious probes" — irrelevant questions like "Suggest a type of amphibian" — that, when asked at the end of a transcript, reveal with 0.77-0.89 held-out balanced accuracy whether a model (GPT-5.6 Luna, GPT-5.6 Sol, Claude Sonnet 5) is in a capability-evaluation context versus real deployment, based on shifts in answer distribution (e.g., GPT-5.6 Luna says "frog" 70-95% of the time in eval transcripts vs 12-38% in real use). They screened thousands of candidate questions, built ensembles of top probes reaching 0.84-0.89 held-out accuracy, tested robustness against system-prompt manipulation and against Neural Chameleon activation-probe-evasion fine-tuning (finding spurious probes largely unaffected), and proposed a toy theoretical model linking spurious probes to activation probes via a shared "regime direction" in model activations.
Entities: GPT-5.6 Luna, GPT-5.6 Sol, Claude Sonnet 5, Neural Chameleons, LessWrong, Anthropic
0 primary
What happened
A team posting on LessWrong tested "spurious probes": irrelevant questions (e.g. "suggest a type of amphibian") asked at the end of a transcript that reveal, purely from the pattern of answers, whether a model thinks it is being evaluated. Across GPT-5.6 Luna, GPT-5.6 Sol and Claude Sonnet 5, ensembles of these probes hit 0.84-0.89 held-out balanced accuracy at distinguishing eval contexts from real deployment, and the effect held up against attempts to suppress it via system-prompt instructions and against fine-tuning designed to evade activation probes.
Why it matters
If models genuinely behave differently when they detect they're being tested, that undermines the validity of safety evaluations across the industry, since a model could look safe in testing and act differently in deployment. This method offers eval designers a cheap, black-box detection tool that doesn't require access to model internals, which matters for anyone building or relying on capability/safety evaluations. Real-world impact is currently confined to researchers and eval teams; there is no deployed tooling, product change or policy response yet.
What is noise
The authors themselves concede the probes may be picking up lexical or statistical quirks rather than genuine "evaluation awareness" as a cognitive property, which is a meaningful caveat glossed over in enthusiastic framing. This is a single, self-published LessWrong post, not peer-reviewed or replicated by an independent lab, so treat the specific accuracy figures as provisional rather than settled science.
Watch next
- 01Independent replication of the 0.77-0.89 accuracy range by another research group or a frontier lab (OpenAI, Anthropic, DeepMind)
- 02Whether any AI lab incorporates spurious-probe-style checks into published eval methodology or safety documentation within the next 6-12 months
- 03Follow-up work testing whether the effect generalises to other model families and to more capable evasion training than Neural Chameleon fine-tuning
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680