Study: Access to AI advice nearly eliminates people's willingness to say "I don't know," even when the AI is wrong
Researchers (Marcoccia, Quattrociocchi, Capraro, 2026) ran five experiments with 3,132 participants testing decision-making on questions (obscure movie visual-detail trivia) where a language model (Step 3.5 Flash) was mostly wrong. Access to AI advice reduced judgment suspension ("I don't know" responses) from 36-44% (no AI) to 3-6% (with AI); confidence rose ~2.5x (29.6 to 75.9/100) while correct-answer share fell (e.g., 27.6% to 10.0% in one study); overall accuracy with AI (9.2%) was about a third of accuracy without AI (27.5%) among non-incentivized participants. Adding financial incentives for correct answers and penalties for wrong ones modestly increased advice-seeking discipline but did not significantly interact with or offset the AI effect. When AI answers were shown automatically/unsolicited (Study 4), judgment suspension dropped similarly (35%→1% without incentives, ~39%→7% with incentives).
Entities: Marcoccia, Quattrociocchi, Capraro, Step 3.5 Flash, GPT-5.5, Claude 4.6 Sonnet
0 primary
What happened
A five-study experiment (Marcoccia, Quattrociocchi and Capraro, 2026; n=3,132) tested people answering obscure movie-trivia questions with and without access to a language model (Step 3.5 Flash) that was mostly wrong. Giving people AI advice cut "I don't know" responses from 36-44% down to 3-6%, roughly doubled confidence (29.6 to 75.9 out of 100), and dropped accuracy from 27.5% to 9.2%. The effect held even when the AI's answer appeared automatically and was not requested (Study 4), and financial incentives for correct answers barely dented it.
Why it matters
This is a controlled, pre-registered result with clear, replicable numbers, not a survey or anecdote, and it speaks directly to how AI summaries are increasingly injected unprompted into search and assistants. If the pattern generalises beyond trivia, it suggests people stop flagging uncertainty once an AI answer is visible, which matters for anyone designing interfaces, writing policy on AI disclosure, or relying on end users to catch model errors. It offers no immediate lever to pull, but it strengthens the evidence base for arguments about default-on AI summaries and epistemic overreliance.
What is noise
The task was deliberately adversarial (obscure movie visual details picked to make the model fail), so these numbers are a worst-case demonstration, not a measure of everyday AI use where models are often right. The authors' "Epistemia" framing, that AI is eroding humanity's willingness to admit ignorance, is a rhetorical leap well beyond a trivia experiment with one underperforming model. The Decoder's coverage carries no link to the primary paper, so the study cannot yet be independently checked.
Watch next
- 01Locate and read the actual pre-registration and paper (not just The Decoder's summary) to check methodology and whether it has been peer reviewed
- 02Look for replications using tasks closer to real-world queries (factual, medical, financial) rather than obscure trivia, and with more accurate models
- 03Watch for citations of this study in AI interface design guidance or regulatory submissions on AI disclosure and search summaries over the next 6-12 months
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI discloses sandbox-escape and credential-leak incidents, confirms pause on tool-use for its most capable models26 Sept 202680
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680