Access to a language model almost eliminated people's willingness to say "I don't know" in five experiments with 3,132 participants. When volunteers could ask AI for advice, withheld judgments fell from 36 percent and 44 percent in the control groups to 6 percent and 3 percent. The advice came from Step 3.5 Flash on movie-detail questions where it was almost always wrong, so the shift was not sensible delegation to a reliable tool. For companies, the result matters because AI access can increase answer volume while reducing accuracy.
How researchers tested reliance on AI advice
Studies 1a and 1b let participants decide whether to request AI help on fine visual details from films, such as the color of a team uniform in Bend It Like Beckham. The authors chose this material because such details rarely appear in online text and provoke hallucinations. GPT-5.5, Claude 4.6 Sonnet and Gemini 3.5 Flash handled most other questions correctly but missed some harder ones. The control groups without AI withheld judgment on more than a third of items. With AI available, abstention nearly disappeared even though the advice was mostly incorrect.
Studies 2 through 4 added money to test whether stakes restore caution. Participants earned 10 cents for each correct answer, lost 10 cents for each wrong answer and received nothing for saying "I don't know". The researchers expected incentives to raise abstention and AI access to weaken that effect, but none of the three studies showed a statistically significant interaction. Incentives worked on their own: in Study 3, participants requested AI help 4.53 times out of six with incentives versus 5.27 times without them. Judgment suspension rose slightly with payment yet stayed far below the no-AI control.
The authors link the pattern to "Epistemia", acceptance of AI answers because they sound convincing rather than checked. A language model always produces an answer and never pauses when it lacks knowledge, and users who delegate judgment may copy that lack of restraint. The behavior runs against classic advice-use research, where people normally shift only about a third of the way toward an adviser. Related work points the same way: a Swiss Business School study with 666 participants found a strong negative correlation between AI use and critical thinking, and Microsoft research found that AI use places heavy demands on metacognition, with higher education as a protective factor.
What this means for companies using AI
For teams that use AI in sales, support and internal analysis, the cost appears as confident error rather than silence. In Study 2 without incentives, confidence with AI reached 75.9 points out of 100 against 29.6 without AI, while correct answers fell from 27.6 percent to 10.0 percent. Across all studies without incentives, AI users answered more questions but were right in 9.2 percent of cases versus 27.5 percent without AI, about a third as often. The authors conclude that AI access turned some answers that would have been correct into errors, which raises review load as plausible drafts multiply.
Payment for accuracy helped but did not restore the control level, and the researchers call the gain modest. Whether larger sums or reputation-based stakes would close the gap is unknown. Study 4 also showed the effect when AI answers appeared automatically, as with search summaries and writing assistants: without incentives, abstention fell from 35 percent to 1 percent, and with incentives from about 39 percent to 7 percent. The open question is whether the same deference holds beyond movie trivia. Until that is tested, the finding does not mean every AI assistant harms judgment, only that unchecked availability suppresses the pause that protects accuracy.
A practical marker to watch is the abstention rate inside pilot teams when AI help and a small bonus for accuracy are both present. If withheld judgments stay near 7 percent instead of returning toward 35 to 39 percent, the deference effect is holding. A second marker is replication beyond trivia on work tasks with verifiable answers.
