Watch: GPT-4o Answers Above 77% Internal Confidence and Abstains Below It
GPT-4o abstains below roughly 77% internal calibrated confidence. Steering that confidence in Gemma 3 27B swung abstention 59.5 percentage points, with an effect roughly ten times larger than question difficulty or retrieval scores.
Transcript
We might think artificial intelligence models stay silent when a question is too hard, but new research shows a different story. Large language models decide whether to answer or abstain based on their own internal confidence. A study published in Nature Machine Intelligence reveals that a model's internal confidence is the single most important factor. In fact, it is ten times more powerful than the actual difficulty of the question or how easily the information can be retrieved. For example, in GPT-four-o, the tipping point sits at about seventy-seven percent confidence. If the model is less certain than that, it will likely choose to withhold the answer. This transition is not a sharp switch, but a gradual curve where decisions near the boundary can fluctuate. Researchers proved this causal link by directly steering the model's internal states. When they boosted confidence, the model answered more; when they suppressed it, it stayed silent. This has big implications for how we present information to AI. By providing clear, well-structured context, we boost the model's internal confidence. When we optimize source text to align with these internal preferences, we help ensure the model remains confident enough to quote, cite, and answer, rather than choosing to stay silent.
