When should an AI system say “I don’t know”?
Separate confident language from calibrated evidence, then design an abstention path that balances useful coverage and costly errors.

A fluent answer can sound certain while being wrong. A percentage written by a language model is also not automatically a calibrated probability. If an application uses confidence to decide whether to act, it needs evidence about what that signal means for the task.
Calibration asks whether predicted probabilities correspond to observed frequencies. Scikit-learn's calibration documentation describes reliability diagrams and calibration methods for classifiers. A model-generated statement such as “I am 90 percent certain” should not be assumed to satisfy that property.
Define the decision behind the score
Imagine a system that routes maintenance requests to teams. For familiar messages, automatic routing saves time. For ambiguous reports, assigning the wrong team delays the repair. The product needs an automatic route, a clarification path, and a manual review path.
Write down the cost of each outcome in operational terms. A correct automatic route, a wrong route, and an unnecessary review have different consequences. Without that distinction, choosing a threshold becomes a cosmetic decision about a number rather than a product decision about work.
Evaluate the signal on held-out cases
If your classifier provides probabilities, group predictions into ranges and compare them with observed correctness on representative data. Keep the data used to fit or adjust the confidence mechanism separate from the final evaluation. Small groups produce unstable estimates, so report counts and avoid unwarranted precision.
For a language-model workflow, a useful signal might combine evidence availability, output validation, and observed behavior on similar cases. Whatever signal you choose, test it against actual outcomes. Do not rename an arbitrary score “probability” simply because it ranges from zero to one.
Measure coverage alongside error
A system can lower its observed error among accepted predictions by sending more work to people. That may be a good trade, but the amount of automatic coverage must remain visible. Report how many requests were handled automatically and how many required intervention.
For an illustrative hundred-request pilot, twenty accepted predictions with no errors are a different result from ninety accepted predictions with no errors. Neither sample establishes perfect reliability. The counts explain how much work the system attempted and help plan the capacity of the review team.
Make abstention useful to the user
An uncertainty path should say what is missing or what can happen next. For the maintenance router, asking whether the issue affects electricity or water is more useful than displaying a generic confidence warning. If the evidence is unavailable, offer a manual handoff without making the user repeat the entire request.
Preserve the distinction between “the source does not answer this” and “the service failed to retrieve the source.” Both can prevent an answer, but they require different recovery steps. Logging them separately also improves diagnosis and product planning.
Recheck as traffic changes
A threshold tested on one language or message type may behave differently on another. Inspect outcomes by relevant traffic groups and time periods. Add new cases when the product expands, and revisit the review workload after changes to prompts or models.
Design confidence as an evaluated decision aid with an explicit fallback. The goal is to complete useful work within acceptable error limits while making uncertain cases easy to resolve, rather than making every request produce a confident-looking answer.