Despite the impressive capabilities of large language models (LLMs), their outputs often exhibit inconsistent correctness and unreliable factual accuracy. In high-stakes domains, overconfident yet incorrect predictions can lead to serious consequences, highlighting the need for robust uncertainty estimation. To address this, we introduce SelectLLM, an end-to-end method designed to enhance the ability of LLMs to recognize and express uncertainty effectively. By integrating selective prediction into finetuning, SelectLLM optimizes model performance over the covered domain, achieving a more balanced trade-off between predictive coverage and utility. Experimental results on TriviaQA, CommonsenseQA and MedConceptsQA show that SelectLLM significantly outperforms standard baselines, improving abstention behaviour while maintaining high accuracy
Bibtex
@misc{
mao2026selectllm,
title={Select{LLM} {\textendash} Calibrating {LLM}s for Selective Prediction: Balancing Coverage and Risk},
author={Yuzhen Mao and Thibaut Durand and Nazanin Mehrasa and Jiawei He and Martin Ester},
year={2026},
url={https://openreview.net/forum?id=JJPAy8mvrQ}
}
Related Research
-
Score-based Outlier Generation via Controlling the Radon-Nikodym Derivative
Score-based Outlier Generation via Controlling the Radon-Nikodym Derivative
A. Mukherjee, T. Milne, K. Y. C. Lui, S. Hazlewood, and J. Liu. IEEE Conference on Decision and Control
Publications
-
Investigating Action Embeddings for More Efficient Off-Policy Evaluation
Investigating Action Embeddings for More Efficient Off-Policy Evaluation
M. Wabartha, K. Wilson, R. David Evans, H. Sharifi, and T. Sylvain. RecSys CONSEQUENCES Workshop
Publications
-
FACTS: Fast, Accurate, and Privacy-Compliant TableSummarization via Offline Template Generation
FACTS: Fast, Accurate, and Privacy-Compliant TableSummarization via Offline Template Generation
Y. Yuan, M. Amin Shabani, and S. Liu. Workshop on Generative AI in Finance at NeurIPS
Publications