Going All-In on LLM Accuracy: Fake Prediction Markets, Real Confidence Signals
arxiv.org·1h
🧠Intelligence Compression
Preview
Report Post

Authors:Michael Todasco (Visiting Fellow at the James Silberrad Center for Artificial Intelligence, San Diego State University)

View PDF

Abstract:Large language models are increasingly used to evaluate other models, yet these judgments typically lack any representation of confidence. This pilot study tests whether framing an evaluation task as a betting game (a fictional prediction market with its own LLM currency) improves forecasting accuracy and surfaces calibrated confidence signals. We generated 100 math and logic questions with verifiable answers. Six Baseline models (three current-generation, three prior-generation) answered all items. Three Predictor models then forec…

Similar Posts

Loading similar posts...