Key Points
Kalshi Research studied more than 2.2 million resolved markets from the platform’s 2021 launch through mid-2026. The study found Kalshi prices become more accurate and better calibrated as markets approach resolution. Trading volume and trader participation improved calibration, though the report cautioned that market depth does not explain everything.
Prediction markets may be more than speculative trading venues when enough traders participate.
A new Kalshi Research study of more than 2.2 million resolved markets found that prices on the regulated exchange increasingly behave like real probabilities as events move closer to resolution.
The working paper by Nicole Kagan and Rubens Baiocchi examined 2,243,741 resolved Kalshi markets across 11 categories from the platform’s 2021 launch through mid-2026.
The study found that Kalshi prices were “extremely well calibrated” in aggregate as markets neared resolution, with Brier scores falling from roughly 0.08 to 0.09 at a three-month horizon to about 0.02 at close. Naive accuracy also rose from 88.3% three months out to 97.2% at close.
Prediction Markets Get Their Credibility Test
The central question in the report is whether a prediction market price can be read literally as a probability. In simple terms, if a contract trades at 70 cents, it should resolve to “Yes” about 70% of the time across a large enough sample.
Kalshi’s data largely supports that claim. The report found prices were well calibrated overall and improved as resolution approached. That gives prediction markets a stronger claim as real-time forecasting tools for traders, policymakers, researchers and journalists.
The finding does not mean every market price should be trusted blindly. The study shows reliability depends on conditions. Time matters, volume matters and the number of participating traders matters.
More Traders Make Forecasts Sharper
The strongest forward-looking finding is that deeper markets produce better signals. The report found Brier scores generally fell as trading volume increased, suggesting that prices sharpen as more activity accumulates.
That relationship was clearest near resolution. Markets with less than $10,000 in event volume had a close-horizon Brier score of 0.0635. Markets with at least $200,000 in volume had a much lower score of 0.0099.
Also Read: Dogecoin Rebounds 10% While ETF Flows Stay Quiet
Trader count showed a similar pattern. At close, markets with fewer than 20 traders had a Brier score of 0.0354, while markets with at least 1,000 traders had a score of 0.0089. The report said participation improves calibration, but also noted that trader count is an imperfect measure because high-profile events can attract more traders while also being harder to forecast.
Not Every Market Is Equally Reliable
The study also found meaningful differences between categories. Economics markets showed a steady and almost linear improvement in Brier scores because they usually resolve against scheduled releases such as CPI, payrolls or GDP data.
Sports and mentions markets were more complicated because information can emerge during the event itself.
The timing of an event also matters. The report said apparent calibration issues in sports and elections changed when researchers measured markets from the true occurrence of the underlying event rather than the administrative market close timestamp.
That distinction is important for the prediction market industry. If market prices are increasingly used as public probability feeds, users need to know when the signal is deep and when it is still noisy.
Kalshi’s study gives prediction markets a credibility boost, but its larger message is more careful. The best prices become useful forecasting signals when enough information, liquidity and traders are present.
As the market grows, the next test will be whether that calibration holds as participation expands beyond today’s most active users.
Read Next: Chainlink Rallies Past $10.60 After Nazarov Discusses Stablecoins At White House





