Papers
Topics
Authors
Recent
Search
2000 character limit reached

When do prophets profit in prediction markets?

Published 7 Jul 2026 in cs.AI, cs.CE, and cs.GT | (2607.06166v1)

Abstract: Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events. Classical theory establishes a clean equivalence between forecasting accuracy and trading profit, but only for the specific automated market maker (AMM) design. However, the largest exchanges today are based on central limit order books in which informed forecasters routinely lose money while uninformed strategies can profit on simple heuristics. We resolve this discrepancy by establishing a formal equivalence between predictive accuracy and profitability. For any strictly proper scoring rule $S$, we exhibit a "proper" betting strategy that depends only on the forecaster's prediction $\mathbf{p}$ and the market price $\mathbf{q}$, and earns positive expected profit whenever $\mathbf{p}$ outperforms $\mathbf{q}$ under $S$ and the market has sufficient liquidity. Moreover, this proper betting is essentially the only strategy with such robust profitability guarantee. The proof rests on a decomposition of expected profit that strictly generalizes the classical AMM guarantee and also explains how strategies can profit without an accuracy edge. Empirically, across thousands of forecasts by AI models, proper betting is the only strategy that reliably converts accuracy into profit, and we further identify systematic forecasting personas and show how the optimal proper strategy varies across them. A month-long live deployment on Kalshi achieves $+80.33\%$ return on investment with a Sharpe ratio of $3.35$.

Summary

  • The paper presents a unified theoretical and empirical framework showing that robust profit in prediction markets requires matching forecast accuracy with optimal bet sizing using proper betting strategies.
  • It rigorously decomposes expected profit into components such as score gap, Bregman divergence, and liquidity loss, challenging traditional heuristics like the Kelly criterion.
  • Live trading and extensive AI forecast backtests validate the approach, achieving +80.33% ROI and a Sharpe ratio of 3.35, underscoring its operational effectiveness.

Formal Summary and Analysis of "When do prophets profit in prediction markets?" (2607.06166)

This paper rigorously examines the disconnect between forecasting accuracy and trading profitability in contemporary prediction markets, especially those implemented with central limit order books (CLOBs) rather than the canonical automated market maker (AMM) structure. The authors develop a unified theoretical and empirical framework showing that robust profitability fundamentally requires strict alignment between forecasting accuracy (measured by proper scoring rules) and bet sizing (dictated by a mapping from forecaster and market beliefs). The study further exposes that, contrary to intuition, neither high forecasting accuracy nor standard betting heuristics (e.g., Kelly criterion) alone guarantee trading gains. The work resolves these inconsistencies by introducing a unique, provably robust "proper betting strategy," thoroughly analyzing its properties, and validating its performance with both large-scale AI forecast backtesting and live market deployment.


Theoretical Framework: Accuracy-Profitability Equivalence and Proper Betting

The core theoretical contribution is a complete characterization of when, and precisely how, forecasting accuracy can be converted to positive expected profit in real-world prediction markets. The authors generalize beyond the AMM paradigm by considering arbitrary price impact functions ρ\rho, covering both order books and AMMs as special cases. For every strictly proper scoring rule SS, there exists a proper betting strategy sG(p,q)=G(p)G(q)\boldsymbol{s}_G(\boldsymbol{p}, \boldsymbol{q}) = \nabla G(\boldsymbol{p}) - \nabla G(\boldsymbol{q}) (where GG is the convex potential associated to SS), guaranteeing strictly positive expected profit whenever a forecaster's accuracy edge over market prices is nontrivial after accounting for liquidity loss.

The expected profit admits a precise decomposition:

π(s,p)=[S(p;p)S(q;p)]+DG(q,p)Lρ(s;q)\pi(\boldsymbol{s}^*, \boldsymbol{p}^*) = [S(\boldsymbol{p};\boldsymbol{p}^*) - S(\boldsymbol{q};\boldsymbol{p}^*)] + D_G(\boldsymbol{q}, \boldsymbol{p}) - L_{\rho}(\boldsymbol{s}^*; \boldsymbol{q})

Here, the score gap quantifies the forecaster's accuracy advantage under SS, Bregman divergence DGD_G captures willingness to differ from the market, and liquidity loss penalizes trading at non-marginal prices. The essential result is that profitability is not determined by accuracy alone; both the divergence from the market and the ability to size bets efficiently matter.

Notably, the authors rigorously prove that, under mild liquidity assumptions, proper betting is the only robustly profitable strategy—up to scaling and constant shifts. Any alternative bet sizing scheme fails to guarantee profit when pitted against even moderately adversarial market conditions. The equivalence between the existence of robustly profitable betting and the strict propriety of scoring rules is also formally established, introducing a new operational criterion for properness.


Empirical Validation: AI Forecasters, Betting Heuristics, and Model Personas

The empirical section leverages thousands of AI-generated market forecasts across diverse domains—sports, politics, finance, and more—using data from the Prophet Arena benchmark. The authors compare proper betting, several canonical heuristics (max-margin, inverse-margin, Kelly, Kelly-alike), and find that only proper betting—where bet sizing derives directly from a proper scoring rule—routinely converts out-of-sample accuracy into positive ROI. Figure 1

Figure 1

Figure 1: Taxonomy of LLM forecasting personas by proportion of forecasts at small margins and accuracy consistency across margins.

Figure 2

Figure 2: Persona classification across models. ROI reported is under the best proper scoring rule: Log (L) or Brier (B).

The empirical decomposition shows that, for models with "dispersed" or "brittle" personas (as classified by forecast margin distributions and accuracy drop-off), profitability can arise from the Bregman divergence term even if the overt score gap (ΔS\Delta S) is negative—a significant qualitative divergence from AMM-era intuition. This is observed in the empirical ROI breakdowns, where highly accurate but insufficiently divergent (conservative) models are outperformed in profit terms by models willing to take large-magnitude disagreements with the market—demonstrating, for the first time at this scale, that both divergence and accuracy are co-determinants of trading performance.

Moreover, the study defines and maps forecasting personas among models, showing that optimal scoring rules, and hence proper betting strategies, differ by model profile. For instance, Brier-based strategies dominate for "conservative" personas with many small-edge bets, while logarithmic weighting is preferred for "dispersed" or loss-prone models. Figure 3

Figure 3: Capital allocation (bet weight in shares) as a function of signed margin and market price (q,y)(q, y) for the three proper betting strategies. Warmer colors indicate larger bets. Grey regions are infeasible (SS0). Contour lines mark constant weight levels.


Live Deployment: Real-Market Impact

The framework's practical viability is tested through live trading with a Gemini 3-based agent on the regulated Kalshi prediction market, executing proper betting (Brier rule) over 26 trading days. The deployment demonstrates +80.33% ROI and a Sharpe ratio of 3.35, even after accounting for slippage, bid-ask spreads, and transaction costs. Figure 4

Figure 4: Live trading statistics for the Gemini 3 agent on Kalshi.

This success underscores the framework's operational utility; nearly all gains derive from predictive accuracy (as measured via empirical Brier score), consistent with theory. Systematic outperformance is documented in detailed case studies, covering both high-volume liquid contracts and edge cases where market prices are internally inconsistent or lag public information. The agent outperforms all heuristic competitors, validating both the necessity and sufficiency of the theoretical prescriptions.


Implications, Limitations, and Future Directions

The research bridges a critical theoretical gap, showing that proper betting is necessary and sufficient for converting statistical accuracy into profitable trading in the presence of realistic market microstructure. This result challenges prevailing industry heuristics, revealing that careless bet sizing can destroy predictive advantage and that market friction fundamentally limits the classical link between information and profit outside AMMs. The introduction of model personas and the rule-dependent nature of profitability further suggest adaptive strategies: scoring rule selection should be tailored to forecaster profile and market regime.

Practically, the findings open avenues for designing more effective AI forecasting systems, including automated market participants, and raise interesting questions for market microstructure and regulatory design—specifically, how market rules should incentivize truthful and informative forecasting. The framework and analysis apply equally to prediction market designers, liquidity providers, and arbitrageurs, and have foundational consequences for the broader field of belief elicitation and information aggregation.

Conclusion

This paper resolves a central puzzle in prediction markets: the mismatch between superior forecasts and trading profits in modern (CLOB-based) venues. Through a rigorous, general theory of proper betting, supported by exhaustive empirical analysis and live deployment, the study establishes that only proper betting strategies, tied tightly to strictly proper scoring rules, robustly transform accuracy into profit. This insight reframes both the design of betting algorithms and the evaluation of forecasting models, with direct consequences for real-world deployment of both AI and human forecasters, as well as for the incentive structure of future prediction markets. The demonstration of persona-dependent strategy selection and the powerful live results mark significant progress and set the stage for further work on adaptive scoring rule choice, risk-aware bet sizing, and the interaction between market design and information aggregation efficacy.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 2 tweets with 16 likes about this paper.