- The paper presents a unified theoretical and empirical framework showing that robust profit in prediction markets requires matching forecast accuracy with optimal bet sizing using proper betting strategies.
- It rigorously decomposes expected profit into components such as score gap, Bregman divergence, and liquidity loss, challenging traditional heuristics like the Kelly criterion.
- Live trading and extensive AI forecast backtests validate the approach, achieving +80.33% ROI and a Sharpe ratio of 3.35, underscoring its operational effectiveness.
This paper rigorously examines the disconnect between forecasting accuracy and trading profitability in contemporary prediction markets, especially those implemented with central limit order books (CLOBs) rather than the canonical automated market maker (AMM) structure. The authors develop a unified theoretical and empirical framework showing that robust profitability fundamentally requires strict alignment between forecasting accuracy (measured by proper scoring rules) and bet sizing (dictated by a mapping from forecaster and market beliefs). The study further exposes that, contrary to intuition, neither high forecasting accuracy nor standard betting heuristics (e.g., Kelly criterion) alone guarantee trading gains. The work resolves these inconsistencies by introducing a unique, provably robust "proper betting strategy," thoroughly analyzing its properties, and validating its performance with both large-scale AI forecast backtesting and live market deployment.
Theoretical Framework: Accuracy-Profitability Equivalence and Proper Betting
The core theoretical contribution is a complete characterization of when, and precisely how, forecasting accuracy can be converted to positive expected profit in real-world prediction markets. The authors generalize beyond the AMM paradigm by considering arbitrary price impact functions ρ, covering both order books and AMMs as special cases. For every strictly proper scoring rule S, there exists a proper betting strategy sG(p,q)=∇G(p)−∇G(q) (where G is the convex potential associated to S), guaranteeing strictly positive expected profit whenever a forecaster's accuracy edge over market prices is nontrivial after accounting for liquidity loss.
The expected profit admits a precise decomposition:
π(s∗,p∗)=[S(p;p∗)−S(q;p∗)]+DG(q,p)−Lρ(s∗;q)
Here, the score gap quantifies the forecaster's accuracy advantage under S, Bregman divergence DG captures willingness to differ from the market, and liquidity loss penalizes trading at non-marginal prices. The essential result is that profitability is not determined by accuracy alone; both the divergence from the market and the ability to size bets efficiently matter.
Notably, the authors rigorously prove that, under mild liquidity assumptions, proper betting is the only robustly profitable strategy—up to scaling and constant shifts. Any alternative bet sizing scheme fails to guarantee profit when pitted against even moderately adversarial market conditions. The equivalence between the existence of robustly profitable betting and the strict propriety of scoring rules is also formally established, introducing a new operational criterion for properness.
Empirical Validation: AI Forecasters, Betting Heuristics, and Model Personas
The empirical section leverages thousands of AI-generated market forecasts across diverse domains—sports, politics, finance, and more—using data from the Prophet Arena benchmark. The authors compare proper betting, several canonical heuristics (max-margin, inverse-margin, Kelly, Kelly-alike), and find that only proper betting—where bet sizing derives directly from a proper scoring rule—routinely converts out-of-sample accuracy into positive ROI.

Figure 1: Taxonomy of LLM forecasting personas by proportion of forecasts at small margins and accuracy consistency across margins.
Figure 2: Persona classification across models. ROI reported is under the best proper scoring rule: Log (L) or Brier (B).
The empirical decomposition shows that, for models with "dispersed" or "brittle" personas (as classified by forecast margin distributions and accuracy drop-off), profitability can arise from the Bregman divergence term even if the overt score gap (ΔS) is negative—a significant qualitative divergence from AMM-era intuition. This is observed in the empirical ROI breakdowns, where highly accurate but insufficiently divergent (conservative) models are outperformed in profit terms by models willing to take large-magnitude disagreements with the market—demonstrating, for the first time at this scale, that both divergence and accuracy are co-determinants of trading performance.
Moreover, the study defines and maps forecasting personas among models, showing that optimal scoring rules, and hence proper betting strategies, differ by model profile. For instance, Brier-based strategies dominate for "conservative" personas with many small-edge bets, while logarithmic weighting is preferred for "dispersed" or loss-prone models.
Figure 3: Capital allocation (bet weight in shares) as a function of signed margin and market price (q,y) for the three proper betting strategies. Warmer colors indicate larger bets. Grey regions are infeasible (S0). Contour lines mark constant weight levels.
Live Deployment: Real-Market Impact
The framework's practical viability is tested through live trading with a Gemini 3-based agent on the regulated Kalshi prediction market, executing proper betting (Brier rule) over 26 trading days. The deployment demonstrates +80.33% ROI and a Sharpe ratio of 3.35, even after accounting for slippage, bid-ask spreads, and transaction costs.
Figure 4: Live trading statistics for the Gemini 3 agent on Kalshi.
This success underscores the framework's operational utility; nearly all gains derive from predictive accuracy (as measured via empirical Brier score), consistent with theory. Systematic outperformance is documented in detailed case studies, covering both high-volume liquid contracts and edge cases where market prices are internally inconsistent or lag public information. The agent outperforms all heuristic competitors, validating both the necessity and sufficiency of the theoretical prescriptions.
Implications, Limitations, and Future Directions
The research bridges a critical theoretical gap, showing that proper betting is necessary and sufficient for converting statistical accuracy into profitable trading in the presence of realistic market microstructure. This result challenges prevailing industry heuristics, revealing that careless bet sizing can destroy predictive advantage and that market friction fundamentally limits the classical link between information and profit outside AMMs. The introduction of model personas and the rule-dependent nature of profitability further suggest adaptive strategies: scoring rule selection should be tailored to forecaster profile and market regime.
Practically, the findings open avenues for designing more effective AI forecasting systems, including automated market participants, and raise interesting questions for market microstructure and regulatory design—specifically, how market rules should incentivize truthful and informative forecasting. The framework and analysis apply equally to prediction market designers, liquidity providers, and arbitrageurs, and have foundational consequences for the broader field of belief elicitation and information aggregation.
Conclusion
This paper resolves a central puzzle in prediction markets: the mismatch between superior forecasts and trading profits in modern (CLOB-based) venues. Through a rigorous, general theory of proper betting, supported by exhaustive empirical analysis and live deployment, the study establishes that only proper betting strategies, tied tightly to strictly proper scoring rules, robustly transform accuracy into profit. This insight reframes both the design of betting algorithms and the evaluation of forecasting models, with direct consequences for real-world deployment of both AI and human forecasters, as well as for the incentive structure of future prediction markets. The demonstration of persona-dependent strategy selection and the powerful live results mark significant progress and set the stage for further work on adaptive scoring rule choice, risk-aware bet sizing, and the interaction between market design and information aggregation efficacy.