Papers
Topics
Authors
Recent
Search
2000 character limit reached

Information Leakage at Population Scale: An Evaluation of the Polymarket Insider-Relevant Subpopulation, 2020-2026

Published 1 May 2026 in q-fin.TR, econ.EM, and q-fin.GN | (2605.00459v1)

Abstract: We carry the deadline-resolved Information Leakage Score (ILS-dl) framework of Nechepurenko (2026a, 2026b) from a single-case proof of concept to a population-scale evaluation across 12,708 Polymarket markets, October 2020 to April 2026. We frame the paper as a scope-discovery study: scaling reveals that the framework's effective domain is materially narrower than initial framing suggested, and the principal obstacle is not score computation but resolution semantics. We report four findings. First, only 88 of 12,708 candidate markets (0.7%) yield computable ILS-dl values; only 1 of 32 markets in the ForesightFlow Insider Cases (FFIC) inventory is in scope, and 14 of 32 FFIC markets are flagged unclassifiable due to genuine resolution-criterion ambiguity. Second, only 12 of the 88 computed markets (13.6%) satisfy anchor-sensitivity, and an independent-second-pass T_event validation reaches 57.8% exact-date agreement, below the 90% ex-ante criterion. Third, raw ILS-dl medians are negative across all six (sub-bucket by period) cells, but a hazard-decay baseline correction we introduce yields a heterogeneous result: regulatory_formal post-2024 shifts to near-zero (-0.21 to -0.02), while regulatory_announcement post-2024 retains a 95% bootstrap CI entirely below zero. Fourth, the constant-hazard exponential of Nechepurenko (2026b) is rejected in favor of Weibull on the pooled post-2024 cell, but a per-subcategory check confirms the preference reflects category mixture rather than within-cell duration dependence. The implication is that detection of informed flow requires methodological refinement on the resolution-typology and score-baseline axes, not only on the score-computation axis where prior work concentrated.

Authors (1)

Summary

  • The paper introduces an ILSdl metric that quantifies pre-event price adjustments attributable to leaked information.
  • The paper employs a multi-stage analytical pipeline, integrating market categorization and LLM-assisted event timestamp recovery.
  • The paper shows that inherent market design ambiguities severely constrain the scalability of insider trading detection.

Information Leakage Detection at Scale in Prediction Markets: An Analytical Evaluation

Introduction

The paper "Information Leakage at Population Scale: An Evaluation of the Polymarket Insider-Relevant Subpopulation, 2020-2026" (2605.00459) presents a comprehensive population-scale evaluation of the Information Leakage Score framework for detecting informed trading in prediction markets, specifically focusing on Polymarket over a six-year interval. Building on prior single-case studies that introduced the deadline-resolved ILS (ILSdl^{\text{dl}}) methodology, the study critically examines the method’s scalability, operational robustness, and domain-specific limitations across 12,708 markets classified as insider-relevant.

Methodological Framework and Scope

The ILSdl^{\text{dl}} metric is designed to quantify the proportion of market-belief adjustment (via price movements) attributable to information reaching the market prior to public event disclosure. This formalism requires precise timestamp recovery for the event ((TeventT_{\text{event}})), robust resolution-typology classification (distinguishing deadline-resolved, event-resolved, or unclassifiable market types), and control for hazard-driven baseline decay in deadline contracts. The empirical analysis restricts focus to three a priori insider-relevant categories: regulatory decisions, corporate disclosures, and military/geopolitical events.

To operationalize the framework at scale, the authors employ a multi-stage pipeline including market categorization, typology assignment, CLOB data validation, and automated LLM-assisted recovery of TeventT_{\text{event}}. Only YES-resolved deadline markets with robust event-timestamp confidence and full price histories are used for ILSdl^{\text{dl}} computation.

Key Empirical Findings

1. Coverage and Scope Narrowness

The dominant finding is the exceptionally narrow empirical coverage of the framework: only 88 of 12,708 candidate markets (0.7%) yield interpretable ILSdl^{\text{dl}} values, with anchor-stable, operationally meaningful results surviving in just 12 markets. Structural ambiguities in market resolution predicates (e.g., in questions involving "strike", "custody", or compound outcomes) render a majority of insider-relevant contracts resistant to discrete event-timestamping. This phenomenon is particularly acute in the ForesightFlow Insider Cases (FFIC) validation set, where only one of 32 documented cases is in-scope for ILSdl^{\text{dl}} computation. The exclusion is not methodological but inherent to market/question design, highlighting that typical real-world insider-market episodes are not cleanly captured by rule-based typologies.

2. Resolution Semantics as Core Bottleneck

Empirical attrition in pipeline coverage primarily arises from the ambiguity of public event timing (TeventT_{\text{event}}). Automated LLM recovery of event times achieves only 57.8% exact-date agreement in independent second-pass validation (far below the target ≥90%\geq 90\%), especially in regulatory markets where multi-stage administrative processes predominate. Anchor-sensitivity filtering (i.e., stability of ILSdl^{\text{dl}} under dl^{\text{dl}}0 perturbations) further reduces the interpretable sample by nearly 87%.

3. Hazard-Decay Adjustment and Distributional Properties

Raw ILSdl^{\text{dl}}1 distributions are uniformly negative across all market sub-buckets and periods, which, absent correction, could misleadingly suggest systematic anti-information price drift. By introducing a per-cell hazard-rate–adjusted baseline (using Weibull fits at the pooled level), the authors demonstrate that the negative central tendency in several subgroups is wholly or mostly attributable to mechanical Bayesian decay in deadline contracts, not to "pricing against outcome." For instance, post-2024 regulatory_formal markets shift from a median raw value of −0.21 to an adjusted value of −0.02 after this correction. However, in regulatory_announcement post-2024 markets, the negative signal persists after adjustment, suggesting genuine pre-event anti-outcome drift not explained by rational decay. The hazard functional-form analysis further reveals that the constant-hazard exponential specification is decisively rejected at the aggregate level in favor of Weibull, attributable chiefly to heterogeneity across rather than within sub-buckets.

4. Implications for Method Development

The scaling exercise reveals that the central challenge is not computational but conceptual—score interpretation is severely gated by resolution semantics and the ambiguity of timestamp assignment in real insider-relevant market populations. The deadline-ILS approach, while interpretable and effective on clean, typologically well-posed single cases (e.g., Iran-Apr30), is operationally viable only on a minute fraction of real markets. The need to integrate multi-anchor event recovery and LLM-augmented, context-sensitive typology classification is clear. The data release and codebase make these findings fully reproducible, and the study sets explicit roadmaps for remediating current limitations.

Comparative Positioning and Theoretical Implications

Contrasted with recent wallet-level (Mitts & Ofir 2026) and account-level (Gómez-Cram et al. 2026) detection paradigms, the market-level ILSdl^{\text{dl}}2 framework is more restrictive but yields high-interpretability diagnostics when applicable. Methods that do not depend on typology classification or exquisite timestamping achieve substantially higher coverage but lack a direct probabilistic attribution of pre-event price movement share. The findings advocate for an ensemble approach, wherein market-level diagnostics are used in tandem with broader account- and wallet-based detectors to optimize detection in regulatory and compliance contexts.

Theoretically, the results underscore that true population-scale insider detection in prediction markets is limited by linguistic and ontological ambiguity in contract formulations—a domain in which current deep learning and LLM toolchains have only partial efficacy. The practical implication is that structural improvements to market/question design, or advanced multi-anchor event recognition (possibly with large-context model assistance), are necessary to operationalize real-time, high-confidence leakage detection at scale.

Statistical and Practical Observations

  • Only 0.7% of the candidate sample yields in-scope ILSdl^{\text{dl}}3; only 13.6% of these are anchor-stable.
  • Bootstrap CIs on hazard-adjusted medians indicate that only the regulatory_announcement post-2024 cell retains a distribution whose negative skew persists after rational-decay adjustment, a fact not explained by market mechanics.
  • The constant-hazard assumption for event-arrival times is structurally invalid for the pooled population, requiring compositional adjustments.
  • Many of the markets with strongest directional ILSdl^{\text{dl}}4 (right or left tail) are characterized by ambiguous resolution semantics, further illustrating the bottleneck.

Future Research Directions

Future work is sequenced on three axes:

  1. LLM-Assisted Re-classification: To expand coverage onto markets currently flagged as unclassifiable or multi-anchored in event semantics.
  2. Multi-anchor and Unstructured Predicate Handling: Methodological generalization to accommodate markets lacking clear, atomic resolution predicates.
  3. Wallet-Level Feature Integration: Extension to real-time joint-feature detection, leveraging operational advances in per-trade data pipelines and integrating additional wallet-level signals (e.g., PIN/VPIN, Kyle's lambda).

Empirical expansion to a minimum of 14 validated positive cases (for acceptable power in population detection) is expected within 60-90 days, contingent on data accumulation.

Conclusion

The study rigorously examines the limits of market-level leakage detection in prediction markets and establishes that, at scale, the pipeline’s primary bottlenecks are not computational but semantic and structural. The research provides substantial advances in methodological transparency, data availability, and diagnostic awareness, while clearly delimiting the boundary between current theoretical promise and operational reality.

The substantive conclusion is that effective, scalable detection of informed trading in prediction markets requires enhanced resolution-typology frameworks, robust multi-anchor event recovery, and hazard-rate–adjusted scoring—integrated via LLM-driven classification pipelines and ensemble detection architectures. The provided data and infrastructure form a robust base for future empirical and methodological developments in the field.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.