- The paper introduces an ILSdl metric that quantifies pre-event price adjustments attributable to leaked information.
- The paper employs a multi-stage analytical pipeline, integrating market categorization and LLM-assisted event timestamp recovery.
- The paper shows that inherent market design ambiguities severely constrain the scalability of insider trading detection.
Introduction
The paper "Information Leakage at Population Scale: An Evaluation of the Polymarket Insider-Relevant Subpopulation, 2020-2026" (2605.00459) presents a comprehensive population-scale evaluation of the Information Leakage Score framework for detecting informed trading in prediction markets, specifically focusing on Polymarket over a six-year interval. Building on prior single-case studies that introduced the deadline-resolved ILS (ILSdl) methodology, the study critically examines the method’s scalability, operational robustness, and domain-specific limitations across 12,708 markets classified as insider-relevant.
Methodological Framework and Scope
The ILSdl metric is designed to quantify the proportion of market-belief adjustment (via price movements) attributable to information reaching the market prior to public event disclosure. This formalism requires precise timestamp recovery for the event ((Tevent​)), robust resolution-typology classification (distinguishing deadline-resolved, event-resolved, or unclassifiable market types), and control for hazard-driven baseline decay in deadline contracts. The empirical analysis restricts focus to three a priori insider-relevant categories: regulatory decisions, corporate disclosures, and military/geopolitical events.
To operationalize the framework at scale, the authors employ a multi-stage pipeline including market categorization, typology assignment, CLOB data validation, and automated LLM-assisted recovery of Tevent​. Only YES-resolved deadline markets with robust event-timestamp confidence and full price histories are used for ILSdl computation.
Key Empirical Findings
1. Coverage and Scope Narrowness
The dominant finding is the exceptionally narrow empirical coverage of the framework: only 88 of 12,708 candidate markets (0.7%) yield interpretable ILSdl values, with anchor-stable, operationally meaningful results surviving in just 12 markets. Structural ambiguities in market resolution predicates (e.g., in questions involving "strike", "custody", or compound outcomes) render a majority of insider-relevant contracts resistant to discrete event-timestamping. This phenomenon is particularly acute in the ForesightFlow Insider Cases (FFIC) validation set, where only one of 32 documented cases is in-scope for ILSdl computation. The exclusion is not methodological but inherent to market/question design, highlighting that typical real-world insider-market episodes are not cleanly captured by rule-based typologies.
2. Resolution Semantics as Core Bottleneck
Empirical attrition in pipeline coverage primarily arises from the ambiguity of public event timing (Tevent​). Automated LLM recovery of event times achieves only 57.8% exact-date agreement in independent second-pass validation (far below the target ≥90%), especially in regulatory markets where multi-stage administrative processes predominate. Anchor-sensitivity filtering (i.e., stability of ILSdl under dl0 perturbations) further reduces the interpretable sample by nearly 87%.
3. Hazard-Decay Adjustment and Distributional Properties
Raw ILSdl1 distributions are uniformly negative across all market sub-buckets and periods, which, absent correction, could misleadingly suggest systematic anti-information price drift. By introducing a per-cell hazard-rate–adjusted baseline (using Weibull fits at the pooled level), the authors demonstrate that the negative central tendency in several subgroups is wholly or mostly attributable to mechanical Bayesian decay in deadline contracts, not to "pricing against outcome." For instance, post-2024 regulatory_formal markets shift from a median raw value of −0.21 to an adjusted value of −0.02 after this correction. However, in regulatory_announcement post-2024 markets, the negative signal persists after adjustment, suggesting genuine pre-event anti-outcome drift not explained by rational decay. The hazard functional-form analysis further reveals that the constant-hazard exponential specification is decisively rejected at the aggregate level in favor of Weibull, attributable chiefly to heterogeneity across rather than within sub-buckets.
4. Implications for Method Development
The scaling exercise reveals that the central challenge is not computational but conceptual—score interpretation is severely gated by resolution semantics and the ambiguity of timestamp assignment in real insider-relevant market populations. The deadline-ILS approach, while interpretable and effective on clean, typologically well-posed single cases (e.g., Iran-Apr30), is operationally viable only on a minute fraction of real markets. The need to integrate multi-anchor event recovery and LLM-augmented, context-sensitive typology classification is clear. The data release and codebase make these findings fully reproducible, and the study sets explicit roadmaps for remediating current limitations.
Comparative Positioning and Theoretical Implications
Contrasted with recent wallet-level (Mitts & Ofir 2026) and account-level (Gómez-Cram et al. 2026) detection paradigms, the market-level ILSdl2 framework is more restrictive but yields high-interpretability diagnostics when applicable. Methods that do not depend on typology classification or exquisite timestamping achieve substantially higher coverage but lack a direct probabilistic attribution of pre-event price movement share. The findings advocate for an ensemble approach, wherein market-level diagnostics are used in tandem with broader account- and wallet-based detectors to optimize detection in regulatory and compliance contexts.
Theoretically, the results underscore that true population-scale insider detection in prediction markets is limited by linguistic and ontological ambiguity in contract formulations—a domain in which current deep learning and LLM toolchains have only partial efficacy. The practical implication is that structural improvements to market/question design, or advanced multi-anchor event recognition (possibly with large-context model assistance), are necessary to operationalize real-time, high-confidence leakage detection at scale.
Statistical and Practical Observations
- Only 0.7% of the candidate sample yields in-scope ILSdl3; only 13.6% of these are anchor-stable.
- Bootstrap CIs on hazard-adjusted medians indicate that only the regulatory_announcement post-2024 cell retains a distribution whose negative skew persists after rational-decay adjustment, a fact not explained by market mechanics.
- The constant-hazard assumption for event-arrival times is structurally invalid for the pooled population, requiring compositional adjustments.
- Many of the markets with strongest directional ILSdl4 (right or left tail) are characterized by ambiguous resolution semantics, further illustrating the bottleneck.
Future Research Directions
Future work is sequenced on three axes:
- LLM-Assisted Re-classification: To expand coverage onto markets currently flagged as unclassifiable or multi-anchored in event semantics.
- Multi-anchor and Unstructured Predicate Handling: Methodological generalization to accommodate markets lacking clear, atomic resolution predicates.
- Wallet-Level Feature Integration: Extension to real-time joint-feature detection, leveraging operational advances in per-trade data pipelines and integrating additional wallet-level signals (e.g., PIN/VPIN, Kyle's lambda).
Empirical expansion to a minimum of 14 validated positive cases (for acceptable power in population detection) is expected within 60-90 days, contingent on data accumulation.
Conclusion
The study rigorously examines the limits of market-level leakage detection in prediction markets and establishes that, at scale, the pipeline’s primary bottlenecks are not computational but semantic and structural. The research provides substantial advances in methodological transparency, data availability, and diagnostic awareness, while clearly delimiting the boundary between current theoretical promise and operational reality.
The substantive conclusion is that effective, scalable detection of informed trading in prediction markets requires enhanced resolution-typology frameworks, robust multi-anchor event recovery, and hazard-rate–adjusted scoring—integrated via LLM-driven classification pipelines and ensemble detection architectures. The provided data and infrastructure form a robust base for future empirical and methodological developments in the field.