- The paper introduces a combined Double ML and Bayesian method to estimate incremental bookings while accounting for substitution effects and cannibalization.
- It leverages geospatial similarity metrics and hierarchical Bayesian modeling to enhance precision in capturing heterogeneous treatment effects across listing segments.
- Empirical results demonstrate adaptive model behavior validated by domain expertise, offering actionable insights for targeted acquisition and marketplace optimization.
Estimating Supply Incrementality in Two-sided Marketplaces: A Causal Machine Learning Approach
The paper "Estimating Supply Incrementality in Two-sided Marketplaces: A Causal Machine Learning Approach" (2606.30999) addresses the challenge of quantifying the incremental value that additional supply creates in heterogeneous, two-sided marketplaces. Using Airbnb as a case study, the authors focus on the causal relationship between the increase in listing supply and marketplace outcomes, specifically total bookings. Supply incrementality is defined as the contribution of new listings to overall bookings, after accounting for cannibalization effects—that is, bookings that would have occurred elsewhere in the absence of new supply.
A crucial difficulty in measuring incrementality arises from the unobservability of the counterfactual scenario (market outcomes absent the new supply), and the endogeneity of supply and demand dynamics. Substitution effects, where bookings shift between similar listings, further complicate estimation. The paper motivates the importance of robust incrementality estimation for optimization of product features, targeted marketing initiatives, and resource allocation.
Figure 1: Diagrammatic representation of supply incrementality, distinguishing between incremental bookings and cannibalized transactions.
Methodological Framework
Double Machine Learning and Geospatial Feature Engineering
The proposed method integrates Double/Debiased Machine Learning (Double ML) [Chernozhukov et al.], a hierarchical Bayesian framework, and geospatial similarity metrics for feature construction. Double ML mitigates confounding by first regressing both supply and bookings on a rich set of supply-side and demand-side covariates, and then estimating the treatment effect on residuals. This approach retains consistency under mild model misspecification conditions and facilitates flexible, high-dimensional modeling.
Geospatial similarity measures are leveraged for feature engineering, effectively reducing the variance of heterogeneous treatment effect estimates by focusing on the most relevant listing segments and minimizing noise from distant or fundamentally dissimilar segments.
Hierarchical Bayesian Modeling
The second-stage model estimates treatment effects (incrementality coefficients) as a function of both intra-segment and inter-segment product features. Priors from established supply incrementality studies are integrated at the listing-group level, with empirical data updating the posterior beliefs when sufficient signal is present. Random effects capture baseline incrementality at the group level, and interaction effects parameterize heterogeneity at finer granularities.
Empirical Results
Posterior Insights and Priors Reconciliation
Posterior estimates for interaction coefficients (λ) are generally aligned with theoretical priors, but demonstrate increased precision, as evidenced by narrower credible intervals. In instances where posterior coefficients diverge in sign or magnitude from priors, these differences correspond to meaningful shifts in demand and supply dynamics, substantiated by empirical data.
Figure 2: Prior vs. posterior distributions for heterogeneous treatment effect interaction coefficients, showing consistency and improved precision.
Aggregating estimates at the product group level, posterior incrementality diverges materially from prior only in the largest supply groups with distinct demand profiles, indicating adaptive model behavior responsive to high-volume data and heterogeneous segments.
Figure 3: Group-level comparison between prior and posterior incrementality estimates, visualizing adaptive Bayesian updating where data is most informative.
Segment-Level Effects and Domain Consistency
Segment-level incrementality estimates display high fidelity with domain expectations; segments characterized by strong demand signals and unique, non-substitutable attributes yield higher incrementality. Validation via domain expertise and auxiliary research corroborates the model output, supporting its reliability for guiding business decisions.
Evaluation Strategy and Limitations
Model evaluation leverages alignment with domain knowledge, robustness checks, and validation against natural experiments (e.g., exogenous supply shocks). In the absence of feasible randomized supply interventions, the approach relies on consistency with established priors and sanity-checks at multiple granularity levels. The authors acknowledge the fundamental limitation of unobservable counterfactual outcomes and emphasize the importance of alternative validation strategies, including potential exploitation of natural supply shocks.
Implications and Future Directions
The approach provides scalable, heterogeneous incrementality estimates that are directly applicable to critical business processes within two-sided marketplaces. Practically, these can inform targeted acquisition, segment prioritization, and platform design. Theoretically, the combination of causal machine learning and Bayesian hierarchical modeling represents a systematic advance for observational causal inference in platforms with complex product heterogeneity. The integration of geospatial similarity metrics underscores the necessity to spatially structure marketplace data, improving estimation precision and treatment effect identification.
The methodology is extensible to other marketplaces—ride sharing, e-commerce, content platforms—where incremental supply decisions drive core outcomes. Future research avenues include dynamic/temporal modeling, alternative similarity frameworks, and development of quasi-experimental validation strategies to further probe causal estimates in high-stakes operational environments.
Conclusion
This paper introduces and empirically validates a methodologically sophisticated framework for supply incrementality estimation in two-sided marketplaces with heterogeneous products. By fusing Double ML, geospatial feature construction, and Bayesian hierarchical modeling, the approach accommodates endogeneity, substitution effects, and leverages pre-existing domain knowledge. The model yields precise, heterogeneous treatment effects, reconciles theoretical priors with data-driven posteriors, and offers practical utility for marketplace optimization. Its extensibility to varied platform types and alignment with theoretical foundations position it as a robust tool for future research and actionable analytics.