Papers
Topics
Authors
Recent
Search
2000 character limit reached

Propensity Score Weighting to Ensure Balance in Key Subgroups or Strata: A Practical Guide

Published 15 Apr 2026 in stat.ME | (2604.14407v1)

Abstract: Propensity score weighting approaches have been widely implemented in clinical research to estimate the effects of a treatment or exposure while mitigating the risk of confounding in the absence of random assignment. In practice, when working with large electronic health records (EHR) or administrative datasets to evaluate health quality outcomes at the institutional level, or evaluate supportive care interventions for a wide range of hospitalized patients, it may be advisable to stratify the propensity score weighting approach by indication, reason for admission, or other clinical risk factors due to the potential for substantial heterogeneity across subgroups of patients with complex care needs. A stratified approach may be appropriate if (i) prognosis differs substantially between patient subgroups such that achieving balance in the composition of these strata between exposure/treatment groups should be prioritized, (ii) likelihood of exposure differs substantially across clinical subgroups, or (iii) the covariate-exposure associations are expected to differ substantially between subgroups (i.e. there are covariate-subgroup interactions in the exposure/treatment propensity model). For example, we may want to evaluate the impact of prophylactic anticoagulant use for venous thromboembolism prevention in elderly patients admitted to hospital for a wide array of conditions. The purpose of this article is to outline an approach to implementing propensity score weighting with stratification by clinical groups. We also provide guidance on best practices with particular focus on EHR and administrative medical data, and population health settings.

Summary

  • The paper introduces a stratified weighting approach that scales weights within subgroups to ensure balance and improve causal effect estimation.
  • It emphasizes rigorous diagnostics using metrics like standardized mean differences and effective sample sizes to validate the weighting strategy.
  • The methodology is practically demonstrated on an oncology dataset, illustrating improvements in subgroup balance and highlighting residual imbalances.

Propensity Score Weighting with Stratification: Methodological Advances and Practical Guidance

Introduction

The paper "Propensity Score Weighting to Ensure Balance in Key Subgroups or Strata: A Practical Guide" (2604.14407) delineates a rigorous framework for implementing propensity score weighting in settings where substantial heterogeneity exists across clinical subgroups within observational data. The authors advance a stratified weighting technique, highlighting its necessity when baseline prognosis, likelihood of exposure, or covariate-exposure relationships vary across strata. They emphasize the importance of achieving balance in both overall and within-stratum distributions of confounders, particularly when the population under study is drawn from large EHR or administrative datasets encompassing diverse patient subgroups.

Methodological Framework

Potential Outcomes and Ignorability

Building on the potential outcomes paradigm, the paper affirms the requirement for both SUTVA and ignorability assumptions to enable unbiased estimation of causal effects from observational data. Under these conditions, the average treatment effect (ATE) can be estimated by comparing outcomes in populations balanced on all measured confounders.

Standard Propensity Score Weighting

Traditional propensity score weighting constructs a pseudo-population where observed covariates are balanced between treatment groups, typically via logistic regression. However, the authors articulate several pitfalls in heterogeneous populations: residual imbalance in key subgroups, poor balance within strata due to inadequate modelling of covariate-subgroup interactions, and distortion in the weighted distribution of strata, all threatening internal and external validity.

Stratified Weighting Approach

The authors propose a stratified implementation, whereby propensity scores are estimated separately within each stratum. This approach allows for distinct modelling considerations pertinent to each subgroup. The weighting schema is augmented with scaling factors to ensure that weighted stratum shares are congruent with the underlying population or target population, thereby maintaining interpretable marginal effect estimates.

The technical implementation involves a two-stage adjustment on the initial weights: first, scaling within each stratum to achieve exposure group balance, second, further rescaling so the overall weighted stratum composition mimics the original data. The approach supports the estimation of marginal ATEs and is adaptable to other estimands, such as ATT, via formulaic modification.

Empirical Demonstration

The paper provides a simulated oncology dataset to illustrate both conventional and stratified weighting methodologies. Key findings include:

  • Unstratified weighting yields only partial balance, especially when strong baseline imbalances exist, and results in an altered marginal distribution of key strata, distorting external validity.
  • Stratified weighting achieves exact balance in stratum proportions and allows for more meaningful marginal comparisons while providing subgroup-specific effect estimates.
  • Despite stratification, residual imbalances on continuous confounders (e.g., age) may persist, necessitating further within-stratum diagnostics.

Effective sample size (ESS) calculations quantify the variance-precision trade-off inherent to weighting, and diagnostics such as standardized mean differences (SMDs) are crucial for confirming balance post-adjustment.

Practical Implications

The stratified approach is particularly relevant in clinical research contexts characterized by multimodal heterogeneity in prognosis and treatment assignment, such as institutional health quality outcomes and supportive care intervention studies. Practitioners are advised to:

  • Conduct balance diagnostics within each stratum, including SMDs and ESS, post-weighting,
  • Examine distributional overlap for potential positivity violations,
  • Carefully specify exposure and outcome models with context-specific confounder selection,
  • Use robust variance estimators or bootstrapping to account for weight-induced heteroscedasticity.

The method integrates seamlessly within the target trial emulation framework, supporting rigorous causal inference from real-world data while addressing the demand for transparency in regulatory and health technology assessment submissions.

Theoretical Extensions and Limitations

Stratified propensity score weighting ensures the internal consistency of treatment effect estimation by preserving subgroup composition and allowing for targeted adjustment of confounder distributions. It is generalizable to matching, covariate balancing weighting, and parametric g-computation. Limitations include the fundamental assumptions of exchangeability, SUTVA, correct model specification, and sufficient overlap—failures in any may bias treatment effect estimates or reduce precision.

Future Directions

The approach sets the stage for further methodological advances in the handling of complex observational health datasets. Potential future developments include:

  • Integration with machine learning-based propensity score estimation to capture more nuanced covariate-subgroup interactions,
  • Extension to time-varying exposures and dynamic treatment regimes,
  • Formal policy-oriented methods to improve transportability of marginal estimands across populations,
  • Investigation into hybrid doubly-robust variants leveraging both stratum-specific outcome regression and propensity score models.

Conclusion

The stratified propensity score weighting methodology offered in the paper provides a rigorous, reproducible framework for maintaining balance in key clinical subgroups when estimating causal effects from observational data. Its practical utility is broad, particularly in heterogeneous health datasets, supporting defensible causal inferences for both regulatory and research applications. As stratification becomes increasingly central to the analysis of real-world evidence, this methodological guide is poised to inform both theoretical development and practical application in causal inference.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.