Confounder Detection via Treatment Intent
- The paper presents CDTI, a method that contrasts treated–untreated pairs to surface candidate hidden confounders by eliciting treatment intent.
- It employs matching strategies such as Z-matching, π-matching, and Z-dominance to suppress observed variables and emphasize unobserved treatment drivers.
- Empirical studies in ICU EHR settings demonstrate plausible candidate confounders, highlighting the need for further validation in causal adjustment.
Searching arXiv for the specified paper and closely related work to ground the article in the current literature. arXiv search: "(Plecko et al., 26 May 2026) Confounder Detection via Treatment Intent" Confounder Detection via Treatment Intent (CDTI) is an observational study design for detecting candidate unobserved confounders by exploiting the fact that any hidden variable that affects treatment allocation must, in some form, have been available to the human decision-maker at decision time. The design constructs carefully matched treated–untreated pairs and then asks why treatment differed, with the goal of eliciting variables outside the recorded covariate set that explain the treatment contrast. CDTI is therefore not a causal estimator by itself. Its role is to surface candidate hidden treatment drivers that may later be measured, proxied, or incorporated into downstream adjustment procedures (Plecko et al., 26 May 2026).
1. Problem formulation and causal motivation
In the CDTI formulation, the treatment is , the outcome is , the observed covariates are , and the unobserved confounders are . The intended causal structure is
with possible dependence between and , represented by a bidirected relation. The observational analyst sees , but not . When affects both treatment and outcome, adjustment on 0 alone is insufficient, so the usual no-unobserved-confounding condition fails (Plecko et al., 26 May 2026).
The ICU example used to motivate CDTI makes this point concrete. Let 1 be mechanical ventilation and 2 be in-hospital mortality. Using observed EHR covariates 3, the paper estimates the effect of treatment on the treated,
4
Across three ICU databases, the back-door-adjusted estimates suggest that mechanical ventilation increases mortality in the treated population. The paper interprets this as a near-certain empirical signature of unobserved confounding, because physicians ventilate patients for reasons that are only partly captured in structured EHR variables (Plecko et al., 26 May 2026).
This design addresses a specific failure mode of observational causal inference: the dataset does not contain all variables that guided treatment choice. Related hidden-confounding work makes the same identifiability obstacle explicit in simpler binary settings. In the three-node DAG 5, 6, 7, the ATE is not identifiable from confounded observational data alone because many different full joint distributions 8 induce the same observed 9 while implying different causal effects (Gan et al., 2020). CDTI starts from the practical premise that, although 0 is not recorded, a human decision-maker may still be able to articulate contrastive reasons for treatment differences (Plecko et al., 26 May 2026).
2. Matching design and stochastic dominance theory
CDTI does not ask experts abstractly which unmeasured variables might matter. It first constructs treated–untreated pairs 1 such that 2 and 3, then queries why treatment differed. The core design choice is the matching strategy 4, which suppresses obvious observed explanations in 5 so that hidden explanations in 6 become more salient (Plecko et al., 26 May 2026).
| Strategy | Pairing condition |
|---|---|
| 7-matching | 8 |
| 9-matching | 0, where 1 |
| 2-dominance | 3 coordinatewise |
| Marginal matching / random baseline | Only 4, 5 |
The central theorem is stated in terms of multivariate stochastic order. Under appropriate assumptions 6, for each strategy 7,
8
9
0
Here 1 denotes multivariate stochastic order: 2 for every coordinatewise non-decreasing 3 (Plecko et al., 26 May 2026).
For scalar 4, the sufficient condition for 5-matching is especially simple. If 6 and
7
is non-decreasing in 8 for every 9, then
0
This is a monotone likelihood ratio argument: within a fixed 1-stratum, treated units have larger posterior 2 than untreated units when treatment probability increases with 3 (Plecko et al., 26 May 2026).
For multivariate 4, the conditions are stronger. The treatment mechanism must be non-decreasing in each coordinate of 5, and 6 must be log-supermodular: 7 or, for twice-differentiable densities,
8
The log-supermodularity condition rules out strong tradeoffs among hidden severity axes; clinically, it encodes that hidden severities tend to cluster positively rather than offset one another (Plecko et al., 26 May 2026).
For 9-dominance, the paper introduces
0
and requires cross-partial conditions: 1
2
Under these plus the 3-matching assumptions, if 4, then
5
The paper interprets these inequalities as balancing a collider channel through conditioning on 6 with a bidirected channel through marginal dependence between 7 and 8 (Plecko et al., 26 May 2026).
3. Expert elicitation, extraction, and success probabilities
The expert interaction is formalized as an extraction-and-selection process. Given a pair 9 with 0 and 1, the expert produces a candidate set of explanations
2
where 3 is the extraction strategy. Under perfect extraction 4,
5
Thus, only variables on which the treated unit exceeds the untreated unit can explain the contrastive treatment decision (Plecko et al., 26 May 2026).
Selection is then modeled as choosing one explanation uniformly among candidates. If 6 denotes the event that an unobserved variable is selected, then
7
Accuracy 8 requires that the selected explanation is unobserved and genuinely satisfies the directional contrast 9. Under perfect extraction,
0
where
1
The design objective is therefore explicit: maximize the number of hidden explanatory differences 2 while minimizing the number of competing observed explanations 3 (Plecko et al., 26 May 2026).
This leads to formal comparisons among strategies. Under perfect extraction, both 4-matching and 5-dominance have 6, so success reduces to
7
The paper proves that
8
Hence, under the stated assumptions, dominance pairs are even better than exact 9-matches (Plecko et al., 26 May 2026).
For 0-matching versus random matching, the paper defines a surrogate utility
1
and marginal predictive strength
2
Under the 3-matching conditions and continuity of each marginal 4,
5
A stated implication is that if
6
then
7
This formalizes the intuition that balancing away observed variation is worthwhile when it removes more distracting 8-signal than hidden 9-signal (Plecko et al., 26 May 2026).
4. Empirical implementation in ICU and EHR settings
The paper’s proof of concept is built around ICU treatment effect estimation from EHR data, with mechanical ventilation as treatment and in-hospital mortality as outcome. In the semi-synthetic MIMIC-III setting, the 12 observed covariates are respiratory rate, mean arterial pressure (MAP), lactate, P/F ratio, 00, 01, oxygen saturation, age, sex, Charlson score, SOFA score, and admission-diagnosis indicators (Plecko et al., 26 May 2026).
Clinical notes are used as a proxy for physician knowledge. For the semi-synthetic environment, the paper extracts binary note-derived concept indicators using UMLS entity linking and negation detection implemented via SciSpaCy. The binary concepts used as latent confounders 02 include pleural effusion, heart failure, dyspnea, pneumonia, pulmonary edema, Bloom syndrome, hypoxia, atrial fibrillation, hypertensive disease, and atelectasis, with 03, all binary (Plecko et al., 26 May 2026).
The semi-synthetic construction starts with real MIMIC-III 04 and text-derived 05, then specifies treatment and outcome models: 06 with 07, and a logistic outcome model with 08 for 09 and direct treatment coefficient 10, so treatment is truly protective. New 11 are then sampled, while analysts observe only 12 (Plecko et al., 26 May 2026).
The synthetic verification uses Gaussian 13 with 14 and logistic treatment
15
When the required assumptions hold, 16-dominance outperforms 17-matching as the dominance gap grows; when they fail, the advantage can reverse. For 18-matching versus random matching, the paper varies
19
and reports that 20-matching beats random matching for 21, while the inequality reverses for 22 (Plecko et al., 26 May 2026).
In the semi-synthetic MIMIC-III experiments, practical matching implementations are based on Euclidean distance on standardized 23, coordinatewise dominance scores, estimated propensity differences, and random baselines. The main finding is an ordering of cumulative success rates: 24-based strategies 25-match and 26-dominance are best, 27-matching and 28-dominance are next, and random matching is worst throughout. The paper also reports that mean success 29 is highest for small propensity gaps and decreases monotonically as the gap widens (Plecko et al., 26 May 2026).
For real-data detection in MIMIC-III, the paper fits a BERT-based treatment predictor 30 for 31 using physiological covariates and text notes, and an xgboost model 32 using only 33, which is used to construct 34-matched pairs. Using 35-matching with 36 selected pairs on a held-out set, and allowing each unit to appear at most three times, the paper performs ablation-based extraction. For each concept 37, it removes mentions from the treated unit’s notes and defines concept impact as
38
Concepts with 39 are kept as candidate confounders (Plecko et al., 26 May 2026).
The top 20 discovered concepts include pleural effusion, hemorrhage, pneumonia, dyspnea, hypoxia, pulmonary edema, pneumothorax, fever, tachycardia, and hypertensive disease. The paper groups them into pulmonary impairment/injury, infection/SIRS, hemorrhage/trauma, cardiac, and non-specific illness severity, and concludes that 40 detected concepts are clinically highly plausible as confounders or confounder proxies for ventilation decisions (Plecko et al., 26 May 2026). However, when these discovered concepts are added to downstream adjustment sets, the ETT changes only slightly and the differences are not statistically significant. The stated explanation is that note-derived binary indicators are sparse and noisy proxies for the underlying latent confounders (Plecko et al., 26 May 2026).
5. Relation to adjacent deconfounding and confounder-selection frameworks
CDTI occupies a distinct position within a broader literature on hidden confounding, confounder selection, and treatment-assignment signals. Its closest relatives address neighboring problems rather than the same design objective.
"Selective deconfounding" in "Causal Inference With Selectively Deconfounded Data" (Gan et al., 2020) assumes a large observational dataset in which treatment and outcome are observed but one key confounder is missing, together with a smaller dataset in which that confounder is revealed. The contribution there is not confounder detection, but selective confounder measurement. The method reconstructs 41 from known 42 and estimated 43, and shows that actively allocating the revealed-44 budget across observed 45 strata can reduce sample complexity (Gan et al., 2020). This is methodologically close to selective measurement design, whereas CDTI is aimed at eliciting previously unrecorded treatment drivers (Plecko et al., 26 May 2026).
"Adversarial De-confounding in Individualised Treatment Effects Estimation" (Chauhan et al., 2022) uses treatment assignment as a proxy for treatment policy and learns disentangled latent representations 46. A treatment classifier connected to 47 through a gradient reversal layer encourages treatment-agnostic balanced confounder representations for ITE estimation (Chauhan et al., 2022). That approach uses treatment-assignment signals to shape latent space, but it does not provide explicit variable-level confounder detection or expert-elicited treatment-intent explanations. CDTI, by contrast, is explicitly contrastive and pair-based (Plecko et al., 26 May 2026).
"ConfoundingSHAP: Quantifying confounding strength in causal inference" (Brockschmidt et al., 11 May 2026) targets a different detection problem: which observed covariates act as confounders. It defines residual confounding bias under restricted adjustment,
48
and constructs Shapley values over adjustment-set coalitions to quantify how much each observed covariate reduces residual bias relative to the full observed set (Brockschmidt et al., 11 May 2026). This distinguishes confounders from instruments or prognostic variables among measured covariates. CDTI instead seeks candidate unobserved confounders that are absent from the original 49 (Plecko et al., 26 May 2026).
"Confounder selection via iterative graph expansion" (Guo et al., 2023) addresses confounder selection through causal-graphical elicitation. Starting from 50, it queries the user for primary adjustment sets and incrementally expands a working ADMG until a sufficient adjustment set is found or ruled out. The method is sound and complete if the user correctly specifies the primary adjustment sets at every step (Guo et al., 2023). This is interactive and expert-driven like CDTI, but the elicited object is structural graph information rather than contrastive treatment intent. A plausible implication is that CDTI could provide candidate variables that later enter a graphical confounder-selection workflow.
"Inferring the Effect of a Confounded Treatment by Calibrating Resistant Population's Variance" (Qin et al., 2023) is relevant at the level of hidden treatment-selection bias rather than confounder discovery. It identifies the conditional average treatment effect on the treated as one of two possible values under nondeterministic treatment assignment, equality of conditional variances of the two potential outcomes in the treatment group, and a resistant population that calibrates 51 (Qin et al., 2023). This quantifies the magnitude of hidden bias caused by latent treatment selection, but it does not detect the underlying confounder or infer treatment intent directly.
6. Interpretation, misconceptions, and limitations
CDTI is best understood as a study design for generating candidate hidden confounders, not as a proof that a detected concept is a true confounder of 52. The paper states explicitly that detecting a candidate hidden treatment driver does not certify that it is a genuine confounder, nor that adding it restores back-door admissibility (Plecko et al., 26 May 2026). This distinguishes CDTI from methods that assume the confounder is already known and merely expensive to measure (Gan et al., 2020).
A second common misconception is to equate treatment intent with any variable predictive of treatment. Related work on observed-covariate confounding makes clear that variables associated only with treatment assignment may instead be instruments, while variables associated only with outcomes are prognostic covariates (Brockschmidt et al., 11 May 2026). CDTI avoids making that identification from treatment prediction alone; it elicits candidate variables because they explain treatment disagreement after observed explanations have been neutralized (Plecko et al., 26 May 2026).
The paper gives explicit conditions under which CDTI can succeed: hidden variables 53 affect treatment monotonically, hidden dimensions are positively associated at fixed 54, matching suppresses observed explanations 55, the expert can articulate contrastive reasons for treatment differences, and elicited concepts are measured or proxied with enough fidelity to be useful later (Plecko et al., 26 May 2026). It can fail when treatment depends on tacit, subconscious, or inarticulable cues; when 56 coordinates trade off rather than co-cluster; when 57 and 58 are too strongly correlated; when matching leaves many observed competing explanations; when extracted variables are noisy or post-treatment; or when the elicited concept is plausible but not actually a confounder of 59 (Plecko et al., 26 May 2026).
Practical limitations are substantial. The method imposes expert burden through repeated pairwise comparisons. Elicited explanations may be vague, high-level, or inconsistent. Measurement error is central: in the ICU application, clinically plausible concepts were detected, but downstream bias reduction remained limited because note-derived binary indicators were sparse and noisy proxies. Any elicited variable used in causal analysis must also be verified as pre-treatment; in note-based settings, timestamps and retrospective documentation matter (Plecko et al., 26 May 2026).
Within the broader causal-inference landscape, CDTI therefore fills a narrow but important role. It is not a replacement for randomized trials, front-door or IV identification, sensitivity analysis, or graphical adjustment criteria. It is a procedure for surfacing what the dataset failed to record by querying differential treatment decisions in carefully selected pairs (Plecko et al., 26 May 2026). This suggests a natural downstream workflow: use CDTI to propose candidate hidden confounders or proxies, then assess those variables within standard adjustment, representation-learning, Shapley attribution, or causal-graphical selection pipelines (Brockschmidt et al., 11 May 2026).