Chapter Four · failure evidence

What Propensity Score Matching got wrong, from 75 dissertations

Across the empirical records, propensity score matching frequently encounters practical limitations including severe data loss from unmatched observations, persistent covariate imbalance, and vulnerability to unmeasured confounding. In addition, propensity models often suffer from misspecification in high dimensional settings and are repeatedly outperformed by linear regression and alternative causal estimators. These records come from PhD theses at 26 institutions, 2021 to 2026. Each links to its thesis. They were extracted by language models reading the full text, so treat each as a lead to read, not a verdict.

Caliper enforcement and strict overlap restrictions cause severe sample loss and participant exclusion

15 theses · 12 institutions

Matching protocols and common support trimming frequently discard large proportions of treated and control participants due to caliper mismatches. This substantial sample attrition leads researchers to reject propensity score matching in favor of methods like regression adjustment or subclassification.

Considered and rejected

Considered and rejected: Decided against HDPS propensity score matching in favor of regression with overlap trimming due to severe sample loss (30-80% in matching vs 1-5% in regression).

Use of antiseizure medications during pregnancy and adverse neonatal outcomes · MSpace - University of Manitoba

Lost to a baseline

Nearest neighbor propensity score matching resulted in substantial sample loss due to unmatched cases, leading the author to switch to subclassification matching.

Three Essays On Noncognitive Factors, Friendship Networks And Education Outcomes · Penn

Lost to a baseline

562 ACDF cases excluded due to lack of suitable CDA propensity score matches (caliper mismatch in 2:1 matching)

Comparative Outcomes in Cervical Spine Surgery: A Dual Analysis of Revision Surgeries and Adjacent Segment Degeneration · Harvard

Lost to a baseline

Chapter 2.4 (Abortion quasi-experiment): Excluded control participants lacking comparable propensity score mix as well as treated women with very high propensity scores outside common support during 1:1 nearest-neighbor matching.

Parenthood and Life Satisfaction: The Consequences of Childbirth, Alternative Pregnancy Outcomes, and Single Parenting on Well-Being in Several Domains of Life · Leibniz Universität Hannover Repository

Lost to a baseline

4 of 24 ONJ cases (and unmatched osteoporosis controls from the pool of 874) could not be matched during propensity score matching, leaving 20 matched pairs.

Estado Cualitativo Y Cuantitativo Oseo Generalizado En La Osteonecrosis De Los Maxilares. Efectos De Los Bifosfonatos. · accedaCRIS

Lost to a baseline

166 patients were dropped during propensity score trimming due to lacking common support/overlap (overall sample decreased from 8,495 to 8,329 matched patients: 13 in stage I–II HR+/HER2−, 4 in stage III HR+/HER2−, 42 in stage I–II HER2+, 70 in stage III HER2+, 14 in stage I–II TNBC, and 23 in stage III TNBC)

Neoadjuvant versus adjuvant chemotherapy for older adults with stage I–III breast cancer · UT Austin

Lost to a baseline

Study 1 (Manuscript #3): 2 male outlier participants excluded during propensity score matching due to age distribution constraints.

In-Group and Out-Group Perspectives on Content of Social Groups · open_UMR Marburg DSpace 10.0

Considered and rejected

Considered and rejected: Rejected propensity score matching in favor of multivariable regression models, citing selection bias from censoring and sample size loss from unmatched subjects.

A COMPARISON OF HOSPITALIZATION OUTCOMES BETWEEN PERITONEAL DIALYSIS AND HOME HEMODIALYSIS PATIENTS IN CANADA · DalSpace

Considered and rejected

Considered and rejected: Rejected propensity score matching (PSM) because it discards unmatched control units and relies on inefficient iterative search, opting instead for entropy balancing.

To Go or Not to Go: A Quantitative Gendered Analysis of Health, Subjective Socioeconomic Status, and Well-Being Outcomes Among Rural-to-Urban Migrants in China · Harvard

Lost to a baseline

13 exposed individuals and 807 non-exposed individuals from the initial IBD cohort were excluded because they could not be matched using propensity score and disease duration calipers (PS caliper ±0.05, disease duration caliper ±0.5).

EVALUATING THE IMPACT OF THE SASKATCHEWAN INTEGRATED CARE FOR INFLAMMATORY BOWEL DISEASE ON DIRECT HEALTHCARE COSTS AND QUALITY OF LIFE · HARVEST

Lost to a baseline

278 towns trimmed from main propensity score analysis because their estimated propensity score fell outside the support of the control group (leaving 53 treated and 102 control out of 439)

Essays on the Political Economy of City Status · DukeSpace

Considered and rejected

Considered and rejected: Decided against using propensity score matching in Paper 2 due to small sample size in subgroups and covariate imbalance causing sample attrition, using multivariable logistic regression instead.

Understanding the adverse impact of centralised care on neonatal outcomes · University of Nottingham Repository

Considered and rejected

Considered and rejected: Rejected full propensity score matching in favor of propensity score covariate adjustment to avoid discarding unmatched cases in small samples.

The Role Of For-Profit Education In Social Stratification And Social Inequality In The United States · Penn

Lost to a baseline

954 treated customers dropped during 1-to-1 propensity score matching with caliper u = 0.05 (sample reduced from 8,730 to 7,776 treated customers)

ESSAYS ON DIGITAL EXPERIENCE · Cornell

Lost to a baseline

171 treated customers dropped during 1-to-1 rolling entry matching without replacement with caliper u = 0.05 (sample reduced from 977 to 806 treated customers)

ESSAYS ON DIGITAL EXPERIENCE · Cornell

Considered and rejected

Considered and rejected: Nearest-neighbor propensity score matching was rejected in favor of kernel matching due to poor common support across treatment and control groups.

Socioeconomic implications of adverse birth outcomes · Leibniz Universität Hannover Repository

High dimensional covariates and model misspecification induce estimation bias and extreme weights

13 theses · 9 institutions

Propensity score models with complex non-linear specifications, high dimensional noise features, or misclassified clustering frequently overfit and produce extreme inverse probability weights. Misspecifying the propensity or outcome model causes persistent estimation bias and degrades the precision of downstream treatment effect estimates.

Tried and failed

generalized additive models applied to propensity score estimation. Outcome: worse than baseline. Reason: did not improve over logistic regression due to high proportion of dummy variables

Disruption analytics in urban metro systems with large-scale automated data · Imperial

Tried and failed

complex non-linear models for propensity score estimation applied to fairness sample re-weighting under MAR. Outcome: worse than baseline. Reason: overfit binary indicators instead of accurately estimating continuous propensity probabilities

ADVANCING ROBUST AND FAIR STATISTICAL AND MACHINE LEARNING MODELS FOR INCOMPLETE DATA · Penn

Tried and failed

covariate matching with high dimensional noise features applied to experimental design treatment assignment. Outcome: worse than baseline. Reason: irrelevant features degrade match quality on important predictive covariates, reducing finite-sample precision

Essays on Experimental Design · MIT

Tried and failed

standard logistic regression for propensity score estimation applied to clustered data with misclassified dropouts. Reason: misclassification of cluster-level missingness induces substantial estimation bias under small cluster sizes

Statistical Methods for Addressing Missing Data and Evaluating Treatment Effect Heterogeneity in Randomized Clinical Trials · Harvard

Lost to a baseline

Logistic regression baseline had less noisy propensity score calibration than BICauseTree due to better data efficiency

Towards trustworthy AI: from local explanations to causal understanding · Oxford

Considered and rejected

Considered and rejected: Rejected propensity score matching (PSM) across full marker sets due to the curse of dimensionality, opting for linear probability model regression adjustment

Markers of a real estate agent's value-add · UT Austin

Tried and failed

naive uncalibrated inverse propensity score weighting applied to treatment effect estimation with misspecified imputations. Reason: asymptotic bias persists and does not diminish with sample size under misspecified initial outcome imputations

Causal Inference in Complex Observational Settings with Applications to Electronic Health Record Data · Harvard

Tried and failed

high-dimensional propensity score with many empirical covariates applied to observational cohort confounding adjustment. Outcome: unstable. Reason: adding over 200 empirical covariates compromised balance and created extreme inverse probability weights

Real-World Bleeding with Ibrutinib in B-Cell Malignancies · Penn

Lost to a baseline

Horvitz-Thompson and Hajec ratio estimators under network interference suffered from significantly larger standard errors/variance compared to the proposed OR and DR estimators due to sensitivity to extreme estimated propensity scores

Essays on Adaptive Methods for Inference and Prediction under Dependence · DukeSpace

Tried and failed

doubly robust propensity score estimation applied to observational pricing data. Outcome: no signal. Reason: treatment assignment had no correlation with observed features, breaking propensity score estimation

A Simulated Annealing Approach to Designing Optimal Decision Trees for Classification, Prescriptive, and Survival Analysis · MIT

Tried and failed

doubly robust nearest neighbor imputation estimator applied to missing survey data estimation. Outcome: did not generalise. Reason: both the outcome model and the propensity/probability model were misspecified simultaneously

Topics in survey design and nearest neighbor imputation for survey data · Iowa State

Tried and failed

truncating propensity score weights applied to time-varying causal effect estimation. Outcome: worse than baseline. Reason: truncating weights distorted covariate balance over sequential treatment paths, amplifying bias and root mean squared error

Toward Robust and Transparent Estimation of the Effects of Time-dependent Interventions · Harvard

Tried and failed

high-dimensional propensity score with tree-based models applied to causal treatment effect estimation. Reason: variance-governing parameters impacting treatment and outcome relationships caused large estimation bias

Machine Learning Methods for Decision Making Inference in Healthcare · Georgia Tech

Propensity score matching is outperformed by ordinary regression and alternative causal estimators

11 theses · 6 institutions

In settings with linear outcome structures or standard observational cohorts, linear regression and doubly robust alternatives achieve lower bias, lower variance, and higher statistical power than matching. Furthermore, difference in differences and regression adjustments often match or exceed the performance of propensity score methods without adding estimation complexity.

Tried and failed

propensity score matching with difference in differences applied to observational policy impact estimation. Reason: matched estimation yielded results nearly identical to standard difference in differences without added benefit

Transforming the Asian Motorcycle City? Evaluating the Travel and Urban Development Effects of the Mass Rapid Transit in Taipei, Taiwan · Penn

Tried and failed

propensity score matching for differential expression applied to single-cell perturbation gene expression detection. Outcome: worse than baseline. Reason: did not outperform standard Student's t-test and Kolmogorov-Smirnov statistical tests

Assessing regulatory function of rare and common variants using expression CROP-sequencing · Georgia Tech

Lost to a baseline

Linear regression achieved lower squared error when the outcome function was strictly linear and well-specified, outperforming matching at N=500.

Matching: The Search For Control · Penn

Lost to a baseline

Sequential matching had lower relative sample efficiency (0.717 to 0.898) compared to OLS regression adjustment under perfectly linear models (LI scenario) at small sample sizes (n=50).

Statistical Analysis and Design of Crowdsourcing Applications · Penn

Lost to a baseline

Under the linear outcome regression and linear propensity score simulation setup (OR1RM1), Hainmueller's entropy balancing achieved an SE and RMSE of 3.46 x 10^-2, outperforming the proposed DR method (3.52 x 10^-2).

Topics on nonparametric calibration, kernel ridge regression imputation\\ and nonparametric propensity score estimation · Iowa State

Lost to a baseline

When the outcome model is strictly linear in observed covariates, TSLS achieves slightly lower absolute bias of the median than the nonparametric full matching IV estimator.

Instrumental Variables and Mendelian Randomization With Invalid Instruments · Penn

Considered and rejected

Considered and rejected: Rejected propensity score matching in Chapter 2.3 in favor of fixed effects linear regressions with within-person centering to examine causal life-course trajectories rather than simply adjusting for pre-event differences.

Parenthood and Life Satisfaction: The Consequences of Childbirth, Alternative Pregnancy Outcomes, and Single Parenting on Well-Being in Several Domains of Life · Leibniz Universität Hannover Repository

Lost to a baseline

Propensity score matching and prognostic score matching showed much higher estimation bias (-42.01% and -201.33%) than BART-CV (-19.50%) on LaLonde PSID-2 observational data

CAUSAL INFERENCE FOR HIGH-STAKES DECISIONS · DukeSpace

Lost to a baseline

Propensity score adjusted conditional log rank test achieved lower statistical power than unadjusted log rank, adjusted Cox score test, and Lin-Wei robust score test across simulation settings, beating only Kong-Slud.

Instrumental Variable and Propensity Score Methods for Bias Adjustment in Non-Linear Models · Penn

Considered and rejected

Considered and rejected: Rejected propensity score matching to generate control groups because DiD handles pre-intervention mean differences and matching introduces regression-to-the-mean bias when selecting extreme initial-year utilizers.

Cost-Effective Management of Diseases: Early Detection and Interventions for Improved Health Outcomes · Georgia Tech

Considered and rejected

Considered and rejected: Rejected propensity score matching in favor of Mahalanobis nearest distance matching due to known methodological pitfalls and bias.

Evaluating the impact of conservation and development interventions on Guatemalan and Colombian forest ecosystems · OpenBU

Propensity score matching struggles with continuous treatments and complex data structures

11 theses · 10 institutions

Standard propensity score matching is poorly suited for continuous exposures, multi-group treatments, or data with high rates of missing covariates. Researchers also reject matching when fixed cohort ratios, differing follow-up times, or clustering variables prevent model convergence and valid paired evaluations.

Considered and rejected

Considered and rejected: Rejected using matched propensity score (PSMATCH) or univariate models alone for disease impact, selecting the doubly robust method for lower variance and narrower confidence intervals

Whole-herd drivers of wean-to-finish mortality under field conditions: a data-driven approach · Iowa State

Tried and failed

paired t-tests for post-matching balance evaluation applied to covariate balance assessment after matching. Reason: matching reduces within-pair variance, causing smaller mean differences to yield misleadingly large t-statistics

Contributions To Multivariate Matching In Observational Studies · Penn

Tried and failed

cohort matching at fixed unbalanced ratios applied to Markov health economic modeling. Reason: significantly longer follow-up times and unbalanced mortality between comparator cohorts prevented valid matching

The future of treatment for avascular necrosis of the femoral head: hip resurfacing arthroplasty health economics and surgical technology · Imperial

Considered and rejected

Considered and rejected: Standard propensity score matching (PSM) was rejected/avoided in favor of IPWRA and BCM due to inconsistency and bias when matching with multiple continuous covariates.

Nutritional and environmental impacts of livestock production systems in Canada: a food systems perspective · MSpace - University of Manitoba

Considered and rejected

Considered and rejected: Rejected propensity score matching for causal inference due to documented drawbacks in matching compared to inverse propensity weighting (IPW).

Causal Inference with Selection Bias and Complex Observation Data · DSpace at SUNY Buffalo

Considered and rejected

Considered and rejected: Rejected standard propensity score matching/univariate covariate bounds because they fail when covariate distributions and overlap are non-linear and multi-group

Tackling Key Challenges to Guide Clinical Decisions in Cardiovascular Diseases · MIT

Considered and rejected

Considered and rejected: Generalized boosting modeling for propensity score matching was rejected because it does not create a 1:1 control group of equal size compatible with the survival and CBA design

Cost–Benefit Analysis of the Windham School District’s Correctional CTE Program · Texas Tech

Considered and rejected

Considered and rejected: Rejected including stratification and cluster variables directly into the propensity score estimation model due to failed model convergence.

Behavioral health services for co-occurring mental health and substance use disorders under the Affordable Care Act · JScholarship

Considered and rejected

Considered and rejected: Rejected using Inverse Probability of Treatment Weighting (IPTW) propensity score matching because bivariate baseline demographic/clinical characteristics did not significantly differ between cohorts

Comparison of medication utilization patterns and healthcare resource utilization outcomes in patients with metastatic castration-resistant prostate cancer receiving novel hormonal therapies · UT Austin

Considered and rejected

Considered and rejected: Propensity score matching was rejected because it is designed for binary treatments, while maternity leave duration was continuous

Unveiling dynamics: examining the gendered association between shifting work patterns, labour market changes, and maternity leave on mental health and well-being. · Oxford

Considered and rejected

Considered and rejected: Rejected using baseline LDL, HDL, and continuous HbA1C in the propensity score model due to high EHR missingness (>30-50%), replacing them with binary measurement-presence indicators

Evaluating the Effects of Pharmaceutical Interventions, Social Policies, and Exogeneous Shocks on People's Health and Behavior · MIT

Propensity score matching fails to achieve acceptable covariate balance and increases model dependence

10 theses · 9 institutions

Applying nearest neighbor matching, multiple matches, or strict calipers often worsens balance across observed covariates and increases sensitivity to specification choices. As a result, researchers frequently abandon propensity score matching in favor of coarsened exact matching or entropy balancing.

Tried and failed

propensity score weighting on baseline covariate levels applied to correcting differential pre-treatment trends. Reason: conditioning on baseline level covariates failed to balance differential pre-treatment trajectory trends across groups

Essays in Labor and Public Economics · Cornell

Tried and failed

1-to-1 nearest neighbour propensity score matching applied to clustered observational data. Outcome: worse than baseline. Reason: dropping poor pairs or matching with replacement degraded covariate balance and increased model dependence

A modified t-test for treatment means in unreplicated classroom comparisons and Matching Methods for reducing data imbalance · Iowa State

Considered and rejected

Considered and rejected: Rejected 1:1 nearest neighbour and optimal propensity score matching due to covariate imbalance; used full matching with probit regression instead.

Effects of the Brain Derived Neurotrophic Factor Val66met Polymorphism on the Structural and Functional Architecture of the Human Brain · YorkSpace

Considered and rejected

Considered and rejected: Rejected propensity score matching and weighting due to inability to achieve adequate covariate balance across treatment and comparison groups, opting instead for baseline covariate regression adjustment in interrupted time series.

Universal Free School Meals: Implementation of the Community Eligibility Provision and Impacts on Student Nutrition, Behavior and Academic Performance · JScholarship

Considered and rejected

Considered and rejected: Rejected using a 0.2 standard deviation propensity score matching caliper because the resulting sample (n=2016) was not balanced across all predictor covariates.

Judgment and Decision-Making in the Context of Health and Law · Cornell

Considered and rejected

Considered and rejected: Rejected standard propensity score matching (PSM) due to model dependence, caliper tradeoffs, and subjective bias from uneven density distributions.

Debiased machine learning causal inference for time-varying social variables · Oxford

Considered and rejected

Considered and rejected: Rejected propensity score matching in favor of coarsened exact matching due to concerns over model dependence and covariate imbalance.

Quasi-Experimental Evaluation of Women's Re-Entry in New Jersey - Through a Black Intersectional Lens · Carleton University Institutional Repository

Considered and rejected

Considered and rejected: Rejected using Propensity Score Matching (PSM), specifically nearest neighbor, because it increased covariate imbalance compared to Coarsened Exact Matching (CEM).

The Effects of Financial and Economic Literacy on Individual Policy Preferences · ResearchWorks

Tried and failed

multi-nearest-neighbor matching with simple ATT estimator applied to causal treatment effect estimation. Outcome: worse than baseline. Reason: increasing number of matches compounded bias under poor covariate balance conditions

Modern Econometric Methods for the Analysis of Housing Markets · Virginia Tech

Considered and rejected

Considered and rejected: Propensity score matching rejected in favor of Mahalanobis distance matching to avoid covariate imbalance and model dependence.

Educational Disparities in Chronic Pain and Life Expectancy: Gaps and Pathways · DSpace at SUNY Buffalo

Propensity score methods cannot correct for unmeasured confounding or structural segregation

7 theses · 5 institutions

Matching and weighting techniques fail when key confounding factors like social networks, activity levels, or unobserved site differences remain unmeasured. In addition, propensity score adjustments break down when treatment and control cohorts are structurally separated across underlying productivity or demographic attributes.

Tried and failed

propensity score weighting and test-then-pool synthesis applied to integrating external heterogeneous control datasets. Outcome: did not generalise. Reason: unmeasured confounders and varying study-specific intercepts caused severe type I error inflation and substantial bias

Methods for the Design and Analysis of Clinical Trials: Uncertainty Directed Randomization and Data Synthesis · Harvard

Tried and failed

adding area-level covariates to propensity score models applied to controlling unmeasured confounding in observational studies. Outcome: no signal. Reason: area-level proxies lacked sufficient granularity to adjust for individual-level unmeasured confounding or shift effect estimates

Leveraging Geographic Information for Causal Inference in Pharmacoepidemiology · Harvard

Considered and rejected

Considered and rejected: Rejected Propensity Score Matching (PSM) for comparing experimental and control groups because groups were not overly imbalanced across baseline variables and PSM could not address unmeasured confounders

Treatment compliance of male perpetrators of intimate partner violence · University of Nottingham Repository

Considered and rejected

Considered and rejected: Rejected using propensity score matching due to inability to match on unobserved social network and information campaign variables.

Analyzing the drivers of agricultural technology adoption among smallholder farmers in Uganda · Oxford

Considered and rejected

Considered and rejected: Rejected propensity score matching because retrospective covariates were limited and omitted-variable bias could not be addressed without an endogenous treatment model.

Unequal starts: the role of different learning environments in the development of inequalities in skills during early childhood · IRIS - UNITN - prod

Considered and rejected

Considered and rejected: Decided against propensity score matching because only summary/aggregate data were provided by registries and unmeasured confounders like activity level remained.

Do dual-mobility cups reduce revision risk in femoral neck fractures compared with conventional tha designs? An international meta-analysis of arthroplasty registries · Oxford

Considered and rejected

Considered and rejected: Rejected using Propensity Score Matching (PSM) for evaluating co-op vs IOF price impacts because PSM fails when participant groups are structurally segregated by productivity.

THREE ESSAYS ON THE ECONOMICS OF SMALLHOLDER AGRICULTURE IN AFRICA: LABOUR ALLOCATION, ADOPTION OF CONSERVATION AGRICULTURE AND CO-OPERATIVE PARTICIPATION · HARVEST

Left open by the authors

Problems the authors named and did not get to.

Left open

Evaluate machine learning classification methods against logistic regression for propensity score estimation with imbalanced group sizes in marginal structural models. Blocker: None

An inspection of measurement error inherited in time-varying latent covariate proxies within marginal structural models · UT Austin

Left open

Log-transform outcome variables and evaluate childless versus mother comparisons via propensity score matching on merged PSID and O*NET data. Blocker: None

Impact of Occupational Flexibility on Labor Market Outcomes of Women Following Childbirth · MIT

Left open

Apply matching methods or propensity score estimation to the survival model dataset to control for strategic selection in leader tenure and concessions. Blocker: None

The Causes and Consequences of Territorial Nationalism · Harvard

Left open

Develop average treatment effect estimation tools for observational studies using kernel ridge regression-based propensity score weighting. Blocker: None

Topics on nonparametric calibration, kernel ridge regression imputation\\ and nonparametric propensity score estimation · Iowa State

Left open

Compare standard propensity score weighting with entropy balancing, matching, and stratification methods in added-growth latent growth models via simulation. Blocker: None

Treatment effect estimation using latent growth models and propensity score weights with heterogeneous growth trajectory shapes · UT Austin

Left open

Implement propensity score matching or multivariate clustering on public Census tract data to select control sites for TOD comparative analysis. Blocker: None

NOT JUST TOD: WHOSE GROWTH? WHOSE NEIGHBORHOOD? – Ethnic-Majority Cohesion, Economic Transformation, and Gentrification Resistance in Hispanic/Latino-Majority Transit-Oriented Neig · Cornell

Left open

Develop and evaluate methods using large language models for propensity score estimation and causal inference on observational data. Blocker: None

Language Models as Opinion Models: Techniques and Applications · MIT

Left open

Implement propensity score matching, stratification, machine learning estimators, and trimming methods for latent variable outcome models in lavaan simulation studies. Blocker: None

Latent variable outcome analysis models in a propensity score framework · UT Austin

Left open

Develop methodology to quantify residual and unmeasured confounding for continuous dose matching and subclassification. Blocker: Lack of concrete statistical approach or mathematical framework specified in the thesis

Observational Data With A Continuous Exposure: Study Design And Outcome Analysis · Penn

Left open

Apply propensity score matching to balance confounding factors when evaluating hospital nursing resources and surgical patient outcomes. Blocker: Access to restricted hospital discharge databases and nursing survey data

ASSOCIATIONS BETWEEN HOSPITAL NURSING RESOURCES AND THE OUTCOMES OF SURGICAL PATIENTS WITH PROLONGED SURGICAL TIME · Penn

Checking a claim in this area?

We can run the same search on any method or claim. If nothing turns up, we will say so, and that proves nothing on its own.