SDS Seminar Series – Chad Hazlett, University of California, Los Angeles

background image
Event starts on this day

Sep

25

2026

Event starts at this time 2:00 pm – 3:00 pm
Cost: Free
In Person (view details)
The Synthetic Confounder: What Pre-Treatment Outcomes Can and Cannot Buy in Longitudinal Causal Inference

Chad Hazlett's headshot

The Fall 2026 SDS Seminar Series continues on September 25th from 2:00 p.m. to 3:00 p.m. with Chad Hazlett (Professor of Political Science and Statistics, UCLA). This event is in-person in the Avaya Room (POB 2.302).    

Title: The Synthetic Confounder: What Pre-Treatment Outcomes Can and Cannot Buy in Longitudinal Causal Inference

Abstract: Longitudinal data are attractive for causal inference about the effects of events on the units that experience them, because repeated pre-treatment outcomes seem to let units serve as their own controls and to carry information about confounders that would otherwise go unadjusted. The methods that exploit this intuition place different gambles. Difference-in-differences (DiD) and fixed-effects (FE) approaches use differencing or unit intercept shifts to remove time-invariant confounding; lagged-dependent-variable (LDV) regression and synthetic control (synth) condition on the pre-treatment outcomes to absorb whatever information they carry; and Arellano-Bond instruments with distant lags. Each is unbiased or consistent under its own demanding assumptions. We begin by showing how fragile those assumptions are under two features of the data we cannot typically rule out: past outcomes may be a cause of later outcomes ("autocausation"), and one or more past outcomes may affect the probability of treatment ("feedback"). Together these render DiD, FE, and Arellano-Bond biased and inconsistent even in the total absence of any confounder. LDV and synth survive that case, but fail once a confounder is introduced, even a time-invariant one, recovering slowly as the number of pre-treatment periods grows.

We introduce the synthetic confounder (synthconf) approach, which is consistent under autocausation and feedback whether there is no confounder, a time-invariant confounder, or a time-varying confounder representable as a rank-one signal f(t) interacting with unit-level confounding U_i through a linear factor model. Its core restriction is a "vanishing lag" assumption: the outcome at time t may directly affect outcomes up to t + ν, but not beyond. Consider the precision matrix of the observables augmented with U as though it were observed: entries for pairs of outcomes farther apart than ν must be zero. Marginalising U, the observed precision is a sparse matrix with a known zero pattern minus a rank-one term, and both parts can be recovered by a convex sparse-plus-low-rank fit, from which the treatment effect is read directly. Inference follows by the delta method, with coverage validated in simulation. We then propose a sensitivity analysis that asks how much unadjusted confounding would be needed to alter the conclusion, benchmarked against the confounding strength of the pre-treatment outcomes and of the estimated confounder path f(t).

Synthconf performs well where existing approaches are biased or inconsistent, but it has important limits of its own: confounding must be rank-one and enter through a linear factor structure, and under realistic noise calibration it can require hundreds of units before it beats the standard methods on RMSE -- a large number relative to many synthetic control applications. We therefore offer synthconf not as a replacement for the existing toolkit but as an addition to it, to be reported alongside LDV, synthetic control, and the differencing estimators in a single transparent display of which assumptions yield which estimates.

Location

POB 2.302

Share


Audience

Other Events in This Series

Oct

4

2024

Seminar Series

SDS Seminar Series – Huiyan Sang, Texas A&M University

GS-BART: Graph Split Additive Decision Trees for Spatial and Network Data

2:00 pm – 3:00 pm • In Person

Speaker(s): Huiyan Sang

Oct

11

2024

Seminar Series

SDS Seminar Series – Mingyuan Zhou, University of Texas at Austin

Building Faster, Better, and Safer Deep Generative Models via Score Identity Distillation

2:00 pm – 3:00 pm • In Person

Speaker(s): Mingyuan Zhou

Oct

18

2024

Seminar Series

SDS Seminar Series – Sherry Zhang, University of Texas at Austin

Pivoting between Space and Time: Spatio-Temporal Analysis with Cubble

2:00 pm – 3:00 pm • In Person

Speaker(s): Sherry Zhang

Oct

25

2024

Seminar Series

SDS Seminar Series – Matt Koslovsky, Colorado State University

Sparse Dirichlet-Multinomial Models

2:00 pm – 3:00 pm • In Person

Speaker(s): Matt Koslovsky

Nov

1

2024

Seminar Series

SDS Seminar Series – Aaditya Ramdas, Carnegie Mellon University

A Game-Theoretic Theory of Statistical Evidence

2:00 pm – 3:00 pm • In Person

Speaker(s): Aaditya Ramdas

Nov

8

2024

Seminar Series

SDS Seminar Series – Myungsoo Yoo, University of Texas at Austin

Dynamic Spatio-Temporal Model Integrating Physics for Fire Front Propagation

2:00 pm – 3:00 pm • In Person

Speaker(s): Myungsoo Yoo

Nov

15

2024

Seminar Series

SDS Seminar Series – Rafael Irizarry, Harvard University

Twenty-Five Years of Data Science: Music, Genomics, and Public Health Surveillance

2:00 pm – 3:00 pm • In Person

Speaker(s): Rafael Irizarry

Mar

7

2025

Seminar Series

SDS Seminar Series - Arun Kuchibhotla, Carnegie Mellon University

Adaptive Inference Techniques for Some Irregular Problems

2:00 pm – 3:00 pm • In Person

Speaker(s): Arun Kuchibhotla

Mar

28

2025

Seminar Series

SDS Seminar Series – Po-Ling Loh, University of Cambridge

Differentially Private M-estimation via Noisy Optimization

2:00 pm – 3:00 pm • In Person

Speaker(s): Po-Ling Loh

Apr

18

2025

Seminar Series

SDS Seminar Series – Richard Samworth, University of Cambridge

How Should We Do Linear Regression?

2:00 pm – 3:00 pm • In Person

Speaker(s): Richard Samworth