PaperScope
LIVE · 2026-09-17 05:40 UTC

When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments

Takes Fujita, Nobutaka Hattori

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.17772 v1
Category
Submitted
2026-09-15

Abstract

AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable. We formulate a causal type discipline for sequential experiments: a versioned representation map, a causal role classifier, a claim-status filter, and an estimand lock. The lock fixes a standardized proximal effect before generated covariates enter the analysis. Under audit correctness and standard identification assumptions, admissible role assignments preserve this estimand. We apply the established conditional-covariance characterization of compression bias to substitution of generated representations for design-relevant states. A standardized decomposition separates compression, conditional-law, and standardization drift. Further results cover mediator adjustment, post-action leakage, marker-intervention conflation, outcome-guided discovery, and state-measurement error. Cluster-level orthogonal estimators distinguish empirical and superpopulation targets under repeated sessions and missing outcomes. Simulations show that refinement helps when it retains design-relevant information, whereas design erasure, leakage, and same-data marker selection can produce bias or undercoverage. The framework places causal semantics and claim status before confirmatory inference with generated representations.

Comment: 29 pages. Ancillary files include simulation code, seeds, and replicate-level results

arXiv abs page · PDF · same-day batch