Theses Doctoral

Coarse Graphical Representations for Causal Inference with High-Dimensional Clinical Data

Anand, Tara Vafai

In biomedical informatics, learning about causal relationships can impact the field and practice of medicine through informing clinical decision-making, health policy, and patient care. The gold standard for determining causal relationships is often said to be a randomized controlled trial, yet such experiments are often too expensive or unethical to conduct in practice and results often suffer from non-generalizability. Observational data, such as electronic health record (EHR) data, hold promise as an alternative source for learning about causality in biomedicine, though many obstacles remain to addressing various forms of bias and data quality challenges. It is well known that causal inferences cannot be determined from observational data alone and require the addition of knowledge in the form of defined assumptions. Popular causal analytical approaches are attached to a rigid set of assumptions, yet causal diagrams, a form for encoding assumptions about the data-generating process, can motivate flexible analytical approaches while ensuring transparency about assumptions made. In practice, the construction of fully specified causal diagrams is challenging in complex, high-dimensional domains like medicine, limiting the adoption of graphical causal
methods in informatics.

This dissertation addresses the challenge of causal diagram construction by introducing and developing a framework for coarse or cluster causal diagrams, where the knowledge requirements for graph construction are relaxed by allowing for nodes that represent sets or clusters of variables rather than individual variables. Partial knowledge of the causal and confounding relationships can then be represented graphically, without committing to uncertain or unknown assumptions at the variable level, and still enabling identification of causal effects and sound analysis. The work makes contributions across theory, methodology, and application.

First, this dissertation develops the theoretical foundations of cluster causal diagrams (C-DAGs), a graphical formalism that represents relationships among clusters of variables while leaving intra-cluster structure unspecified. I show how C-DAGs can allow for a sparser encoding of knowledge, representing a class of causal diagrams sharing this knowledge, and I establish their formal semantics. In particular, I prove that d-separation is sound and complete for probabilistic inference over clusters, that Pearl’s do-calculus rules are valid for causal identification using C-DAGs, that the standard identification algorithm remains sound and complete in this setting, and that counterfactual reasoning can be performed at the cluster level. Together, these results demonstrate that C-DAGs support valid inference at all layers of Pearl’s Causal Hierarchy.

Second, this dissertation empirically evaluates the practical use of coarse causal diagrams for causal effect estimation. Using simulated clinical data, I examine how different cluster causal diagram constructions, with varying degrees of coarseness, affect identification and accuracy of estimated causal effects. I then apply the framework to real-world observational data to study the comparative safety and effectiveness of antihypertensive and antihyperglycemic medications. These case studies illustrate how explicit articulation of assumptions through coarse causal diagrams can clarify the validity of popular approaches such as large-scale propensity score adjustment, reveal sensitivity to cohort definitions, and enable alternative identification strategies when standard assumptions are violated.

Finally, this dissertation addresses the challenge of learning coarse causal diagrams from data when prior knowledge is insufficient for their construction by knowledge. I introduce novel graphical equivalence classes for cluster-level causal discovery in both Markovian and non-Markovian systems. I develop sound causal discovery algorithms that recover learnable causal relationships among clusters of variables from observational data, extending classical constraint-based discovery methods to this abstracted setting.

Taken together, this work advances the theory and practice of graphical causal inference in biomedical informatics by providing principled tools for constructing, reasoning with, estimating from, and learning causal diagrams in high-dimensional domains. By bridging formal causal inference theory with applied clinical analysis, this dissertation aims to make causal graphical approaches more feasible, transparent, and impactful in real-world informatics applications.

Files

  • thumbnail for gsas-dissertations-000289.pdf gsas-dissertations-000289.pdf application/pdf 1.52 MB Download File

More About This Work

Academic Units
Biomedical Informatics
Thesis Advisors
Hripcsak, George
Degree
Ph.D., Columbia University
Published Here
June 17, 2026

Notes

Causality, Artificial Intelligence, Computer Science, Medical Informatics, Graph Theory