Theses Doctoral

Causal Generative Modeling: Theory and Practice

Xia, Kevin M.

Causal inference is a fundamental field of study integral to developments in philosophy, statistics, econometrics, natural sciences, and, more recently, machine learning and the data sciences. A defining feature of human intelligence is the ability to reason about cause and effect. A causal understanding of reality provides explanations for outcomes that help determine responsibility, implement policies, remove bias, and generalize to broader settings. Modern AI systems such as chatbots and art generators are built on a foundation of generative models, capable of learning the data distribution so well that they can reproduce novel but realistic samples. Nonetheless, most of these models only learn correlational patterns rather than understanding the causal reasoning behind why those patterns exist. While these two fields of study emerged separately, it is expected that, like humans, a truly intelligent AI cannot only produce the powerful results of deep learning models, they are able to do so with a causal understanding of the underlying mechanisms of reality.

This dissertation moves towards a unification of these two fields under the emerging topic of causal generative modeling. The first part of the thesis covers the theoretical foundations of what makes generative models causal, including the tension between the expressiveness of deep learning models and the necessity of causal assumptions and constraints for inference. A novel class of causal generative models called the Neural Causal Model (NCM) is introduced, which serves as the backbone for the theoretical developments and implementation realizations in this work. The NCM is shown to have the capabilities to incorporate causal constraints as an inductive bias and to identify, sample, and estimate interventional and counterfactual quantities in a generative manner. The NCM framework generalizes prior developments of causal generative models in many directions such as expressiveness of causal queries, capability of incorporating a variety of causal constraints, ability to address unobserved confounding, ability to leverage multiple data inputs, proven theoretical guarantees, and algorithmic results on solving classical causal tasks.

The second part of the thesis dives deep into causal abstraction theory for the purposes of improving the performance of causal generative models in high-dimensional settings like with image data. While training models for such settings can be challenging, low-level data (e.g., pixels) can often be abstracted into high-level concepts (e.g., image of dog). A new framework of abstractions is formalized that allows direct comparisons between the quantities induced by different causal models that are, in reality, models describing the same variables at different granularities. This enables the capability of performing causal inferences across granularities, allowing one to work in a more abstract space than what is provided in the data. This concept of working in an abstract space is connected with the machine learning concept of representation learning. NCMs are then shown to leverage abstraction theory to incorporate representation learning aspects and demonstrate applicability in scalable causal generative modeling. These results broadly contribute to bridging the gap between the causal capabilities of human reasoning and the celebrated applications of deep learning models.

Files

  • thumbnail for gsas-dissertations-000268.pdf gsas-dissertations-000268.pdf application/pdf 4.03 MB Download File

More About This Work

Academic Units
Computer Science
Thesis Advisors
Bareinboim, Elias
Degree
Ph.D., Columbia University
Published Here
June 17, 2026

Notes

causation, machine learning, deep learning, causal inference, generative modeling