2026 Theses Doctoral
Learning graphical models with intractable likelihoods
This dissertation presents theoretical and methodological developments for a range of statistical problems that involve graphical structures. Such graphical structures naturally arise by introducing dependent latent variables and modeling conditional independence structures. However, the resulting complexity often leads to a likelihood involving an intractable normalizing constant. As this hinders model estimation as well as even understanding fundamental identifiability, it is important to develop a theoretical understanding for such graphical models. We answer this general question by focusing on three popular statistical problems that allow both supervised and unsupervised settings.
First, we consider the problem of conducting frequentist inference under Gaussian mixture models with dependent latent labels. Despite its popularity, most theoretical research for this problem assumes that the labels are either independent and identically distributed, or follow a Markov chain. It remains unclear how the fundamental limits of estimation change under more complex dependence. Here, under the spherical two-component Gaussian mixture model, we first show that a naive estimator based on misspecified likelihood can be utilized for labels with arbitrary dependence. Additionally, under labels that follow an Ising model, we establish the information-theoretic limitations for estimation, and discover an interesting phase transition as dependence becomes stronger.
Second, we consider another unsupervised learning problem and propose a complex model with a deep latent structure that consists of multiple binary latent variables. This model is motivated by deep generative models in machine learning (such as Deep Belief Networks) that exhibit a similar deep architecture with layer-wise graphical dependence. The proposed model, termed Deep Discrete Encoders, enjoys additional statistical guarantees including identifiability and interpretability. Theoretically, we propose transparent identifiability conditions for the proposed model, which imply progressively smaller sizes of the latent layers as they go deeper. Identifiability ensures consistent parameter estimation and inspires an interpretable design of the deep architecture. Computationally, we propose a scalable estimation pipeline of a layer-wise nonlinear spectral initialization followed by a penalized stochastic approximation EM algorithm. We illustrate the proposed method via conducting simulation studies and illustrations on three real datasets to perform hierarchical topic modeling, image representation learning, and educational response time modeling.
Finally, we consider the problem of establishing limit distributions in a high-dimensional hierarchical Bayesian linear regression problem. Unlike the existing literature that focuses on posterior contraction, we work under a non-contracting regime where neither the likelihood nor the prior dominates the other. This is motivated by modern high-dimensional datasets with a bounded signal-to-noise ratio. Here, we take a first step towards understanding limit distributions for one-dimensional projections of the posterior, as well as the posterior mean, in such regimes. Analogous to contractive settings, the resulting limiting distributions are Gaussian, but they heavily depend on the chosen prior and center around the Mean-Field approximation of the posterior. We study two concrete models of interest to illustrate this phenomenon — the white noise design, and the (misspecified) Bayesian model. As an application, we construct credible intervals and compute their coverage probability under any misspecified prior.
Subjects
Files
-
gsas-dissertations-000574.pdf
application/pdf
2.07 MB
Download File
More About This Work
- Academic Units
- Statistics
- Thesis Advisors
- Gu, Yuqi
- Degree
- Ph.D., Columbia University
- Published Here
- June 24, 2026
Notes
Statistics, Probability, Machine Learning
Additional thesis advisor(s): Mukherjee, Sumit