Theses Doctoral

Learning representations under data heterogeneity: Scalable self-supervised models for generalizable neural encoding and decoding

Zhang, Yizi

Large neuroscience and animal behavior datasets are increasingly heterogeneous, enabling us to learn the bidirectional relationship between the brain and behavior under diverse experimental conditions, modalities, and tasks. However, such heterogeneity also makes it challenging to learn representations that generalize beyond a single dataset or setting. This thesis develops methodologies to address these challenges, with a focus on neural encoding and decoding and their applications in brain–computer interfaces (BCIs).

The first part of the thesis focuses on classical probabilistic and statistical approaches for predicting animal behavior from neural activity recorded using high-density Neuropixels probes. We propose two statistical models for neural decoding. The first utilizes a Gaussian mixture model in which mixture assignments are conditioned on the behavior of interest and allowed to vary over time. The second imposes a low-rank constraint on the weight matrix of a linear regression model to capture shared spatial and temporal firing patterns across neural data from multiple recording sessions. While these methods provide interpretability and statistical grounding, they remain limited in important ways. In particular, they are either restricted to modeling data from a single recording session or rely on inductive biases that fail to capture the full variability present in large, heterogeneous neural datasets.

To overcome these limitations, the second part of this thesis develops scalable deep learning approaches, in particular, transformers, for modeling large, heterogeneous neural datasets. Unlike classical models with fixed inductive biases, transformers provide a flexible function class capable of capturing data variability across sessions, modalities, and experimental conditions while learning shared, invariant latent structure that generalizes across datasets. We introduce a unified self-supervised training framework for extracting robust and transferable structure from diverse, largely unlabeled neural data. By leveraging large-scale pretraining across heterogeneous data sources, the resulting models achieve improved sample efficiency and stronger cross-dataset generalization. Through several case studies, we demonstrate that self-supervised pretraining consistently yields representations that outperform task-specific models when adapting to new datasets with minimal supervision.

The preceding parts focus on neural representation learning under extreme data heterogeneity. However, understanding the relationship between the brain and behavior also requires robust representations of complex behaviors themselves. To this end, we propose a self-supervised framework that pretrains vision transformers on large-scale, unlabeled video data capturing animal behavior. We show improved performance in extracting behavioral features that exhibit stronger correlation with neural activity in downstream analyses. In a separate line of work, we extend our framework to language, which can also be viewed as a form of behavior, and utilize existing pretrained large language models (LLMs) that encode language priors. By combining transformers pretrained on cross-species neural data (human and non-human primates) with pretrained LLMs, we translate neural activity into text, with the goal of restoring speech for individuals with paralysis. Together, we demonstrate that large-scale self-supervised learning provides a powerful framework for modeling variability in both neural activity and behavior, enabling more generalizable neural encoding, decoding, and their applications in BCIs.

Files

  • thumbnail for gsas-dissertations-000374.pdf gsas-dissertations-000374.pdf application/pdf 14.3 MB Download File

More About This Work

Academic Units
Statistics
Thesis Advisors
Paninski, Liam
Degree
Ph.D., Columbia University
Published Here
July 15, 2026

Notes

Computational Neuroscience, Foundation Models, Brain-Computer Interfaces