Theses Doctoral

Understanding Contextual and Hierarchical Auditory Processing in the Human Brain through Deep Neural Networks

Mischler, Gavin

Humans rely heavily on our auditory capabilities to communicate, and the brain has a fascinating ability to transform continuous streams of sound into meaningful representations of speech and music. This is done through a hierarchical processing pathway that spans from the primary auditory cortex to higher-order language and association areas. At each stage of this hierarchy, the brain integrates incoming acoustic information with contextual cues to construct robust and rich perceptual experiences. Despite decades of research characterizing the neural correlates of auditory processing, the computational principles that govern how the brain encodes context across different timescales and domains remain poorly understood. A central barrier has been the lack of computational models with sufficient capacity to capture the nonlinear, hierarchical, and context-dependent transformations performed by the auditory system.

In this dissertation, we leverage deep neural networks (DNNs) as computational models to investigate how the human brain processes speech and music in context. We combine high-resolution neural recordings, including intracranial electroencephalography (iEEG) from neurosurgical patients and scalp electroencephalography (EEG), with deep learning models to probe the neural encoding of contextual auditory information across the cortical hierarchy.

First, we seek to understand how the brain deals with dynamic acoustic environments, where the surrounding context is rapidly changing. We use DNNs to model neural adaptation to sudden changes in background noise during speech listening. We show that DNNs significantly outperform classical linear models at predicting neural responses and reveal that the models achieve noise-robust encoding through a combination of adaptive gain control and spectro-temporal noise filtering, with distinct filtering mechanisms emerging along the cortical hierarchy.

Second, we investigate the alignment between large language models (LLMs) and the brain's language-processing pathway. By examining a diverse set of LLMs with similar parameter counts, we find that as model performance on language understanding benchmarks improves, LLMs become more predictive of neural responses and converge toward the hierarchical feature extraction pathways of the brain, with contextual encoding emerging as a critical factor driving this alignment.

Third, we examine how the brain encodes the hierarchical structure and context of music, and how this encoding is shaped by musical expertise. Using representations from a generative music transformer, we find that deeper, more contextual model layers are more predictive of neural responses, with this correspondence significantly enhanced in expert musicians and lateralized to the left hemisphere. Intracranial recordings further reveal an anatomical gradient, with sites progressively farther from primary auditory cortex encoding musical context more strongly.

Together, these studies demonstrate that deep learning models, from task-specific DNNs to large language models and generative music transformers, provide a powerful lens for understanding how the human brain processes auditory information in context. Our findings reveal shared computational principles across speech and music processing, including hierarchical feature extraction, contextual integration, and adaptive encoding, and point toward a convergence between artificial and biological systems in how they represent complex auditory stimuli.

Files

  • thumbnail for gsas-dissertations-000439.pdf gsas-dissertations-000439.pdf application/pdf 3.6 MB Download File

More About This Work

Academic Units
Electrical Engineering
Thesis Advisors
Mesgarani, Nima
Degree
Ph.D., Columbia University
Published Here
August 5, 2026

Notes

Neurosciences, Computational neuroscience, Neural networks (Neurobiology), Deep learning (Machine learning), Electrophysiology