Theses Doctoral

Deep Learning on Low-Dimensional Manifolds: From curve classification to manifold denoising

Wang, Tingran

Deep learning has achieved remarkable success across a wide range of domains, including natural language processing, computer vision, and multimodal tasks. Neural networks now tackle increasingly complex challenges, such as generating coherent text, producing photorealistic images from descriptions, and integrating information across multi-modalities, redefining the boundaries of artificial intelligence. These achievements are driven by two pivotal factors: the scaling of model sizes and datasets to unprecedented levels, and the development of novel architectures that capitalize on the inherent structure of data. Leveraging massive datasets and sophisticated training techniques, modern neural networks have attained levels of expressiveness and adaptability that were once inconceivable.

Despite these advances, fundamental questions remain about the principles governing the success of deep learning models. What allows these networks to excel at generalizing across diverse and noisy datasets? How do neural networks adapt to the geometric and statistical properties of real-world data? A central hypothesis in this domain is the manifold assumption: that high-dimensional datasets often reside on or near low-dimensional manifolds. This insight suggests that data, while appearing complex, may possess an intrinsic structure that neural networks can exploit to learn efficiently. Understanding these structures is key to uncovering how neural networks navigate the vast search space of high-dimensional data representations.

This thesis investigates the interplay between neural networks and structured data, particularly datasets with low intrinsic dimensions and manifold structures. While high-dimensional data is challenging to model due to noise and variability, many natural datasets—such as images and scientific measurements—exhibit geometric constraints that define their effective dimensionality. These manifold-like structures provide a simplified yet powerful framework for understanding the principles of learning, making them an ideal foundation for theoretical and empirical studies. Through this lens, we examine how network architecture, data structure, and resource allocation shape the performance of neural networks, bridging the gap between foundational insights and practical applications.

We begin by exploring the theoretical underpinnings of neural network performance on data that lies on low-dimensional manifolds. Using curve classification as a case study, we provide provable guarantees on how network architecture, depth, width, and training sample size contribute to successful learning. These results illuminate critical dependencies between data geometry and network design, offering insights into the theoretical capabilities of neural networks.

Building on this theoretical foundation, we turn to empirical investigations, focusing on the denoising problem. Data residing on a manifold is often corrupted by noise, which distorts its intrinsic structure and poses a significant challenge for recovery. To enable controlled empirical study of this problem, we develop a principled method for generating synthetic manifolds with rigorously characterized geometric properties using Gaussian processes—a contribution that addresses the lack of datasets with known geometric structure in the field. Through systematic experiments using these generated manifolds, we evaluate the ability of neural networks to restore clean data from noisy observations. We analyze how network design choices, including depth, width, and sample size, influence denoising performance. Our findings demonstrate that well-designed networks can effectively recover manifold-structured data while preserving its geometric properties, shedding light on the practical considerations for network design in such tasks.

By combining theoretical and empirical perspectives, this work bridges the gap between foundational insights and real-world applications, offering a comprehensive understanding of how neural networks interact with low-dimensional manifold data. It advances the broader understanding of when and why neural networks succeed, emphasizing the critical roles of data structure and network design in achieving robust and efficient learning. Through this focused exploration, we aim to lay a foundational framework for addressing the complexities of deep learning on structured data.

Files

  • thumbnail for gsas-dissertations-000579.pdf gsas-dissertations-000579.pdf application/pdf 2.14 MB Download File

More About This Work

Academic Units
Electrical Engineering
Thesis Advisors
Wright, John N.
Degree
Ph.D., Columbia University
Published Here
August 26, 2026

Notes

Deep Learning, Machine Learning, Optimization, Gaussian processes, Manifold