Theses Doctoral

Statistical Foundations for Diagnostics and Efficiency in Data-Driven Optimization

Wang, Tianyu

Modern machine learning tools and data-driven algorithms offer tremendous opportunities for decision-making. However, their real-world deployment remains limited without a principled framework that integrates learning algorithms with downstream decisions tailored to specific problem instances. This gap raises several fundamental questions: How should we evaluate and select among competing learning algorithms under decision-focused objectives? How can we design optimization methods that remain reliable when data are limited, noisy, or subject to change?

This thesis aims to bridge this divide by developing a statistical foundation for evaluating and designing data-driven optimization methods. It is organized around two central themes. The first part of the thesis focuses on diagnostics, developing principled tools and empirical benchmarks to assess the performance of data-driven decision methods, guide method selection, and diagnose failure modes under limited data. The first part of the thesis (Chapters 2-4) focuses on developing principled tools and empirical benchmarks to assess the performance of data-driven decision methods, guide method selection, and diagnose failure modes under limited data. The second part (Chapters 5-6) focuses on developing optimization algorithms that are not only theoretically grounded but also computationally scalable and tailored to the structure of real-world problems.

The first half of the thesis focuses on the diagnostics of the performance of data-driven optimization methods. Standard model evaluation tools from machine learning provide little insight into how different data-driven decision-making algorithms will actually perform in operational tasks, since the ultimate goal in various high-stakes decision scenarios is not to generate accurate predictions. Moreover, we lack empirical benchmarks to understand what algorithm components work best in practice. In Chapter 2, we develop a new framework called Optimizer’s Information Criterion (OIC). OIC provides a rigorous and computationally efficient way to compare a wide range of data-driven optimization methods based on their decision quality rather than predictive accuracy. In Chapter 3, we further extended this work to interval estimates, exploring which interval construction approaches most effectively reflect decision reliability with high confidence. These tools allow practitioners to assess the reliability of a method under varying data conditions. In Chapter 4, we advance the evaluation foundation of data-driven optimization algorithms with applications in trustworthy machine learning problems to provide systematic diagnostics of core design choices. These efforts include designing the first comprehensive open-source software package implementing robust optimization methods (dro) and the first large-scale benchmark for evaluating these methods (WhyShift). Together, my work provides rigorous theory and practical tools for evaluating and deploying data-driven decision-making algorithms.

The second half of the thesis is devoted to designing efficient optimization methods. While diagnostics guide us in selection, it is arguably of even more primitive importance to design data-driven methods that perform well across diverse conditions, such as complex cost structures, limited or noisy data, and changing environments. Specifically, we demonstrate how problem-specific structure can be exploited to improve optimization algorithm design. The performance of data-driven optimization methods varies considerably across problem instances, driven by heterogeneity in both data distributions and cost functions, which makes it essential to develop instance-specific approaches that are efficient to both the problem structure and the data environment. In this part of the thesis, we study algorithms tailored to such structures and show that they can achieve significant statistical and computational gains over general-purpose methods, both asymptotically and in finite-sample regimes. In Chapter 5, we leverage side information from data-driven optimization problems to help understand when and how methods provide measurable efficiency gains compared with the empirical optimization solution. In Chapter 6, we employ the parametric approximations of data distributions to create new effective formulations of distributionally robust optimization that remain tractable and reliable when cost functions are highly complex.

Files

  • thumbnail for gsas-dissertations-000570.pdf gsas-dissertations-000570.pdf application/pdf 17 MB Download File

More About This Work

Academic Units
Industrial Engineering and Operations Research
Thesis Advisors
Iyengar, Garud N.
Degree
Ph.D., Columbia University
Published Here
June 24, 2026

Notes

Operations research, Machine learning, Mathematical optimization, Artificial intelligence, Mathematical models--Evaluation

Additional thesis advisor(s): Lam, Henry K.