2026 Theses Doctoral
Generative AI Methods for Estimating How Unstructured Treatments Affect Customer Behavior
Unstructured data—such as images, text, audio, and video—are pervasive in marketing and increasingly central to understanding customer behavior. From product packaging and advertisements to social media content and headshots, unstructured data influence managerially relevant outcomes like sales, click-through rates, and brand attitudes. Yet estimating the causal effects of features derived from unstructured data on behavior remains a fundamental challenge, because such treatments are latent, high-dimensional, and often confounded with other latent features.
This dissertation develops and applies novel methodological frameworks based on generative AI to address this challenge. Each chapter leverages deep generative models—Generative Adversarial Networks and diffusion models—in combination with tools from causal inference and interpretable machine learning to estimate how latent features derived from unstructured data affect customer behavior. Together, the two chapters demonstrate complementary approaches: one focused on testing specific causal hypotheses about a single latent attribute, and the other on scalably discovering a taxonomy of predictive features.
In Chapter 1, we document an understudied form of discrimination based on femininity expressed in headshots. To do so, we use a Generative Adversarial Network to create realistic headshots of people and then manipulate their femininity independently of other attributes like pose, general facial expression, background, hairstyle, and clothes. Then, we devise an experimental design that allows us to identify the separate and combined causal impact of femininity and gender identity (proxied by gender pronouns) on real-life outcomes. In a field experiment within a naturalistic advertising environment, we find that prospective customers of an education company discriminated against femininity at various stages of the purchase funnel, to a largely similar extent for men, women, and non-binary people.
The findings inform managers and policymakers of an important form of discrimination that would otherwise be underestimated by disregarding headshots. Methodologically, we introduce a novel framework for testing specific hypotheses pertaining to causal effects of treatments which are derived from unstructured data. The approach involves using natural text to readily identify hypothesis-relevant features from a generative AI model (in a particularly disentangled and interpretable representation space) and then creating realistic stimuli that vary controllably in those features.
In Chapter 2, I develop a novel methodological framework to automatically discover what interpretable features make visual designs in a given domain successful. I first leverage a deep generative text-to-image AI model that adopts the role of designer and enables visual designs to be described by low-dimensional design representations. Then, I apply a novel adaptation of cutting-edge “mechanistic interpretability” methods—specifically “sparse autoencoders” typically applied to large language models—to scalably discover a taxonomy of interpretable and managerially relevant features predictive of success from these design representations.
Finally, I generate image redesigns by manipulating features of interest to help managers scalably pilot data-driven design changes. I apply this framework to discover how book cover redesigns predict sales on Amazon.com using a unique dataset I collected of over 160,000 books. I discover a diverse set of interpretable features related to illustration, typography, composition, and layout. I then create realistic cover redesigns predicted to improve sales by manipulating those features (e.g., redesigns with lower contrast and less separation of text and graphical elements). In a holdout analysis with a rich set of control variables, including just 30 of these discovered features (out of 9,728) improves variation explained in sales by nearly as much as prices and by more than reviews. Back-of-the-envelope calculations suggest that a large publisher could leverage this subset of features to increase annual revenue for the whole publisher by over $9.1 million, reflecting a change in sales equivalent to introducing an 8.5% price discount. In a lab study, I find causal evidence that the proposed methodological framework can redesign covers to significantly improve preferences, and that generative AI can help level the playing field in the publishing industry.
Subjects
Files
This item is currently under embargo. It will be available starting 2031-04-20.
More About This Work
- Academic Units
- Business
- Thesis Advisors
- Toubia, Olivier M.
- Degree
- Ph.D., Columbia University
- Published Here
- August 12, 2026
Notes
Marketing, Economics, Machine Learning, Generative AI, Causal Inference