Theses Doctoral

Aligning Words With More Human Worlds

Wang, Sky C.

This thesis begins from a simple premise—that building "better" language technologies first requires a better understanding of the humans they are designed to serve. This means attending not only to what people say, but to how they reason, feel, and differ from one another across contexts. It means recognizing that human judgments are nuanced, sometimes inconsistent, and often shaped by values that are implicit rather than stated. And it means designing methods that can uncover and incorporate these complexities, so that the language models that underlie these technologies do not simply mimic surface patterns, but reason in ways that are grounded in the lived realities of their users.

We begin with measurement. In a series of case studies using text as data, we explore what language reveals about psychological and social human behavior. First, we examine the structure of emotional responses to music: by modeling listener reactions at scale, we uncover the latent perceptual and affective factors that shape why music feels the way it does—advancing a long-standing measurement challenge in affective science. Next, we turn to cross-cultural measurement. Through a comparative analysis of social norms in Chinese and American cultures, we align and interpret judgments of everyday actions across these cultures, identifying not only where norms diverge but also elucidating the underlying reasons underpinning those divergences. We then demonstrate the critical role of qualitative research in surfacing gaps invisible to large-scale data. Through in-depth interviews with Deaf and hard-of-hearing (DHH) users, we examine failures in automatic music captioning—technology designed to assist, but often misaligned with the needs of those it claims to serve. These insights expose the assumptions embedded in such systems and motivate concrete design changes that reflect the real-world preferences of traditionally marginalized users. Finally, we revisit how we gather and structure human feedback. Standard approaches to model alignment most commonly rely on binary comparisons, which flatten nuanced human judgments. We develop a fine-grained feedback framework in which annotators highlight specific liked or disliked spans of a response and explain why. This approach captures richer supervision signals at little additional annotation overhead—and sets the stage for more sample-efficient, human-centered alignment in language model post-training.

The second half of the thesis turns from understanding humans to incorporating that understanding into model design. That is: once we have better characterized the nature of human behavior, how can we build systems that more faithfully reflect and respond to that understanding? We pursue this through three complementary modes of steering. First, we explore norm-conditioned prompting, where we use cultural norm data to guide large language models toward more realistic and contextually grounded outputs. Rather than treat cultural variation as noise, we design prompting strategies that foreground it—encouraging models to generate conversations that reflect different normative perspectives. This enables the controlled generation of social behaviors and norms in conversation, and allows researchers to simulate and interrogate how norms may play out across these conversational contexts. Second, we turn to supervised fine-tuning (SFT) to train models that directly encode contextually grounded behavior. We demonstrate this in two settings: first, in training models to perform cross-cultural norm entailment, where the model must reason about when two norms align or conflict and explain why—a task that explicitly tests a model's ability to engage with cultural nuance and value tradeoffs; and second, in designing a music captioning system for DHH users, directly informed by qualitative studies. This closes the loop from user-centered research to system design, embedding domain-specific knowledge of user needs into model behavior.

Finally, we investigate direct alignment—approaches for aligning model behavior with fine-grained human feedback without requiring reward modeling or reinforcement learning. We begin by developing interpretable linear probes on model hidden states, showing that models implicitly encode rich internal signals about their own behavior, including whether a generation is hallucinatory. These probes allow us to localize attributes such as hallucination, verbosity, or vagueness at the span level. We then propose Lamarckian Preference Tuning, using these span-level signals to regenerate improved model outputs—transforming localized edits into global preference pairs. This technique enables data-efficient direct preference optimization, leveraging internal model knowledge and user feedback alike to steer behavior in interpretable, controllable ways.

Together, these contributions chart a path toward building language technologies that are not only powerful, but meaningfully responsive to the people who use them.

Files

  • thumbnail for gsas-dissertations-000107.pdf gsas-dissertations-000107.pdf application/pdf 2.98 MB Download File

More About This Work

Academic Units
Computer Science
Thesis Advisors
Muresan, Smaranda
Degree
Ph.D., Columbia University
Published Here
May 13, 2026

Notes

Natural language processing (Computer science), Artificial intelligence, Machine learning, Sociology

Additional thesis advisor(s): Yu, Zhou