Essay

The Representation Crisis - How LLM-Based Synthetic Users Obscure Rather Than Illuminate User Understanding

The proliferation of LLM-generated synthetic users in design and research creates a fundamental crisis of representation that undermines the very purpose of user-centered design. This analysis exposes the clarity deficit

The proliferation of LLM-generated synthetic users in design and research creates a fundamental crisis of representation that undermines the very purpose of user-centered design. This analysis exposes the clarity deficit inherent in synthetic user generation and its profound implications for design validity.

The rapid adoption of Large Language Model-generated synthetic users represents one of the most profound methodological shifts in user-centered design since the emergence of digital interfaces. Proponents herald these systems as democratizing user research, reducing costs, and accelerating design cycles. Critics dismiss them as shallow approximations that cannot capture the complexity of human experience. Both perspectives miss the fundamental issue: synthetic users create a crisis of representational clarity that undermines the epistemological foundations of user-centered design.

This analysis argues that the problem with LLM-based synthetic users extends far beyond questions of accuracy or authenticity. The core issue lies in what we term the “clarity deficit” —the systematic obscuring of the relationship between synthetic representations and actual user populations. This opacity creates a cascade of methodological problems that threaten the validity of design decisions, the integrity of user research, and ultimately, the quality of human-computer interaction.

The stakes of this analysis extend beyond academic methodology to practical design outcomes. When synthetic users replace or supplement real user research without adequate transparency about their representational basis, we risk creating what appears to be user-centered design while actually designing for algorithmic artifacts with unknown relationships to human needs, behaviors, and contexts.

The contemporary enthusiasm for LLM-generated synthetic users reflects a seductive proposition: if we can generate user representations that appear realistic, comprehensive, and behaviorally plausible, why not use them to replace or supplement expensive, time-consuming traditional user research? This proposition, however, rests on a fundamental category error that conflates representational plausibility with representational validity.

Modern LLMs excel at generating user profiles, personas, and behavioral scenarios that appear convincing to human evaluators. These synthetic users exhibit internal consistency, plausible demographic combinations, and coherent behavioral patterns that satisfy our intuitive expectations about human diversity and complexity. This plausibility creates what we term the “coherence illusion” —the mistaken belief that internally consistent synthetic users necessarily represent valid samples from real user populations.

The trap lies in conflating two distinct qualities:

Narrative Coherence : The internal consistency and plausibility of individual synthetic user profiles Representational Validity : The accuracy with which synthetic users reflect actual user population characteristics, behaviors, and needs