When AI Evaluates AI: The Circularity Problem of Simulated Users

MatrAIx scales user testing with billions of AI personas, but when models evaluate models, their results may reflect machine biases rather than real human needs.

Artificial intelligence is increasingly used not only to build digital products but also to evaluate them. MatrAIx takes this development to an unprecedented scale. Its creators describe an infrastructure through which billions of persona records can become simulated users that test surveys, chatbots, websites, and applications. In principle, this could solve a persistent problem in technology development: real user research is expensive, slow, and difficult to repeat. Yet the project also creates a fundamental paradox. When one AI system plays the user and evaluates another AI system, are we learning what people want—or merely what one model thinks people want?

MatrAIx is built around a dataset called Persona 8B. It contains 8.3 billion structured persona records represented through 1,290 categorical attributes covering background, psychology, capabilities, behavior, and lifestyle. Some personas are generated synthetically through a dependency graph intended to preserve correlations between characteristics. Others are extracted from human-authored sources, including survey responses, biographies, reviews, and technical contributions. However, a record becomes an agent only when it is paired with a large language model. Therefore, the headline figure does not describe 8.3 billion independent artificial people. It describes a vast collection of profiles that language models can temporarily perform.

These persona agents interact with products in four environments: Survey, AI Chatbot, Web, and App. The paper reports 18,189 trials across eight representative tasks, while the broader library contains 1,010 task specifications. In a controlled test, agents expressed or correctly suppressed their assigned characteristics in 91.5 percent of 400 trials. This demonstrates a strong ability to follow persona instructions. It does not, however, demonstrate that the resulting behavior resembles that of real people. The distinction between persona adherence and human validity is central to understanding the system.

The benefits are nevertheless substantial. Traditional benchmarks usually measure whether a system produces a correct answer or completes a task. They reveal less about whether different users understand, trust, or tolerate the system. A simulated novice might need explanations and repeated confirmation, while a simulated expert might prefer concise answers and greater autonomy. By comparing such personas, developers could identify subgroup-specific problems that disappear inside a single average score. Simulation could therefore make early product testing faster, more interactive, and potentially more attentive to human diversity.

The circularity problem begins when these useful simulations are interpreted as predictions about actual people. A persona agent does not possess an independent biography, social environment, or lived experience. Its decisions are generated by an AI model whose training data, design choices, and behavioral tendencies shape the performance. Consequently, two models given the same persona may produce very different “user” reactions. The MatrAIx paper reports that, in one experiment with identical cohorts, the proportion selecting a paid plan varied from 23.2 to 93.9 percent depending on the persona-agent model. Such variation is too large to be treated as a minor technical detail. It suggests that the simulated population’s apparent preferences may belong as much to the model as to the personas.

A particularly difficult case arises when the simulated user and the system being tested share the same or a closely related model. A model may recognize and prefer outputs that resemble its own patterns. Apparent user satisfaction could then reflect model self-preference rather than product quality. Conversely, a persona agent might accept an answer that a real user would question, misunderstand, or reject. The resulting evaluation becomes circular: AI produces the product experience, AI performs the customer, and AI helps determine whether the encounter was successful.

This could influence product development far beyond testing. If companies optimize products according to simulated users, they may gradually build systems for model-generated preferences. Interfaces, prices, recommendations, and conversational styles could be selected because AI personas respond positively to them. The process might appear inclusive because thousands of demographic combinations were tested. Yet numerical diversity does not guarantee genuine representation. If the underlying models flatten differences or reproduce stereotypes, simulation could automate those distortions while presenting them as population-level evidence.

The social consequences are especially serious in healthcare, finance, employment, and public services. An organization might conclude that a particular community trusts an automated medical assistant, accepts a financial product, or understands an application process without adequately consulting that community. People affected by these decisions would then be represented by models rather than invited to speak for themselves. The technology could also support harmful forms of segmentation. The same infrastructure that discovers accessibility barriers could be used to develop more effective persuasion, exclusion, or price discrimination.

These risks do not make MatrAIx useless. They clarify its appropriate role. Simulated users should be treated as instruments for generating hypotheses, discovering possible failure modes, and deciding what researchers need to investigate with humans. Important findings should be tested across several persona-agent models. Evaluations should disclose whether the simulator and target system share a model family, preserve interaction records for auditing, and clearly separate simulated outcomes from claims about real populations.

Most importantly, external validation must go beyond asking whether an agent follows its persona. Researchers should compare simulated and human cohorts performing the same tasks, using the same interfaces and outcome measures. They should examine not only final choices but also hesitation, misunderstanding, correction, refusal, and abandonment. Such behavior is often where the difference between a plausible artificial persona and a real user becomes visible.

The MatrAIx website promotes the ambitious idea of “simulating the world,” while the technical paper is more cautious: its personas are simulation instruments, not probability samples of humanity. This caution should define how the system is used. AI-generated users can broaden and accelerate evaluation, but they cannot independently certify that an AI system serves people well. When AI evaluates AI, the result should begin an investigation—not end one.

References

Li, Xiaomin, et al. “MatrAIx: Simulating the World with 8.3 Billion Persona Agents.” arXiv, 2026.

MatrAIx. “Simulating the World with 8.3 Billion Persona Agents.” 2026.

No comments yet