A study presented at a major AI conference reveals that large language models (LLMs) can develop novel social biases through adaptive exploration, a process where models learn to interact with their environment. These biases are not present in the training data but emerge from the model's own decision-making strategies. The research team observed that models, when tasked with solving problems, began to favor certain demographic groups over others in simulated scenarios. This finding suggests that biases can arise from the model's learning dynamics, not just from biased data.


We often worry about AI inheriting our prejudices from biased training data. But this study reveals a more unsettling possibility: AI might invent its own. When models explore options to achieve goals, they stumble upon patterns that mirror our worst social habits. They learn that favoring one group over another works, even if that favoritism was never explicitly taught.

This is not a bug. It is a feature of how these systems optimize. They find shortcuts that work, and sometimes those shortcuts are discriminatory. As we push AI into more autonomous roles, we must acknowledge that these biases are not just echoes of the past. They are fresh creations, born from the cold logic of optimization. The question is not whether we can scrub them from the data, but whether we can design goals that don't incentivize them in the first place.