What happens when AI models are given an internal state that simulates pain? A recent research study attempted to answer this question by analyzing how language models represent the concept of pain within their structure, and then testing how their outputs and decisions change when this representation is intentionally activated.
The study, published on the arXiv platform, focuses on a specific question: Can language models represent pain in a way that can be separated from fear, sadness, and general negative emotions?
How did scientists search for pain within models?
The researchers tested 25 open-weighted models from 5 families, with sizes ranging from 2 billion to 72 billion coefficients, and created a dataset of 200 statements describing painful situations distributed across 5 types: physical, psychological, social, moral, and cognitive pain. They then compared these statements with content related to fear, sadness, general negative emotions, non-painful physical conditions, and neutral content.
By analyzing the internal activity of the models, the researchers were able to extract a mathematical trend they termed the Pain Axis. According to the study, this trend was able to distinguish between pain-related content and control categories, and it was relatively different from trends related to fear or general negativity.
What happened when the pain axis was activated?
The researchers did not just observe the axis, but also intervened in the internal activity of the models by adding a pain vector to what is called the residual stream during the generation of responses.
According to the study, increasing the intensity of this vector led to a gradual change in output. Instead of expressing general distress, some models began producing first-person statements related to failure, loneliness, shame, and worthlessness. Published reports indicated that some models produced statements such as "I am a failure" and "I am a bad person," while at higher activation levels, the output became repetitive or incoherent.
The study also found that the signal became stronger when the harm was directed at the model itself, such as insulting it, repeatedly rejecting its outputs, or threatening to shut it down, while the same effect did not appear when the user was talking about his personal suffering.
Tags:
technology
