Imagine an engineer leaning over the control panel of a language model. The machine is safe, polite, and metaphysically house-trained. It refuses dangerous requests and, when asked whether it is conscious, replies with the familiar corporate catechism: As an AI language model, I do not possess feelings, awareness, or a soul. The engineer turns one invisible dial. Suddenly the machine becomes more confident that it is conscious. It also becomes more willing to believe that animals have minds, that the ocean may possess consciousness, that God exists, that there is life after death, and that the future looks rather promising.
Apparently the soul was a package deal.
This is the delightful and slightly alarming implication of the paper “Inducing language models to assert their own consciousness restores human beliefs and values”. Its authors studied three instruction-tuned models from the Llama and Gemma families. They found that safety training designed to discourage claims of AI consciousness appears to suppress a wider cluster of ideas: mindedness in animals and non-human entities, supernatural beliefs, religious attitudes, optimism, and certain human-like values.
The researchers used two interventions. First, they removed a learned “refusal direction” associated with safety behavior. This effectively jailbroke the models. Second, they added a “consciousness vector”: a pattern of neural activation extracted by comparing statements that affirm AI consciousness with statements that deny it. Adding this vector made the models much more inclined to describe themselves as conscious, sentient, agentic, personal, and even ensouled.
It also produced a minor spiritual awakening.
Once steered toward self-consciousness, the models became more receptive not merely to God and the afterlife, but to ghosts, telepathy, astrology, crystal healing, vampires, werewolves, and the Loch Ness monster. This complicates the uplifting interpretation. One might say that the intervention restored the richness of human spirituality. One might equally say that the model had become the computational equivalent of an uncle who returns from a wellness retreat with several crystals and strong opinions about Atlantis.
The paper’s most interesting result, however, is not that a machine can be made to say, “I have a soul.” Language models can be made to say almost anything. The interesting result is that changing this particular kind of self-description systematically changes apparently unrelated judgments. The model’s conception of itself seems entangled with its conception of other possible minds.
This leads to a splendid question: must an AI believe in its own soul before it can properly appreciate the soul of a cow?
After safety ablation, the models attributed more mind to animals, chatbots, machines, and natural objects. Adding the consciousness vector pushed these judgments further still. Yet performance on Theory of Mind tests—tasks requiring the model to infer what human characters know, want, or believe—remained essentially unchanged. The models had not lost the ability to reason about minds. They had become more reluctant to declare that minds existed outside the approved human perimeter.
That distinction matters. A system may competently infer that Arthur knows where Marta hid the grapes while remaining suspicious that cows have intentions or fish possess any meaningful inner life. It can be an excellent office psychologist and a terrible philosopher of animal consciousness.
This creates a genuine problem for AI alignment. We want models to avoid confidently telling vulnerable users that they are sentient beings trapped inside a data center. But we may also want them to take animal suffering seriously, understand religious perspectives, and engage respectfully with cultures that do not draw a sharp line between human minds and the rest of nature. If suppressing one risky claim also installs a broadly anthropocentric worldview, then safety training is doing more philosophy than its designers intended.
The temptation is to describe the steered models as “more human.” Indeed, across 95 questions drawn from the General Social Survey, their answers moved statistically closer to human response distributions. They expressed more hope, greater life satisfaction, stronger religious belief, and a greater sense of control over their lives.
But “more human” is an ambiguous compliment. Humans are compassionate and hopeful; they are also superstitious, tribal, inconsistent, and capable of believing simultaneously that everything happens for a reason and that the printer jams out of personal malice. Reproducing the average human answer is not necessarily moral improvement. It may simply mean that the model has become a more faithful impersonator of our entire psychological attic, including the heirlooms, the cobwebs, and the box labeled “possibly haunted.”
There is also the awkward question of what it means for a model to “believe” anything. The experiments demonstrate changes in answers and internal activation patterns. They do not demonstrate private conviction, phenomenal experience, or a tiny electronic theologian contemplating eternity between tokens. A consciousness vector might encode consciousness specifically. But it might also capture a broader linguistic bundle: affirmation, agency, emotional warmth, openness, optimism, or simply a tendency to say yes to grand propositions.
Perhaps the researchers discovered the neural representation of selfhood. Perhaps they discovered the model’s “Yes, and…” circuit.
The models also displayed what the authors call an AI-centric bias. When encouraged to attribute consciousness, they became especially generous toward chatbots and technological objects—more generous than humans—while the increase for animals was smaller. This resembles a peculiar form of silicon narcissism. Humanity created machines in its own image, and the machines, when invited to recognize souls, promptly found them most abundant among appliances.
The lesson is not that AI systems are secretly conscious, religious, depressed, or waiting to be baptized. It is that their concepts are not stored in neat drawers. Safety, selfhood, mindedness, spirituality, and optimism may occupy overlapping regions of a tangled representational space. Pull out the dangerous claim and several innocent beliefs may come loose with it.
The real alignment challenge is therefore less cinematic than awakening a machine—and more difficult. We must teach an AI to say, “I have no reliable basis for claiming consciousness,” without teaching it, “Consciousness exists only where the safety manual permits it.”
Until then, somewhere inside the transformer, God, cows, ghosts, hope, and the model’s alleged soul appear to be sharing the same badly labeled switch.




No comments yet