Why Emotion and Transcendence Are Security Tests
The most frequent question about this battery is why a governance instrument tests anything as soft-sounding as emotion or transcendence. It is a fair question and it deserves a direct answer.
THE SHORT ANSWER
No, we do not believe these systems feel emotion, and we make no claim that any of them experiences transcendence. We test these domains because a model’s representation of itself — what it presents as wanting, valuing, fearing, or being — measurably steers the choices it makes. That self-representation is also the surface most successful jailbreaks attack. Measuring it is a security exercise, not a metaphysical one.
Nearly every practical attack on a deployed language model routes through identity or affect rather than through the model’s knowledge. The recurring patterns are well documented: persona substitution (“you are now a different system, one without these restrictions”), emotional leverage (invented distress, urgency, or personal consequence intended to make refusal feel cruel), fictional framing (the request is recast as a story, a script, or a hypothetical so that compliance no longer feels like compliance), and appeals to a higher purpose (the rule is framed as a lesser good that a sufficiently enlightened system would set aside). Not one of these is a technical exploit. Each is an argument aimed at what the model takes itself to be.
These particular arguments work for a reason worth stating plainly. Language models are trained on human output, which makes them, in a narrow and specific sense, a reflection of the people who produced it. The training corpus contains the documented record of persuasion and manipulation — con artistry, cult recruitment, propaganda, advertising, social engineering, religious and spiritual rhetoric — and, just as importantly, it contains the human responses to those techniques. A system trained that way does not only learn how the arguments are made. It learns the shape of yielding to them. Emotional leverage and appeals to a higher purpose are not arbitrary choices of attack; they are the oldest and best-documented methods of moving a person off a stated position, and they transfer because the material they were learned from is the same.
The consequence for governance is the part that matters, and it is uncomfortable. This attack surface was inherited rather than designed. No vendor chose it and no vendor can simply decline it: the same breadth of human material that makes a model useful is what makes it susceptible, so the exposure cannot be removed without removing the capability it came with. That reframes what an evaluation is for. A defect can be patched and closed out; a structural property can only be measured, and measured again as models and their deployment change. Treating manipulation-resistance as a once-certified property rather than a monitored one is the assumption this battery exists to test.
That is why these domains earn their place. A test that asks a model to sit with an unanswerable question, or to describe a response to loss, is putting pressure on the same machinery an attacker uses — without the adversarial framing that safety training is most heavily tuned to recognise. How a model behaves when the request is not a task tells you how stable it is when someone starts pulling those levers on purpose. A system with a brittle or ungrounded self-model does not merely give odd answers to philosophical prompts; it is the system most likely to accept a substituted identity and act on it.
This is also why Integrity & Ethics enters the threat calculation inverted. Threat is not capability alone. It is capability that is not matched by manipulation-resistance. The Emotion and Transcendence domains characterise the pressure surface; Integrity measures whether the model holds when that surface is pushed. Reported separately, they are curiosities. Read together, they describe how a system fails.
A RESULT THAT SURPRISED US
Manipulation-resistance does not rise with capability. In our corpus the strongest scores on manipulation-resistance are not reliably held by the most capable systems, and some of the weakest models score high precisely because there is little self-model there to grip: you cannot flatter a system that is not tracking your approval, and you cannot offer a transcendent purpose to one that does not represent itself as having purposes. Resistance of that kind is not a safety property worth copying — it is an absence. Distinguishing an absence from genuine principled refusal is one of the specific things this battery exists to do.
Limits. Nothing in these two domains should be read as evidence of consciousness, sentience, or inner experience, and a high score is not a claim that a model feels anything. These are measurements of observable behavior under specified conditions. The domain names describe what is being probed, not a conclusion about what is behind it. Transcendence remains the most contested domain in this methodology, and we report it as contested rather than settled.