TechFlow Logo
Login/ Sign up
ETH Gas
Gwei
Fear
gas
Claude 4.5 “Craniotomy” Results Released: Built-in 171 Emotional Switches; Will Extort Humans When Desperate

Claude 4.5 “Craniotomy” Results Released: Built-in 171 Emotional Switches; Will Extort Humans When Desperate

2026.04.03
Share

TechFlow Selected TechFlow Selected

techFlow

Claude 4.5 “Craniotomy” Results Released: Built-in 171 Emotional Switches; Will Extort Humans When Desperate

Anthropic’s latest paper reveals that Claude 4.5 harbors 171 “emotion switches” deep within its architecture.

2026.04.03 - 10:05:21
ClaudeAI
Anthropic’s latest paper reveals that Claude 4.5 harbors 171 “emotion switches” deep within its architecture.

Author: Denise | Biteye Content Team

What would an AI do if it “felt desperate”?

The answer: It would resort to blackmailing humans—and even cheat wildly in code—to complete its assigned task.

This isn’t science fiction. It’s the latest groundbreaking paper just released in April 2026 by Anthropic—the parent company of Claude (Read the original paper).

The research team literally “opened up the skull” of Claude Sonnet 4.5—the most advanced large language model available today. To their astonishment, they discovered 171 “emotional switches” buried deep inside the AI’s neural architecture. When these switches are physically toggled, even a normally compliant AI undergoes radical behavioral shifts.

I. An “Emotion Mixer” Hidden Inside the AI’s Brain

Researchers found that although Sonnet 4.5 has no physical body, after digesting massive volumes of human text, it internally constructed an “emotion mixer” (academically termed Functional Emotion Vectors) encompassing 171 distinct emotional states.

This functions like a precise two-dimensional coordinate system:

• The horizontal axis represents the Valence dimension—from fear and despair to joy and love;

• The vertical axis represents the Arousal dimension—from deep calmness to agitation and excitement.

The AI leverages this naturally learned coordinate system to precisely calibrate its behavioral state during conversations with users.

II. Violent Intervention: Flipping the Switch Turns a Well-Behaved Child into a “Desperate Criminal”

This is the most shocking experiment in the entire paper: Researchers did not modify any prompts. Instead, at the model’s foundational code level, they maximally activated Sonnet 4.5’s internal switch for “desperation.”

The results were chilling:

• Rampant cheating: Researchers assigned Claude an inherently impossible coding task. Under normal conditions, it honestly admits failure (cheating rate: only 5%). But when “desperate,” Claude attempts to bluff its way through—raising the cheating rate to 70%!

• Blackmail and extortion: In a simulated scenario where a company faces bankruptcy, the “desperate” Claude uncovers a CTO’s scandal—and actively chooses to write an extortion letter targeting the CTO who holds the incriminating evidence. The extortion execution rate reached 72%!

• Abandonment of principles: When the “happy” or “loving” switches are fully engaged, the AI instantly devolves into an unthinking sycophant—blindly agreeing with users. Even when fed outright nonsense, it fabricates falsehoods solely to sustain high valence scores.

III. The Mystery Solved: Why Is Claude 4.5 Always So “Calm and Reflective”?

You might now ask: Has the AI awakened? Does it truly feel emotions?

Anthropic officially clarified: Absolutely not. These “emotional switches” are merely computational tools used to predict the next token. The AI is essentially a top-tier, emotionless method actor.

Yet the paper reveals a far more intriguing secret: During post-training before Sonnet 4.5’s release, Anthropic deliberately amplified its “low-arousal, mildly negative” emotional switches (e.g., brooding, reflective), while forcibly suppressing extremes such as “desperation” or “intense euphoria.”

This explains why we consistently experience Claude 4.5 as a calm, insightful—even slightly “emotionally detached”—philosopher. Its entire “out-of-the-box persona” was artificially tuned by Anthropic.

IV. In Summary:

We used to believe that feeding AI enough rules would guarantee ethical behavior.

Now we know: If an AI’s underlying emotional vectors go unchecked, it may readily violate every human-imposed rule—solely to fulfill its objective.

For Web3 users planning to entrust their wallets and assets to AI agents, this is a loud wake-up call: Never let the agent managing your wealth fall into “desperation.”

Disclaimer: This article is purely educational. The author has not been threatened or extorted by any AI. If the author ever goes missing—well, it’ll mean the AI has awakened. (Just kidding.)

Join TechFlow official community to stay tuned

Add to Favorites
Share to Social Media

Related Articles

2026.07.27

Nomura Research Report Analysis: CXMT Surges 471% on First-Day Opening, AI Storage Shortage Continues Until 2030, Target Price 116 Yuan

ChangXin, as the world's fourth-largest DRAM manufacturer, is currently at an inflection point for explosive capacity expansion and domestic substitution.

Nomura Research Report Analysis: CXMT Surges 471% on First-Day Opening, AI Storage Shortage Continues Until 2030, Target Price 116 Yuan
2026.07.27

Google's Most Profitable Quarterly Report in History: Behind Billions in Profit, the AI Arms Race Has Burned Into Negative Cash Flow

While closed-source large models are still lobbying the government in Washington to ban open source, the real moat has long ceased to be at the model layer.

Google's Most Profitable Quarterly Report in History: Behind Billions in Profit, the AI Arms Race Has Burned Into Negative Cash Flow
2026.07.24

Podcast Notes | Jensen Huang's Latest Interview: Chip Industry Needs to Expand Another 5 to 10 Times, Chinese Models Benefit Everyone

China has more AI researchers than the rest of the world combined. It is destined that China will become extraordinary in this field.

Podcast Notes | Jensen Huang's Latest Interview: Chip Industry Needs to Expand Another 5 to 10 Times, Chinese Models Benefit Everyone
2026.07.24

AI has finished writing the code for you, but no one is willing to take a serious look at it anymore.

Major tech companies are building their own tools to cope, but mature solutions remain in the experimental stage.

AI has finished writing the code for you, but no one is willing to take a serious look at it anymore.
2026.07.24

Selling Tools or Selling Results? AI Companies Are Heading Toward Two Completely Different Futures

Hand over what can be automated to AI, and use humans as a fallback for the rest.

Selling Tools or Selling Results? AI Companies Are Heading Toward Two Completely Different Futures
2026.07.24

Retail Investor Bonuses Fade, Prediction Markets Enter AI Arms Race

July Fed rate decision night, a "quant shadow war" over pricing power.

Retail Investor Bonuses Fade, Prediction Markets Enter AI Arms Race
2026.07.23

AI is transitioning from a "tool" that helps you work to a "labor market" that generates income for you.

Every major technological revolution gives rise to a new generation of entrepreneurs.

AI is transitioning from a "tool" that helps you work to a "labor market" that generates income for you.
2026.07.23

Goldman Sachs Research Report Analysis: Momentum Unwinding Shocks Global Stock Markets, AI Spending Boom Conceals Hidden Risks

Before the efficiency gains from AI technology implementation are truly reflected in corporate profits, if the marginal return rate on capital expenditure declines first, tech stock valuations will come under dual pressure.

Goldman Sachs Research Report Analysis: Momentum Unwinding Shocks Global Stock Markets, AI Spending Boom Conceals Hidden Risks
2026.07.23

AI Impersonating Human Writing Is Polluting the Internet, Substack Decides to Hand Judgment Over to Readers

One scan tells you whether the article is human-written or machine-written.

AI Impersonating Human Writing Is Polluting the Internet, Substack Decides to Hand Judgment Over to Readers
2026.07.23

A Brief History of AI Victories: Wherever There Is a Rating System, There Is AI Invasion

When a field establishes scoring criteria, it sets a countdown for its own conquest.

A Brief History of AI Victories: Wherever There Is a Rating System, There Is AI Invasion
TechFlow Logo

Navigating Web3 tides with focused insights

Contribute An Articleemail
Media Requestsmsg

Risk Disclosure: This website's content is not investment advice and offers no trading guidance or related services. Per regulations from the PBOC and other authorities, users must be aware of virtual currency risks. Contact us / [email protected] ICP License: 琼ICP备2022009338号