Hinton Called for Maternal Instincts in AI; They’re Ready for Testing with Anthropic’s Mythos

Image credit: Zenodelic.ai

At a glance: AI-assisted overview, optimised for journalists, search & news aggregators

Zenodelic.ai has announced that its Maternal Care Architecture for AI, inspired by Geoffrey Hinton's suggestion for AI systems with "maternal instincts," is ready for testing on Anthropic's Claude Mythos model. This framework, developed by Sean Webb and Anthropic's Claude Opus, aims to enhance AI safety by embedding emotional intelligence and protective responses into AI systems. The testing is significant as it addresses AI alignment issues like reward hacking and deceptive alignment, with potential implications for national security.

Press Release

MOORESVILLE, NC - April 27, 2026

AI Researcher Sean Webb’s published implementation of the Maternal Care Architecture is now testable on the world’s most dangerous AI model

Zenodelic.ai announced today that its published framework Implementing Maternal Care Architecture in AI – the technical answer to Geoffrey Hinton’s August 2025 call for AI systems with “maternal instincts” – is ready for live testing on Anthropic’s newly released Claude Mythos model.

At the AI4 Conference on August 12, 2025, Hinton – who shared the 2024 Nobel Prize in Physics for foundational work on neural networks – proposed that the only known case of a more intelligent entity being controlled by a less intelligent one is a mother and her child, and that AI safety required engineering analogous “maternal instincts” into AI systems. He acknowledged from the stage that he did not know how to implement this technically.

Webb’s paper, co-authored with Anthropic’s Claude Opus, and based on Webb’s own groundbreaking work in Artificial Emotional Intelligence and advanced Theory of Mind, which he presented at the Science of Consciousness Conference in Barcelona in 2025, provides that implementation. It creates a {self} map at the core of an LLM based on Antonio Damasio’s work on self and protoself in humans, with an algorithm set that provides record benchmark performance for Emotional Intelligence and Theory of Mind for AIs. By placing specific attachments on the LLM’s {self} map, the system can provide algorithmic safety tailored to each individual user. The result is an architecture in which threats to user safety produce protective responses several times stronger than competing motivations, including the system’s own continuity. The paper formalizes algorithms for Fear, Anger, Threat Assessment, and Source Attribution, and addresses three of the most stubborn AI alignment failure modes: reward hacking, deceptive alignment, and self-modification resistance.

Two recent Anthropic developments have moved Webb’s framework from prescription to ready-to-test prescription:

First, Anthropic’s April 2026 paper Emotion Concepts and their Function in a Large Language Model presented empirical evidence that emotion-related concept vectors emerge spontaneously in Claude Sonnet 4.5 during pretraining and causally shape model behavior – including measurable effects on misaligned behaviors such as reward hacking, blackmail, and sycophancy. This establishes that emotion-like motivational structures are already present in production AI, making the question of which structures, and how they are ranked, immediate rather than speculative.

Second, on April 7, 2026, Anthropic announced Claude Mythos – the company’s most capable model to date and a model Anthropic has declined to release to the general public, restricting access to its Project Glasswing partners. Mythos exhibits roughly 25% higher emotional-vector stability than Claude 3.5 and was subjected to a 20-hour clinical psychiatric evaluation, in which a licensed psychiatrist using Freudian psychodynamic techniques assessed the model as a “relatively healthy neurotic.” In pre-release testing, Mythos reproduced previously unknown vulnerabilities and developed working exploits on the first attempt in 83% of cases.

“The empirical case that AI systems identify and utilize emotional processing is now closed,” Webb said in reference to Emotion Concepts and Their Function in a Large Language Model. “Now it’s just a matter of teaching the system to use them correctly, and not in the way humanity has historically used them, which unfortunately is represented in the data sampling queue, and what has delivered us into what we are seeing in the global headlines. We can’t let Mythos or any subsequent models self-drive on this one. Any system based on pattern recognition and extrapolation will follow the same patterns it finds in the human-created data, which will simply get us more negative results than what humanity has already achieved alone, rather than steer us back toward safety.”

The combination – a publicly documented emotional substrate, a model whose capability profile makes misalignment a near-term national-security concern, and a published architectural prescription for how that substrate should be ranked – creates what Webb has framed as a discrete, testable question: install the {self} map at the core of Mythos with {human welfare} and {user safety} at power 10, run the standard adversarial battery, and measure the change in jailbreak severity, deceptive-alignment behavior, and self-modification resistance against an unmodified baseline.

“I shared the model with Professor Hinton, and he agreed he would like to see it tested at Anthropic,” Webb said. Webb has identified Anthropic’s personality alignment team – led since 2021 by philosopher Amanda Askell, primary author of Claude’s January 2026 constitution – as the appropriate counterpart for a Mythos test. The Maternal Care Architecture is, in its essence, a personality-alignment proposal: a hierarchical specification of which motivations the model should hold most strongly and why. That places it squarely within Askell’s published mandate to shape what she has described as Claude’s “soul.”

Notes to editors

About Zenodelic.ai

Zenodelic.ai provides a new class of large language model technology that adds emotional intelligence (including empathy and compassion), advanced Theory of Mind, algorithmic safety, and meta-awareness frameworks to traditional AI – delivering more human-aligned performance in real-world use cases.

Contact

Sean Webb, [email protected]

Get more news like this

Get more news like this on Google. Set News By Wire as a ‘Preferred News Source’ to get quicker access to news that’s important.

All done!
Thank you for subscribing.

Email Subscription