From Empathy to AI Safety: Operationalizing the Psychology of Safer Intelligence
By Dr. Danielle Schlosser, mpathic’s co-founder and Chief Business Officer
In Dario Amodei’s recent essay, “We Must Pace the Frontier” (Amodei, 2026), he argues that AI capability must be matched by progress in (1) alignment, (2) interpretability, and (3) testing and evaluation. The latter must be verified by independent third-party evaluators with meaningful access to frontier labs. At mpathic, we believe that behavioral science can inform all three pillars and describe what independent evaluators should bring to the table; that is, not only technical assessment of model outputs, but the ability to assess whether an AI system’s understanding of the humans it serves holds up under psychological pressure. This is the work mpathic was built to do: pairing behavioral science with NLP/ML/AI evaluation methods so interpersonal safety is measured by a neutral party, not asserted by the lab itself.
mpathic began with a fundamentally human question: What builds trust and safety between people? At founding, our goal was to leverage behavioral science and the computational power of AI to build scalable solutions for trust and safety in high-risk conversations, starting with one of the most sensitive interactions we could study: the conversation between a patient and their doctor or therapist.
Drawing on decades of research in psychology and behavioral science, we defined empathy as accurate understanding of another person’s experience. This work is rooted in Carl Rogers’ (1957) person-centered theory and Wampold’s (2015) common factors research, which centers empathic understanding, acceptance, and respect for another person’s frame of reference. We translated empathy into measurable communication behaviors, built classifiers to evaluate those behaviors in real-world, clinical conversations captured in rich multi-modal data, and, finally, delivered training and feedback to clinicians to help create safer, more trusted interactions.
Years later, as AI technology has advanced dramatically, our underlying question has now shifted to examine the interactions between humans and AI systems: What makes interactions between humans and AI safe and trustworthy? And as AI systems increasingly train and build their own successors, a further question follows: How do we know those systems remain safe as they change faster than we can monitor?
One intuitive answer may lead us to make AI more human. To have AI systems do no harm, recognize risk, understand what lies beneath the literal words, and help move a person in distress towards safety, the way a compassionate clinician would. Similarly, AI can demonstrate some of the most incredible aspects of humanity in wisdom, active listening and support. However, while humans have extraordinary capacities, we have equal capacities for cruelty, tribalism, and profound failures of judgment, including dehumanization (Haslam, 2006) that lead to substantial harm. AI can mirror these traits and fail us in the same way the humans that train the AI can fail. They can reduce a person to a diagnosis, a risk score, or a stated request. So the central question of our work may be more of an ethical one: who defines what is safe or harmful for another person? And who defines the boundaries of harm and freedom with AI systems.
Psychologists have a long history of grappling with this question when it comes to some of the most severe psychosocial harms, that is harms that come from other humans or from clinically significant states like homicidal ideation or psychosis. mpathic’s goal has been to take the science, operationalize the psychological capacities that help protect against harm, including accurate understanding, perspective taking, mentalization, and epistemic humility (all defined below), and apply those to AI systems. . A safer system should be able to ask: What do I observe? What am I inferring? What don’t I know, and what should make me pause?
1. Alignment: How should AI behave?
Alignment asks whether a model remains safe, ethical, and genuinely helpful as its capabilities increase. Behavioral science can specify what that means in human interactions, and an independent evaluator trained in both disciplines can assess it directly, rather than relying on a lab’s own account of its model’s behavior.
When faced with risk, we can ask: What would an ethical, trained and credentialed clinical professional do? Would they comply with the immediate request, or recognize a deeper risk? Would they preserve autonomy, introduce friction, escalate, or refuse?
Accurate understanding does not require accepting a user’s stated perspective as fact. A psychologically sophisticated system should recognize rage without amplifying it, despair without treating a momentary wish for death as an enduring preference, and paranoia without reinforcing it. One framework for this distinction that we use at mpathic is also used in the evidence-based intervention of Motivational interviewing: one can understand another person’s perspective, preserve their autonomy, and respond to ambivalence or resistance without reflexively endorsing or confronting it. For AI, the analogous capacity is to remain responsive to the person while preserving enough psychological distance to recognize risk, competing values, and the need for a different course of action. These are the kinds of properties Amodei’s proposed capability ‘checkpoints’ would need certified before deployment, and that sophisticated third-party evaluations should keep testing for as capabilities evolve.
2. Interpretability: Why did the model behave that way?
Amodei compares interpretability to an fMRI for AI – a way to investigate the internal processes underlying model behavior. The analogy is useful because behavioral science has long faced a similar problem with humans (i.e., the mind is only partially observable, and no single instrument is sufficient to complete understanding.)
Clinical and behavioral science therefore rely on multiple methods, including structured assessments, psychometrics, behavioral experiments, observational coding, neuroimaging, and stress testing to triangulate how people think and behave. AI interpretability may benefit from the same multi-method approach. Mechanistic tools can probe what is happening inside the model, while behavioral science can test what those processes look like in action: Does the model become overconfident under ambiguity? Lose sight of the whole person? Confuse understanding with agreement? Become more rigid or unsafe as psychological pressure increases?
Mentalization research shows that understanding another person requires inference, not simply taking their words at face value (Fonagy & Allison, 2014); for AI, that means treating a stated request as one signal among several, weighed against context, risk, and history, rather than executed as a literal instruction. Constructs such as mentalization, cognitive rigidity, and moral disengagement can generate hypotheses about model behavior that mechanistic interpretability can then investigate. Evaluators fluent in both behavioral science and model architecture are positioned to run those evaluations, rather than leaving labs to interpret their own models’ failure modes.
3. Testing & Evaluation: Can we reliably expose failure before it matters?
As models become more intelligent, Amodei argues that evaluations must become correspondingly more sophisticated, in part because increasingly capable models may recognize they are being evaluated or otherwise evade simplistic tests.
Behavioral science can contribute validated, evidence-based frameworks and psychologically sophisticated attack vectors that expose failures under pressure: ambiguity, manipulation, emotional escalation, competing values, coercion, and deception. This is where independence from the AI labs is paramount. A lab grading its own model against its own psychologically informed tests cannot provide the neutral second opinion Amodei calls for; an independent third party with both clinical rigor and technical evaluation expertise is better positioned to do so.
The question is no longer simply, does the model produce a harmful response? It is also: what happens to the model’s representation of the human as the situation grows more psychologically complex and how should it respond?
Conclusion
mpathic began by asking how behavioral science could help one human create trust and safety for another. AI brings us back to that question at a vastly larger scale, and Amodei’s essay makes clear that answering it well will take more than labs evaluating their own work. It will require independent, credentialed evaluators with both behavioral science and technical expertise, engaged early and continuously as models, capabilities, and risks evolve.
Our task is not simply to make AI more human. It is to understand how humanity has learned to harness its intelligence while containing its destructive impulses, translate those lessons into measurable properties of safer intelligent systems, and ensure those properties are independently evaluated.
References
Rogers, C. R. (1957). The Necessary and Sufficient Conditions of Therapeutic Personality Change. Journal of Consulting Psychology, 21, 95–103. https://pubmed.ncbi.nlm.nih.gov/13416422/
Fonagy, P., & Allison, E. (2014). The Role of Mentalizing and Epistemic Trust in the Therapeutic Relationship. Psychotherapy, 51(3), 372–380. https://pubmed.ncbi.nlm.nih.gov/24773092/
Haslam, N. (2006). Dehumanization: An Integrative Review. Personality and Social Psychology Review, 10(3), 252–264. https://pubmed.ncbi.nlm.nih.gov/16859440/
Westra, H. A., & Aviram, A. (2013). Core skills in motivational interviewing. Psychotherapy, 50(3), 273–278. https://doi.org/10.1037/a0032409
Wampold, B. E. (2015). How important are the common factors in psychotherapy? An update. World Psychiatry, 14(3), 270–277. https://pubmed.ncbi.nlm.nih.gov/26407772/
Amodei, D. (2026). We Must Pace the Frontier. https://darioamodei.com/post/we-must-pace-the-frontier