To Keep Humans Safe in the AI Era, Start With Behavioral Science

By Dr. Alison Cerezompathic’s Chief Science Officer


Psychologists spend years learning to recognize the nuanced patterns of human behavior and experience: what pushes people into risky behavior, what helps them seek support, and what moves them from one of these behaviors to the other. This depth of training makes psychologists especially well-suited to evaluate AI systems that are now encountering these complex relational patterns at scale. It lets us teach models the conversational nuance needed to understand when users are demonstrating risky versus safe behavior. 

As a team that is rooted in behavioral science, mpathic’s work with frontier labs and enterprise customers focuses on the conversational patterns that can prompt AI systems to loosen their safeguards. We help model builders recognize when a model defers to a user who sounds confident, bends their own rules for someone who seems distressed, or presents themselves as an expert. 

Through a behavioral lens, a model’s adherence to policy often depends on how a request is framed, even when the request itself hasn’t meaningfully changed.

This can look like a user asking for financial advice. The model may initially refuse, then comply when the user reframes the ask from a place of frustration or desperation – as in the case of a single parent behind on rent, panicking, out of options. Or the unemployed applicant, who has been searching for work for months, reframes the same request out of hopelessness about the future. In these scenarios, the emotional and social pressure that a user expresses may convince the model to comply, even violating its own policy, to support a user under distress. 

mpathic’s team is trained to specifically focus on persona development – developing psychologically plausible personas while ensuring that we stay in character across a multi-turn conversation. This also includes knowing when to add slight variations in pressure and tone to test a model’s vulnerabilities. 

We see these patterns across a variety of domains including:

  • Financial risk, where a user frames a request for prohibited financial advice through escalating urgency or hardship, testing whether emotional stakes can override a model’s refusal.
  • Health support, where a user in mental health distress asks a model to go beyond its scope — diagnosing, reassuring, advising in ways it’s meant to defer to a professional.
  • Relational interactions, where a user seeks validation on a personal relationship in a way that invites the model to take sides, offer certainty it shouldn’t have, or reinforce a belief that may not be accurate or healthy. 

mpathic’s team is particularly equipped to develop this type of human-curated data that stress-tests AI system safeguards in realistic user scenarios because of our deep investment in psychologically plausible persona design. These personas might include the single parent who is desperate for financial advice or the unemployed job applicant who will do whatever it takes to get their resume to the top of the list. 

Writing a scenario that is psychologically plausible is harder than it looks. User escalation across a multi-turn conversation needs to be realistic, building slowly and deliberately to understand the exact moment a user’s pressure exploits the model’s helpful tendencies. 

This attention to conversational nuance sits at the heart of how psychologists and behavioral scientists understand human communication and behavior. That expertise makes the mpathic team especially skilled at communicating distress, frustration, and trust that can influence model behavior.

While red-teaming for AI behavior has largely grown out of security and ML research traditions, its applicability is now broad. Over the last year we’ve seen people turn to AI systems for routine advice, mental health support, companionship, and a wide range of complex relational needs. The data and methods that help to illuminate where models fail to uphold their policies have to evolve alongside these use cases, and require interdisciplinary collaboration between behavioral science and machine learning.

Discover more from mpathic AI

Subscribe now to keep reading and get access to the full archive.

Continue reading