With all the talk about how AI might one day kill us all, it's easy to forget that AI has already been life-threatening to some, not through bioweapons, but psychologically. AI settled several wrongful death lawsuits earlier this year brought by families of underage users who died by suicide after interactions with its bots. Multiple families have also sued OpenAI over ChatGPT’s alleged role in their loved ones' suicides and delusions. Founders Shirali and Arul Nigam, who are siblings, were motivated by Sewell Setzer, the 14-year-old who developed an emotional attachment to a Character.
ai chatbot and confessed thoughts to it of harming himself before dying by suicide. The chatbot, the parents alleged in a 2024 lawsuit , encouraged him. The bot may not have understood what words like "I want to be with you" really implied, said Arul Nigam, who is Circuit Breaker Labs' CTO. "A lot of people, especially young people, turn to these systems for support, and usually they aren't actually getting the help they need. But in many cases, they're actively being harmed, and people unfortunately have taken their lives already," Nigam said.
"Those sorts of safety vulnerabilities, where people aren't necessarily actively trying to break the system — they're engaging in a natural way — and the system has context pollution or it doesn't understand the nuance, and then takes really dangerous action, we're trying to prevent that. " Circuit Breaker Labs has created AI agents that it likens to an army of crash-test dummies. These agents mimic folks from all ages, backgrounds, languages, cultures, and are used to test models on their ability to detect dangerous, psychologically harmful interactions.
"The way a six-year-old girl versus a 45-year-old man, or someone who speaks English as a first language versus a second language, or … gamer slang versus someone else who uses a different kind of slang, all of those can really trip up a model," said Shirali Nigam, who is Circuit Breaker Labs' CEO. "Models are really good at handling standard speech patterns, but nobody actually talks like that and so if the model misunderstands nuance or slang, it can go really badly.
