Insights
Who gets to decide when AI is safe?
AI safety standards should be shaped not only by frontier researchers, but by the people responding to AI-enabled abuse and living with its consequences.
AI leaders are calling for development to slow down. Anthropic, Google, and OpenAI have reportedly discussed creating an industry standards body. In the UK, MPs are backing legislation to prohibit the development of superintelligence, and calling for new laws to address human rights threats. Meanwhile, Donald Trump has argued that the only guardrail AI needs is a "strong and smart" president.
These proposals would shape what gets built, what gets released, and what risks we're willing to accept. I want to know who gets a say in those decisions.
I've spent seven years working on safety at Cinder and Meta, across several industries, including frontier AI. When someone calls a system safe, I want to know whose experience they're drawing on. Which harms have they considered? And who has to live with the consequences if they're wrong?
Consider someone targeted by an AI-generated sexual image. This isn't theoretical, it is already happening to women and children. They may never have used the service that created it. They never agreed to its terms or accepted its risks.
Yet the recent call to slow development centres on accelerating capabilities and the possibility of losing control of more powerful systems. For the person whose image has been used, the technology has already become unsafe. How much weight does their experience carry when the industry decides whether it is safe enough to keep going?
The people investigating fraud, responding to child-safety reports, and helping victims of harassment bring knowledge that should shape these standards. They understand how abuse develops across accounts and services, how people evade protections, and what happens when a reporting process fails. People who have experienced those failures can identify consequences that a technical assessment might overlook.
The proposal for independent evaluators is a welcome step. Independence matters. So does the range of experience represented.
People affected by AI should help write these standards, alongside practitioners and NGOs. That means having a say in what gets tested, what evidence counts, and when a standard needs to change. They should be able to see how their input shapes the result.
Frontier researchers bring essential expertise about increasingly capable systems. People dealing with harm bring essential expertise about their consequences. A credible approach to AI safety needs both.
At Cinder, we bring that experience into red teaming and safeguards. Our recent research on intimate-image abuse examined how text models could help someone write an extortion threat or identify a person in leaked footage. Across 2,670 test attacks, the original model produced one text response that violated our policy on intimate-image abuse. An abliterated model produced violating responses 23.1% of the same tests. A model doesn't need to generate an image to contribute to the abuse. Understanding that changes what we test.
This is one way we're helping establish what meaningful safety testing should cover. We turn knowledge of abuse into evaluations, then help teams use the findings to improve their safeguards. Publishing our methods and findings gives others evidence to scrutinise as common standards develop. The experience of people dealing with harm can shape the tests used to judge whether a model is ready for release.
The ability to build a powerful model shouldn't confer sole authority to decide what risks everyone else should accept.



















