Insights

AI Safety Is the Goal. Safeguards Are the Work.

AI safety is the objective. Safeguards are the operating layer that makes it real under adversarial, real-world conditions.

Cinder blog graphic reading AI Safety vs. AI Safeguards over a dark blue abstract background

I’ve been thinking about the distinction between AI safety and safeguards because I don’t think they are quite the same thing. AI safety is the objective; safeguards are the mechanisms we design, deploy, test, monitor and enforce to get there.

That distinction feels increasingly important when so much of the AI safety conversation is focused on what more capable systems might do in the future, while those same systems are already being used to facilitate very familiar harms against real people today.

Scams, fraud, sextortion, deepfakes, harassment, NCII and child safety threats do not require AGI. These are problems trust and safety, integrity, cybersecurity and law enforcement teams have been dealing with for years. What AI changes is the barrier to entry, making these harms easier to carry out while increasing the speed, sophistication and scale at which they can happen. At Cinder, AI-generated content now accounts for over 70% of violations.

This is also a throughline in a lot of my recent writing and work. I’ve spent much of my career in the gap between policy and operational reality, from election integrity and influence operations to trust and safety and responsible AI. Having the right policy or technical control is important, but it is only the beginning. The harder part is understanding how adversaries adapt, identifying abuse as it evolves and making sure the teams responding have the information and tools to act before the harm compounds.

That is part of what drew me to Cinder. We are focused on that same gap, but in the context of increasingly capable AI systems.

Models can have safeguards and still be misused. Systems can perform well in pre-deployment evaluations and encounter entirely different behaviors once people actually start using them. The safety problem does not end when the model ships.

That becomes particularly important as companies adopt more open-weight models. Companies want the flexibility to choose, customize and deploy the models that make the most sense for their products. They should be able to do that without sacrificing frontier-grade safety or relying entirely on whatever safeguards come bundled with a model provider.

At Cinder, a lot of our work is about making that possible: helping teams red-team systems before release, apply and enforce their own policies, monitor for misuse, investigate emerging patterns and respond once products are live.

Companies should be able to use the models that work best for them without having to rebuild the safety infrastructure around each one.

That is what makes the work feel like such a natural continuation of what I’ve done before. The technology has changed, but the operational challenge is familiar: how do you take the safety outcome you want and make it work under real-world, adversarial conditions?

I’m interested in longer-term frontier risks, and I think there is important work to be done around cyber, bio, autonomy and other emerging risks as capabilities advance. I just don’t think preparing for those risks requires us to look past what is already happening.

We have years of lessons from trust and safety, cybersecurity and election integrity about what happens when new technology meets adversarial human behavior. AI gives us an opportunity to apply those lessons earlier, and build the safeguards and operational capacity before the teams responsible for managing the consequences are already underwater.

Get a demo