Skip to main content

Mozilla protected your data; now we protect your LLM

J
By Joe McBride

Today Mozilla 0DIN is launching 0DIN AI Defense, a line of self-hostable AI security solutions built from real attacker data, starting with the SusFactor jailbreak detection model. The early-access beta opens now.

Small, potent and fast

Three factors set it apart.

First, at only 560M parameters, it can run locally in your environment. It fits comfortably on cheap commodity hardware and it's fast enough to analyze prompts in a realtime process. No data ever has to leave your infrastructure.

Second, it is trained on real attacks from the 0DIN Threat Feed. We retrain the model continuously as new submissions arrive, so your AI defense can always stay on top of the latest offensive tactics.

Does it hold up? On the public JailbreakBench benchmark, SusFactor blocks all but 12.1% of attacks while holding a 9.0% false alarm rate, one of the lowest false alarm rates of any detector in our evaluation that actually catches attacks. For comparison, WildGuard sits at 45%, Mistral at 39%, and Llama Guard 3 at 23%. It detects attacks by intent rather than by surface technique, so it stays consistent across attack types where technique-based detectors like Meta's Prompt Guard 2 collapse. It also holds its own against 7-8 billion parameter generative guards more than 10x its size. Full results and the details of our evaluation methodology are coming soon in the companion research paper.

Finally, it's lightning fast. We consistently clocked it at sub 50ms on recommended CPU-based hardware under specified conditions. Run it in production on your most critical workloads with confidence.

Explore AI security with the Scanner Datasheet

The datasheet offers insight into the challenges and solutions in AI security.

Download Datasheet

The wall every team hits

Every privacy-oriented team shipping an LLM feature into production hits the same wall: there is no guardrail they can both trust and run themselves. The options are cloud-only, slow, opaque, or untested against the attacks people are actually running today. You either send your users' prompts to someone else's API, or you pay for expensive GPU time, or you go without.

We built SusFactor to close that gap, and in doing so we started a new line of products we call 0DIN AI Defense.

Why an AI bug bounty is building Defense

0DIN began on offense. For two years, more than 2,000 researchers in the 0DIN Bug Bounty program have been finding real jailbreaks and prompt injections against production AI systems, and every submission flows into our Threat Feed jailbreak dataset.

But offense was never the whole story. Our customers do not want a list of ways their model can be broken. They want actionable tools to help them ship AI-based systems safely, and to hand their own customers something that says, credibly, here is how we defend this.

That is the gap Defense fills. You cannot build a great guardrail without having attacked at scale first. Most detection models are trained on synthetic data, educated guesses about what an attack might look like. SusFactor is trained on the real thing. That closes the loop from offense to defense: the same program that finds the attacks now trains the model that detects them.

What SusFactor does

SusFactor is 0DIN's own classifier. It reads each incoming prompt and outputs a single suspiciousness score, from 0 for benign to 1 for likely malicious. Your system compares that score against a threshold you set and decides what to do: let it through, log it, or reject it before it ever reaches your LLM.

Safeguard Your GenAI Systems

Connect your security infrastructure with our expert-driven vulnerability detection platform.

Where Defense goes next

SusFactor is the first model in Defense, not the last. Next come response-side scoring and multi-turn detection, reporting and attack-intel enrichment, and attack categorization mapped to your risk model.

It is also built to meet you where your stack already lives, starting with the LiteLLM proxy. We cover the full integration story, what is available today and what is on the roadmap, in an upcoming post.

The through-line is simple: offense feeds defense, and every new attack our researchers find makes the models that protect you sharper.

Get into the early-access beta

SusFactor is available now as an early-access beta. We are opening it in limited waves, so signing up adds you to the waitlist and we will reach out as spots open.

If you are putting GenAI into production and want a guardrail you can trust and run yourself, get in line.

Join the early-access beta waitlist

A note about data privacy and the limited beta instance. Do not send personal data or real production traffic. By default, your trial prompts are not used to train our models nor are they retained.

Secure People, Secure World.

Discover how 0DIN helps organizations identify and mitigate GenAI security risks before they become threats.

Request a Demo