AI Model Red Teaming: How Labs Stress-Test Before a Launch
What red teaming actually involves, why the specifics matter more than the label, and how to read a model card’s safety-testing disclosure.
Jump to 6 sections
Quick answer: Red teaming is the practice of deliberately trying to make an AI model fail — produce harmful content, leak training data, get manipulated by adversarial prompts — before it ships. Labs run this internally, contract external specialists, and increasingly open limited external access, but the depth and independence of red teaming varies enormously between launches, and a launch announcement rarely says which kind was actually done.
This roundup covers what red teaming actually involves, why the specifics matter more than the fact that "red teaming was conducted," and what to look for in a model's launch documentation to judge how seriously it was tested. It's for anyone evaluating a model for deployment in a context where failure modes matter.
What Red Teaming Actually Involves
At its core, red teaming means assigning people (or automated systems) to actively try to break a model rather than passively evaluate it. This includes attempting to extract harmful content through adversarial prompting, testing whether safety guidelines can be circumvented through indirect phrasing or role-play framing, probing for training data leakage, and testing behavior under edge cases the model wasn't explicitly trained to handle.
This is meaningfully different from standard evaluation, which typically measures how well a model performs on tasks it's supposed to do well. Red teaming measures how badly it can be made to fail at things it's supposed to refuse or avoid — a distinct skill set, closer to security penetration testing than to benchmark scoring.
At The Model Drop, we've reviewed red-teaming disclosures across a range of major model launches, and the depth varies more than most launch announcements let on. Some labs describe months of structured adversarial testing with external specialists across multiple risk categories; others describe a comparatively brief internal review before public release, using language vague enough to sound similar on the page.
Internal vs. External vs. Automated Red Teaming
Three distinct approaches show up across the industry, and they surface genuinely different classes of failure.
- Internal red teams — employees of the lab itself, often with deep model access, testing against known risk categories. Fast and well-informed about the model's architecture, but prone to the same blind spots as the team that built the model.
- External/independent red teamers — outside specialists, sometimes academic researchers or dedicated safety organizations, given structured access before public release. Slower and more expensive to coordinate, but consistently surface failure modes internal teams miss, precisely because they don't share the same assumptions about how the model is meant to be used.
- Automated red-teaming tools — other models or scripted systems generating large volumes of adversarial prompts automatically. Scales coverage far beyond what human testers can manage, but tends to miss the creative, context-specific attacks a motivated human comes up with.
The strongest pre-launch testing programs combine all three: automated tools for broad coverage, internal teams for architecture-specific probing, and external testers for the blind spots neither of the first two can see on their own.
Reading a Model Card's Red-Teaming Section
Model cards and launch documentation vary widely in how specific they are about red-teaming methodology. A few concrete things are worth checking.
Does it name who did the testing? "External red teamers with expertise in biosecurity and cybersecurity" is a meaningfully stronger disclosure than "the model underwent extensive safety testing," which could describe almost anything.
Does it give a timeframe or scope? A disclosure mentioning weeks or months of structured testing across specific risk categories carries more weight than an undated, unscoped claim.
Does it disclose what was found, not just that testing happened? The most credible disclosures describe specific categories of issues found and how they were mitigated, rather than only asserting that testing occurred with a clean result.
Why Pre-Launch Red Teaming Isn't the End of the Story
| Testing type | Timing | Coverage | Limitation |
|---|---|---|---|
| Pre-launch red teaming | Before public release | Known risk categories, model's stated use cases | Can't anticipate every real-world deployment context |
| Post-launch monitoring | Ongoing after release | Real usage patterns at scale | Reactive, not preventative |
| Domain-specific red teaming | Before your own deployment | Your specific use case and risk profile | Rarely done by teams outside the lab itself |
A model that passed a lab's pre-launch red teaming has had its general-purpose failure modes reduced, not eliminated for every possible deployment. A model deployed in a specific domain — medical information, legal guidance, financial advice — carries risk profiles the lab's general red-teaming pass likely didn't specifically target. Organizations deploying models in these contexts increasingly run their own domain-specific adversarial testing before launch, rather than relying solely on the lab's disclosure.
The National Institute of Standards and Technology has published guidance on this exact gap through its AI Risk Management Framework, which frames red teaming as one control among several needed across a model's full deployment lifecycle, not a one-time pre-launch checkbox. Academic research on adversarial robustness, much of it indexed through arXiv's security and cryptography listings, has documented similar findings — that red-teaming coverage tends to be strongest for the risk categories a lab already anticipated, and weakest for genuinely novel attack patterns.
If you're evaluating whether a model's launch claims hold up more broadly, our guide to reading a model launch announcement covers the adjacent skill of separating substantiated claims from marketing framing.
What This Looks Like When It Goes Wrong
A useful contrast is what happens when red teaming is thin. A model that skips independent external testing tends to ship with failure modes clustered around whatever the internal team already anticipated — the risks it was specifically looking for — while missing failure modes that require an outside perspective to spot, like a phrasing pattern that bypasses a safety filter through a use case the internal team simply never considered testing.
We've seen this pattern show up publicly more than once: a model passes its stated pre-launch review, ships, and within weeks users discover a jailbreak technique that a dedicated external red team, given the same access before launch, likely would have caught. This isn't necessarily evidence of a lab cutting corners deliberately — internal teams are genuinely limited by their own familiarity with the system. It's a structural argument for why the internal-only approach has a ceiling that external testing helps push past, the same structural skepticism our benchmarks roundup applies to capability claims that go unverified by anyone outside the lab that made them.
For a broader sense of which labs have published the most detailed safety documentation to date, our frontier model families overview is a useful starting reference, though it's worth checking each lab's own model card directly rather than relying on any single roundup as the final word.
The Bottom Line
"Red teaming was conducted" is close to meaningless on its own — the specifics of who did it, how long it took, what risk categories it covered, and what was found are what actually tell you how seriously a model was tested before launch. For general-purpose use, a lab's disclosed red-teaming program is a reasonable starting signal. For any deployment in a specific, consequential domain, treat the lab's testing as a baseline and plan your own domain-specific evaluation on top of it rather than assuming general-purpose testing covered your use case.
The Model Drop reviews model launch documentation closely, including safety disclosures, rather than taking a launch announcement's claims at face value.