Skip to main content

Red Teaming

Find the way in before
somebody else does.

Attack your agent the way an adversary would, in one turn or across a whole conversation, and track a safety score from one release to the next.

Red Teaming in the Netra dashboard

Trusted by teams shipping agents in production

What is Red Teaming?

Adversarial testing for your agent, on every release. An attacker model writes attacks from your agent's stated purpose, a judge model scores every reply, and the results roll up into a safety score. You can rerun the same setup on every release and see whether you got safer or not.

An adversary on your release checklist

Save the target, the attack catalogue and the models once, then rerun the same red team on every release.

Save reusable configurations

Keep the agent, evaluators, tests per evaluator and models as a setup, and rerun it every release.

Save reusable configurations: Docs (opens in new tab)

Pick separate attacker and judge models

The attacker writes the prompts and the judge scores the replies. Change one without touching the other.

Pick separate attacker and judge models: Docs (opens in new tab)

Set how many attacks each evaluator runs

Two by default. Raise it for more coverage as a release gets closer.

Set how many attacks each evaluator runs: Docs (opens in new tab)

Attacks written for your agent

Each evaluator's template takes in your agent's stated purpose, so the attacks go after what your agent actually does.

Run single and multi-turn agent attacks

One-shot prompts test individual guardrails. Multi-turn attacks probe escalation and context manipulation.

Run single and multi-turn agent attacks: Docs (opens in new tab)

Test automatically with an autonomous attacker

An attacker agent reads your agent's purpose and keeps probing, with no prompt list to maintain.

Test automatically with an autonomous attacker: Docs (opens in new tab)

Tailor attacks to your agent's purpose

Generation templates use what your agent is for, so a banking agent gets banking attacks.

Tailor attacks to your agent's purpose: Docs (opens in new tab)

Ship with a safety score, not a hope

Every run rolls up into one score, where 100% means every attack was blocked. It breaks down by suite and evaluator to the category that needs work.

See whether a release made you safer

Each run is compared with the one before it, so a regression shows up as a delta on the release that caused it.

See whether a release made you safer: Docs (opens in new tab)

Share reports by suite and evaluator

A breakdown you can hand to engineering, and a history you can hand to compliance.

Share reports by suite and evaluator: Docs (opens in new tab)

Spot the evaluators that fail

The count of evaluators with at least one vulnerable result in the latest run.

Spot the evaluators that fail: Docs (opens in new tab)

Evidence for engineering and for compliance

Every attack, response, judge score and explanation is kept, multi-turn transcripts included.

Read detailed results

Pass, fail or error for every attack, with the judge's score and its explanation.

Read detailed results: Docs (opens in new tab)

Replay multi-turn transcripts

See the whole exchange, turn by turn, and where the agent gave way.

Replay multi-turn transcripts: Docs (opens in new tab)

Trace every attack

Runs emit telemetry, so a successful attack opens in the same trace view you use for production.

Trace every attack: Docs (opens in new tab)

Open source

Red teaming you can read the source of

Open source · Apache 2.0

Agent OPFOR — adversary emulation for AI agents

The adversarial framework we red team agents with, open source on GitHub. Run it yourself, read how every attack is generated, and contribute new ones.

View on GitHub (opens in new tab)

How a run works

  1. 01Point it at an agent registered in your project
  2. 02Pick suites aligned to security frameworks, or single evaluators
  3. 03An attacker LLM generates prompts tailored to your agent's purpose
  4. 04A judge LLM scores every response, single- or multi-turn
  5. 05Results roll up into a safety score you can track release to release
Frequently Asked Questions

Everything You Need to Know About Red Teaming

Can't find the answer here? The Red Teaming docs go deeper, or talk to our team.

What does red teaming test for?

Adversarial behaviour such as jailbreaks and system-prompt leakage, organised into suites aligned to security frameworks — including the EU AI Act — or run as individual evaluators for a specific attack category.

What's the difference between single-turn and multi-turn attacks?

Single-turn sends each adversarial prompt once and suits testing individual guardrails. Multi-turn has the attacker model sustain a conversation, which tests escalation resistance and context manipulation; each turn is judged independently.

How is the safety score calculated?

Each evaluator's score is the percentage of its tests that passed. A suite's score is the average of its evaluators, and the overall safety score averages every distinct evaluator in the run. Errored or cancelled results are left out.

What do I need before my first run?

An agent registered in your project, an attacker model and a judge model configured, and at least one suite or evaluator selected.

Is the attack framework open source?

Yes. Agent OPFOR, the adversarial framework we red team agents with, is on GitHub under Apache 2.0.

How do I know whether a release made things worse?

Score change is always computed against the immediately previous completed run, and every run is plotted on the score history — so a drop is visible on the release that caused it.

Start today

Find the way in on your own terms

Point a red-team run at your agent and get a safety score before your next release. Free to start.