Skip to main content

Prompt Management

The instruction is the product.
Treat it that way.

Version every prompt, promote it with a label instead of a release, and prove it is better before a single customer sees it.

Prompt Management in the Netra dashboard

Trusted by teams shipping agents in production

What is Prompt Management?

One home for every prompt your agents run. Each prompt is stored as an immutable version. Labels such as production and staging point at the version that is live, your agent fetches it by name at runtime, and every trace records which version answered.

Let the people who own the behaviour own the prompt

Product managers and domain experts iterate in the dashboard; engineers stop being a bottleneck on copy changes.

Manage prompts in one library

Browse, open, fork or start a prompt from a single list per project. Nothing lives in a string constant any more.

Manage prompts in one library: Docs (opens in new tab)

Test with different model configurations

Choose a provider and model, then tune temperature, max tokens and top P. The settings save with the version, so it reproduces exactly.

Test with different model configurations: Docs (opens in new tab)

Prove a prompt is better before it reaches customers

Run the new version through evaluations and simulated conversations first — promote on evidence, not on a hunch.

Find which model works best for a prompt

Compare latency, cost, tokens and scores per model, with a radar across evaluators, and choose on quality, speed and cost together.

Find which model works best for a prompt: Docs (opens in new tab)

Test prompts with evaluators of your choice

Score every run with LLM-as-Judge, latency, cost, token count, JSON validation or regex, each with its own pass criteria.

Test prompts with evaluators of your choice: Docs (opens in new tab)

Compare prompts side by side

Put two versions next to each other and review a GitHub-style diff before anything is published.

Compare prompts side by side: Docs (opens in new tab)

Change a prompt without shipping a release

Prompts are fetched at runtime by name and label, so a wording fix goes live in minutes instead of waiting on the next deploy window.

Version management

Every publish becomes an immutable version with its history and metadata. Draft from any version without touching what is live.

Version management: Docs (opens in new tab)

Promote prompts without code changes

Point production, staging or your own labels at a version. Your app fetches by label, so promoting never needs a deploy.

Promote prompts without code changes: Docs (opens in new tab)

Roll back a bad prompt before it costs you a day

Every version is immutable and every deployment is a label — move production back one version and the regression is gone.

Roll back a bad prompt before it costs you a day: Docs (opens in new tab)

Move fast without adding latency

Opt-in client-side caching serves prompts from memory, so centralising them costs you nothing on the hot path.

Move fast without adding latency: Docs (opens in new tab)

Know which prompt version caused the regression

Every trace carries the prompt version behind it, so a quality drop bisects to the exact change instead of a week of guessing.

Trace every answer to its prompt version

Filter traces, dashboards and evaluations by the prompt version behind them, and see which change moved a metric.

Trace every answer to its prompt version: Docs (opens in new tab)

Compare versions on live traffic

Send prompt_version as a span attribute and compare versions side by side in traces and dashboards.

Compare versions on live traffic: Docs (opens in new tab)

See reliability across versions

Every stress test gets a health verdict, and the score history shows whether each version made things better.

See reliability across versions: Docs (opens in new tab)

Set up alerts for changes

Get notified when a prompt is published or a label moves, so nothing changes in production unannounced.

Set up alerts for changes: Docs (opens in new tab)

Fetch prompts from the SDK you already trace with

One call returns the version a label points at, and needs nothing beyond Netra.init(). Your agent gets the prompt it should run, and a default if the network blinks.

  • Python SDK
  • TypeScript SDK
  • OpenAI
  • Anthropic
  • Google Gemini
  • Mistral
  • AWS Bedrock
  • Groq

Fetch by name and label

get_prompt(name, label) returns the version a label points at, and uses production when you don't name one.

Docs (opens in new tab)

Fail safe by design

The prompt client is built never to throw. Errors are logged and a safe fallback comes back, so a network blip never becomes an outage.

Docs (opens in new tab)

Cookbook

Pick the winner on evidence

Cookbook · Evaluation

A/B test two prompt or model configurations

Run the same test cases against each configuration, read the scores side by side, and weigh quality against cost and latency before you promote.

Open the cookbook (opens in new tab)

What you'll build

  1. 01Add Answer Correctness and Conciseness evaluators from the library
  2. 02Build an evaluation from a handful of real traces
  3. 03Trigger one test run per configuration from the SDK
  4. 04Compare evaluator scores and the traces behind them
  5. 05Weigh quality against cost and latency, then decide
Frequently Asked Questions

Everything You Need to Know About Prompt Management

Can't find the answer here? The Prompt Management docs go deeper, or talk to our team.

How does my application get the current version of a prompt?

Through the SDK: get_prompt(name, label) fetches the version a label points at, and defaults to production when no label is given. It needs nothing beyond your usual Netra.init().

What happens if Netra can't be reached when my app fetches a prompt?

The prompt client is designed never to throw. Errors are logged and a safe fallback is returned — None for an unknown name, an empty object on a network error — so keep a default prompt in code for that path.

Can I edit a prompt after it's published?

No. Published versions are immutable, which is what makes a trace reproducible. Create a draft from any version, iterate, and publish it as a new version with its diff.

How do labels work?

A label is a pointer to a version — production, staging, latest or anything you define. Promoting or rolling back is moving the label; your application code never changes.

How do I test a prompt before I publish it?

Run a stress test from Prompt Studio: pick up to five models and 1–100 runs per model, switch on evaluators such as LLM-as-Judge, latency, cost, token count, JSON validation or regex, and read the per-model results and health verdict. It works on drafts as well as published versions.

How do I compare quality across prompt versions in production?

Send the prompt version as an attribute on your spans and it becomes a dimension you can filter and compare on across traces and dashboards. Stress-test history also plots reliability scores version over version.

Start today

Stop waiting on a deploy to fix a sentence

Move your prompts into Netra, version every change, and promote the one that proves itself. Free to start.