All integrations

Promptfoo Integration with DefectDojo

Promptfoo Integration with DefectDojo

Promptfoo is an open source command-line tool and library from Promptfoo, Inc. for evaluating and red-teaming applications built on large language models. Teams define prompts, test cases, and providers (the model or application under test), then run evaluations that score each output against assertions, or run adversarial red-team scans that probe for harms such as prompt injection, data leakage, and unsafe content. Results are written as JSON with the -o flag, and Promptfoo Cloud stores evaluations that can be read over its API.

Promptfoo Integration with DefectDojo

We added Promptfoo to our AI release process because a chatbot that leaks system prompts or follows injected instructions is a security bug, and it needs the same handling as any other bug. Importing Promptfoo results into DefectDojo turns each failing probe category into a Finding on the Asset that owns the model integration. Repeated attacks against the same weakness collapse into one Finding with an occurrence count, so a red-team run of several hundred probes produces a list our engineers can actually work through. When a guardrail change lands, the next import shows which weaknesses stopped reproducing.

Why Promptfoo Matters

LLM applications fail in ways traditional scanners were never built to test. A model can be coaxed into revealing data, ignoring instructions, or producing content the business cannot ship, and those failures depend on the prompt, the model, and the guardrails together.

  • Red-team plugins generate adversarial inputs per harm category, so testing covers more than a few hand-written prompts.
  • The same configuration can be run against several providers, which makes it practical to compare a model upgrade before rollout.
  • Each failure carries the attack input, the model output, and the grader's reason, which is the evidence a reviewer needs.
  • On its own, a results file is one run. It does not tell you whether a weakness was already known, accepted, or fixed last sprint.

Advantages of This Integration

What running Promptfoo through DefectDojo gives us:

  • One Finding per weakness per target. Failures for the same red-team plugin against the same provider are aggregated into a single Finding, with the occurrence count reflecting how many attempts succeeded and the most severe rung kept.
  • Stable deduplication across runs. The scan type hashes only the title and component (the target model), so changing attack wording between runs does not create duplicate Findings.
  • CWE context. Plugins are mapped to CWEs (for example CWE-1427 for prompt injection), which lets AI findings sit in the same CWE-based reports as the rest of the portfolio.
  • A lifecycle for guardrail fixes. Reimporting into the same Test mitigates weaknesses that no longer fail and reactivates any that come back.
  • Severity-based SLAs. Promptfoo's red-team severity flows straight into DefectDojo, so a critical jailbreak gets the same SLA clock as a critical CVE.
  • API sync in DefectDojo Pro. The Promptfoo connector reads evaluations directly from Promptfoo Cloud on a schedule.

How This Integration Works

DefectDojo imports Promptfoo results with the Promptfoo Scan scan type, which reads the Promptfoo results schema (results.version 3).

1. Produce a results file. For an adversarial scan:

promptfoo redteam run -o results.json

For a standard evaluation:

promptfoo eval -o results.json

2. Import the file. In the UI, open the Engagement, choose Import Scan Results, select Promptfoo Scan, and upload results.json. With the API (Community Edition and DefectDojo Pro):

curl "https://YOUR_INSTANCE/api/v2/import-scan/" 
  -H "Authorization: Token $DD_API_TOKEN" 
  -F "scan_type=Promptfoo Scan" 
  -F "file=@results.json" 
  -F "product_name=support-assistant" 
  -F "engagement_name=LLM Red Team" 
  -F "auto_create_context=true"

DefectDojo Pro users can run the same import from a pipeline with Universal Importer:

universal-importer import 
  --defectdojo-url "https://YOUR_INSTANCE.cloud.defectdojo.com/" 
  --scan-type "Promptfoo Scan" 
  --report-path "./results.json" 
  --product-name "support-assistant" 
  --engagement-name "LLM Red Team" 
  --auto-create-context

3. Or connect Promptfoo Cloud (DefectDojo Pro). The Promptfoo connector imports red-teaming and evaluation findings from Promptfoo Cloud. Enter https://api.promptfoo.app as the Location, put your Promptfoo API token in the Secret field, and optionally set a Minimum Severity. DefectDojo reads every stored evaluation the token can see, creates a Record for each target application (provider) that was probed, and aggregates failing probes per weakness, so repeated evaluations of the same weakness stay one Finding.

Promptfoo inverts the usual scanner logic, and the parser follows it. A result with success: true means every assertion passed (for a red-team probe, the model defended the attack) and is not imported. Only success: false results become Findings. Results that errored at the provider (failureReason 2) are skipped because the test never ran.

Data Granularity: What Gets Imported

DefectDojo Field Source in Promptfoo Results Notes
Title harmCategory and pluginId For example Hate (harmful:hate); plain eval failures read Failed assertion: <metric or type>
Severity metadata.severity critical, high, medium, low map directly; eval failures without severity default to Medium
Description Plugin, harm category, goal, target, grader reason Also includes the attack input (vars or prompt) and the model output
CWE Plugin or harm category 89 (SQL injection), 78 (shell injection), 1427 (prompt injection or extraction), 200 (PII, privacy), else 1426
Component Name Provider label or id The target model or application
Vuln ID From Tool pluginId Falls back to the failed assertion type
Unique ID From Tool Weakness plus provider Used for aggregation within one file
Occurrences Count of failed attempts Highest severity among them is kept
References shareableUrl Link to the shared Promptfoo view, when present
Tags promptfoo, plugin id, harm category Useful for filtering by harm
Finding type Static All Promptfoo findings are marked static
Deduplication Hashcode Title and component name only

Description and severity are deliberately left out of the hash. The attack text changes every run, and aggregate severity can shift as the set of failed attempts changes.

Use Cases

Before a model upgrade: A team swapping one hosted model for another runs the same red-team configuration against both providers. Because the component is the provider, DefectDojo shows each model's weaknesses side by side on the same Asset.

In a release pipeline: Every build of an assistant runs a red-team scan and reimports the results into a fixed Test. New weaknesses appear as new Findings, and a release check can query the DefectDojo API for open Critical or High items.

Tracking guardrail work: When engineers add an input filter for prompt injection, the next reimport mitigates the CWE-1427 Findings that stopped reproducing, which gives the AI team a record of what the fix actually closed.

Central AI risk reporting: Organizations with several LLM features import each into its owning Asset, so security leadership can report on open AI weaknesses with the same metrics used for application vulnerabilities.

Operational Tips

  • Reimport into one Test per target configuration. The hashcode depends on the provider name, so renaming a provider label creates new Findings rather than matching old ones.
  • Plain promptfoo eval failures carry no severity and import as Medium. Use red-team runs, or adjust severity after import, if eval assertions are guarding security behavior.
  • Set minimum_severity on import (or Minimum Severity on the connector) if low-severity harm probes would crowd out the issues your team acts on first.
  • Read the model output in the description before closing a Finding. A grader can be wrong, and false positive marking keeps that decision on record.
  • Use the plugin and harm category tags to route Findings: privacy failures to one owner, prompt injection to another.
  • An import with zero Findings is a valid result. It means every probe passed or errored, not that the file was ignored.