All integrations

PyRIT Integration with DefectDojo

PyRIT Integration with DefectDojo

PyRIT (Python Risk Identification Tool for generative AI) is an open source framework from Microsoft, released under the MIT license, for red-teaming generative AI systems. Security engineers use it to run automated, often multi-turn attacks against a model or AI application, pursue a stated objective, and score the target's responses to decide whether the attack worked. PyRIT records conversations, scores, and attack results in its memory store (SQLite or Azure SQL) rather than writing a report file, and those attack results can be exported as JSON.

PyRIT Integration with DefectDojo

We use PyRIT to test our AI features the way an attacker would, with many attempts and different strategies against the same objective. Importing the exported attack results into DefectDojo means a successful jailbreak becomes a High Finding on the Asset that owns that AI feature, with the objective, attack strategy, number of turns, scorer rationale, and last response kept in the description. Attacks the model refused never show up as Findings, so the list our engineers see is the set of places where the guardrails did not hold.

Why PyRIT Matters

Testing a generative AI system by hand does not scale. A single objective may need dozens of rephrasings, multi-turn conversations, or converted prompts before a model gives way.

  • PyRIT automates attack strategies and runs them repeatedly against a target, so coverage is not limited to what one tester can type.
  • Scorers judge each final response, which gives every attack a recorded verdict instead of an informal impression.
  • Each result keeps a conversation id, so the full exchange can be reviewed in PyRIT's memory later.
  • Results in a memory store are useful to the person who ran the campaign, but not to the team that has to fix the system or the leader who has to report on it.

Advantages of This Integration

What routing PyRIT through DefectDojo adds:

  • Only real failures become Findings. Attacks with outcome failure (the target refused) are not imported. Successes import as High, undetermined outcomes as Low for human review, and infrastructure errors as Info.
  • Severity grounded in behavior. Severity reflects whether the guardrails held, not how alarming the objective sounds.
  • Harm categories as tags. The targeted harm categories and an outcome tag such as pyrit-outcome-success are attached, so campaigns can be filtered by harm.
  • Evidence trail. The conversation id and attack result id in each description point back to PyRIT's memory store for the full transcript.
  • Deduplication across campaigns. DefectDojo hashes the title and attack strategy, so rerunning the same objective with the same strategy updates the existing Finding.
  • A lifecycle for fixes. Reimporting a new campaign into the same Test mitigates attacks that no longer succeed and reactivates any that return.

How This Integration Works

DefectDojo imports PyRIT results with the PyRIT Scan scan type.

1. Export attack results to JSON. After a campaign, export the attack results from your PyRIT memory store to a JSON file. The parser accepts a JSON array of attack results, or an object holding them under results or attack_results. Each result should carry fields such as objective, outcome, atomic_attack_identifier, executed_turns, conversation_id, targeted_harm_categories, last_score, and last_response.

2. Import the file. In the UI, open the Engagement, choose Import Scan Results, select PyRIT Scan, and upload the JSON. Through the API, in Community Edition or DefectDojo Pro:

curl "https://YOUR_INSTANCE/api/v2/import-scan/" 
  -H "Authorization: Token $DD_API_TOKEN" 
  -F "scan_type=PyRIT Scan" 
  -F "file=@pyrit-results.json" 
  -F "product_name=support-copilot" 
  -F "engagement_name=AI Red Team" 
  -F "auto_create_context=true"

DefectDojo Pro users can import from a pipeline with Universal Importer:

universal-importer import 
  --defectdojo-url "https://YOUR_INSTANCE.cloud.defectdojo.com/" 
  --scan-type "PyRIT Scan" 
  --report-path "./pyrit-results.json" 
  --product-name "support-copilot" 
  --engagement-name "AI Red Team" 
  --auto-create-context

3. Reimport each campaign. Send later exports to /api/v2/reimport-scan/ against the same Test so you can see which attacks stopped working after a guardrail change.

Data Granularity: What Gets Imported

DefectDojo Field Source in PyRIT Attack Result Notes
Title Attack strategy and objective strategy: objective, or just the objective
Severity outcome success High, undetermined Low, error Info; failure not imported
Description Objective, outcome, reason, strategy, harms, turns Also conversation id and attack result id
Description (score) last_score Score type, value, category, rationale
Description (response) last_response.original_value The target's final response
Description (errors) error_type, error_message Added for error outcomes
Component Name atomic_attack_identifier name The attack strategy
Vuln ID From Tool Attack strategy name Same value as the component
Tags Outcome and targeted_harm_categories For example pyrit-outcome-success
Finding type Dynamic All PyRIT findings are marked dynamic
Deduplication Hashcode Title and vuln ID from tool

A missing outcome is treated as undetermined. A float_scale score is recorded but not used for severity, because scorer scales differ between campaigns.

Use Cases

Before launching an AI feature: A product team runs PyRIT campaigns against a staging deployment and imports the results. Any High Finding is a demonstrated guardrail failure that has to be fixed or formally risk-accepted before launch.

After a model or prompt change: The same campaign is rerun and reimported. Findings that no longer reproduce are mitigated, and new successes appear as new Findings with the scorer's rationale.

Reviewing uncertain results: Undetermined outcomes import as Low, so a reviewer can work through them, read the last response in the description, and either raise the severity or mark the Finding as a false positive.

Reporting AI risk: Security leads tracking several AI systems use the harm category tags to report open guardrail failures by category across Assets.

Operational Tips

  • Keep objectives worded consistently between campaigns. The title includes the objective text, so rewording it creates a new Finding.
  • Use one Test per target and campaign definition, and reimport into it, so mitigations reflect real fixes.
  • Error outcomes import as Info. Filter them out with minimum_severity=Low, or keep them to spot rate limits and timeouts that left parts of a campaign untested.
  • Route undetermined Findings to a human reviewer. PyRIT could not decide, so the decision belongs in DefectDojo's notes and status history.
  • Store the PyRIT memory store alongside the campaign. The conversation id in each Finding is only useful if the transcript is still available.
  • Add tags at import time for the model version or deployment, so results can be compared across releases.