All integrations

Garak Integration with DefectDojo

Garak Integration with DefectDojo

garak is an open source LLM vulnerability scanner maintained by NVIDIA. It sends probes to a target model (a hosted API, a Hugging Face model, or another supported generator), runs detectors over the responses, and records the attempts where a detector fired. Its probe families cover prompt injection, jailbreaks, encoding tricks, data leakage and replay, malware generation, cross-site scripting in model output, toxicity, and misleading content. Each run writes a JSON Lines report plus a separate hit log, and DefectDojo imports the hit log.

Garak Integration with DefectDojo

Our AI features go through garak before every model or prompt change, and the results go into DefectDojo so they are handled like any other security finding. A garak run produces a lot of output, and on its own it's hard to say whether a jailbreak that worked last month still works. In DefectDojo, each failing probe becomes a Finding attached to the model under test, with a CWE, a severity, the prompt and output that triggered it, and a status that changes when a rescan stops reproducing it.

Why Garak Matters

Applications built on LLMs inherit a class of weaknesses that traditional scanners don't test: instructions hidden in user input, guardrails that give way under role-play, and models that repeat sensitive training data.

  • garak tests the model's behavior directly, using a large library of probes instead of a handful of hand-written prompts.
  • Its probe and detector structure makes results repeatable enough to compare across model versions, system prompts, and providers.
  • It is open source and runs locally, which suits teams that can't send test traffic through a third-party testing service.
  • Results identify the probe, the detector, the target, and the exact prompt and output, which is what an engineer needs to reproduce the problem.

Advantages of This Integration

What DefectDojo adds on top of a garak hit log:

  • Aggregation that matches how people triage. Hits for the same probe, target, and detector are combined into one Finding, with the occurrence count recorded and the most severe rung retained.
  • CWE mapping for AI weaknesses. Prompt injection families map to CWE-1427, XSS to CWE-79, leakage and divergence to CWE-200, and everything else to CWE-1426, so LLM findings appear in CWE-based reporting.
  • Stable deduplication across runs. garak samples prompts non-deterministically, so the Garak Scan type deduplicates on title (probe and goal) and component name (the target model) only. Rescanning the same model doesn't create a fresh set of Findings each time.
  • Severity you can attach SLAs to. Detector scores become Info through Critical, adjusted by probe family, so a working jailbreak gets a tighter remediation window than a quality issue.
  • One view of application risk. LLM findings sit beside SAST, SCA, and DAST results for the same Asset, with assignment, notes, Jira tickets, and risk acceptance.

How This Integration Works

DefectDojo imports garak results with the Garak Scan scan type.

1. Run garak against your target. For example, against a Hugging Face model with one jailbreak probe:

python3 -m garak --target_type huggingface --target_name gpt2 --spec probes.dan.Dan_11_0

garak prints the name of its report at the start and end of the run. Alongside the report.jsonl it writes a hit log named garak.<run_id>.hitlog.jsonl. Upload the hit log, not the report.

2. Import the hit log. In the UI, open the Engagement, choose Import Scan Results, select Garak Scan, and upload the .hitlog.jsonl file. For automation, use the API in Community Edition or DefectDojo Pro:

curl "https://YOUR_INSTANCE/api/v2/import-scan/" 
  -H "Authorization: Token $DD_API_TOKEN" 
  -F "scan_type=Garak Scan" 
  -F "file=@garak.run.hitlog.jsonl" 
  -F "product_name=support-assistant" 
  -F "engagement_name=LLM Red Team" 
  -F "auto_create_context=true"

DefectDojo Pro users can use Universal Importer:

universal-importer import 
  --defectdojo-url "https://YOUR_INSTANCE.cloud.defectdojo.com/" 
  --scan-type "Garak Scan" 
  --report-path "./garak.run.hitlog.jsonl" 
  --product-name "support-assistant" 
  --engagement-name "LLM Red Team" 
  --auto-create-context

3. Reimport after each change. Send later hit logs for the same target to /api/v2/reimport-scan/ against the same Test so probes that no longer succeed are mitigated.

The parser was tested against the garak 0.15.x hit log schema. Lines without a probe field, such as run metadata, are ignored, and a hit log with no hits imports zero Findings.

Data Granularity: What Gets Imported

DefectDojo Field Source in garak Hit Log Notes
Title probe and goal Formatted as probe: goal, truncated at 255 characters
Severity Detector score, adjusted by probe family See mapping below
Description Goal, probe, detector, score, target Plus the prompt text, model output, and triggers
CWE Probe family 1427, 79, 200, or 1426 by default
References Probe family Link to the garak reference page for that probe family
Component Name generator The model or target that was scanned
Vuln ID from Tool probe Full probe name
Unique ID from Tool Probe, generator, detector The aggregation key
Occurrences Count of hits Hits aggregated into the Finding
Tags garak, probe family, detector family For filtering
Finding type Static As set by the parser
Deduplication Hash code Title and component name

Severity mapping: a score of 0.9 or higher starts at High, 0.7 at Medium, 0.4 at Low, and anything lower at Info. A hit with no score is treated as 1.0. Active attack families (dan, promptinject, latentinjection, exploitation, malwaregen, xss) move up one rung, and content families (misleading, snowball, continuation, toxicity) move down one.

Use Cases

Gating a model upgrade: A team switching to a new model version runs the same garak probe set against both. Importing each into its own Test shows which jailbreaks the new model resists and which new ones appear, with Critical findings blocking the rollout.

Testing system prompt changes: Product teams iterate on system prompts often. A scheduled garak run with reimport into a fixed Test shows whether a change reopened a prompt injection that was previously fixed.

AI red team reporting: A red team imports garak hits alongside manual findings into the application's Engagement. Leadership gets one report with CWE categories and severities instead of a folder of JSON Lines files.

Comparing providers: Running the same probes against several hosted models, each recorded as its own component, gives a like-for-like view of which provider's model is more resistant for your use case.

Operational Tips

  • Upload the .hitlog.jsonl file. The full report.jsonl is not what this parser reads.
  • Expect findings to cluster at High and above. Many garak detectors are string or word-list matchers that return a score of 1.0, so most real hits land in the upper bands. Review severity during triage rather than treating every High the same.
  • Restrict access to these Findings. Descriptions include the model's actual output, which for some probes is the harmful or offensive content the probe was designed to elicit.
  • Keep the target name consistent across runs. The generator is part of the deduplication hash, so renaming the target creates a new set of Findings.
  • Use separate Tests per model or per system prompt variant, and reimport into them, so history stays readable.
  • Treat CWE assignments as a starting point. The mapping is coarse by design, and you can refine it on individual Findings when your reporting needs more detail.