ModelScan Integration with DefectDojo
ModelScan Integration with DefectDojo
ModelScan is an open source scanner for serialized machine learning models, created by Protect AI (now part of Palo Alto Networks). It inspects model files saved in formats such as pickle, TensorFlow SavedModel, Keras, and NumPy for operators that can execute code when the model is loaded, and grades each unsafe operator it finds from Low to Critical. It can write its results as JSON, which is the format DefectDojo imports.
ModelScan Integration with DefectDojo
We started running ModelScan once our data science teams began pulling pretrained models from public hubs into production services. A model file is code as far as the loader is concerned, and nobody was reviewing it like code. Importing ModelScan JSON into DefectDojo gives each unsafe operator a Finding tied to the specific model file on the right Asset, with the severity ModelScan assigned, so model risk sits in the same triage queue as our dependency and container findings.
Why ModelScan Matters
Several popular model serialization formats can carry instructions that run on load. A model downloaded from a public repository, or tampered with in storage, can execute code on whatever machine loads it, from a data scientist's laptop to a production inference service.
- It looks inside the model file for operators capable of code execution, rather than relying on where the file came from.
- It covers the formats teams commonly exchange, including pickle-based files, TensorFlow SavedModel, Keras, and NumPy.
- Each issue names the module and operator, which tells a reviewer exactly what the model would call on load.
- Its severity scale distinguishes clearly dangerous operators from lower-risk ones, which helps when a scan of a model directory returns many results.
- A ModelScan report covers one run. Tracking which models were reviewed, replaced, or accepted needs a platform that keeps history.
Advantages of This Integration
What running ModelScan through DefectDojo adds:
- Severity carried through. ModelScan's CRITICAL, HIGH, MEDIUM, and LOW map directly to DefectDojo severities, so SLA rules apply to model findings without extra configuration.
- Per-file tracking. The model file path is stored as the Finding's file path and component, so you can see which model artifacts carry which operators.
- Stable deduplication. DefectDojo hashes ModelScan findings on title, file path, and the operator recorded as the tool's vulnerability ID, so rescanning the same model does not create duplicates.
- Lifecycle on reimport. When a team replaces a pickle model with a safer format and reimports, the old Findings are mitigated. A model that reintroduces an operator reactivates its Finding.
- Documented exceptions. Some internal models need an operator ModelScan flags. Risk acceptance with an expiry date records that decision instead of leaving the Finding open indefinitely.
- One view of AI supply chain risk. Model findings sit alongside SCA and container findings for the same Asset, which is useful for reviews of services that serve models.
How This Integration Works
DefectDojo reads ModelScan output with the ModelScan Scan scan type.
1. Produce a JSON report. Point ModelScan at a model file or directory and write JSON:
modelscan -p ./models -r json -o modelscan.json
2. Import it. In the UI, open the Engagement, choose Import Scan Results, select ModelScan Scan, and upload the file. For automation, the API works in Community Edition and DefectDojo Pro:
curl "https://YOUR_INSTANCE/api/v2/import-scan/"
-H "Authorization: Token $DD_API_TOKEN"
-F "scan_type=ModelScan Scan"
-F "file=@modelscan.json"
-F "product_name=recommendation-service"
-F "engagement_name=Model Intake"
-F "auto_create_context=true"
DefectDojo Pro users can run the same import with Universal Importer:
universal-importer import
--defectdojo-url "https://YOUR_INSTANCE.cloud.defectdojo.com/"
--scan-type "ModelScan Scan"
--report-path "./modelscan.json"
--product-name "recommendation-service"
--engagement-name "Model Intake"
--auto-create-context
3. Reimport when models change. Send later scans of the same model directory to /api/v2/reimport-scan/ against the same Test, so replaced models close their Findings.
The parser reads the issues array. Each issue becomes one Finding, and the ModelScan version from the report summary is recorded in the description.
Data Granularity: What Gets Imported
| DefectDojo Field | Source in ModelScan Report | Notes |
|---|---|---|
| Title | Operator and module | "Unsafe operator 'X' from module 'Y'" |
| Severity | Issue severity |
CRITICAL, HIGH, MEDIUM, LOW map directly; anything else becomes Medium |
| Description | Issue description, operator, model file, scanner, ModelScan version | Operator shown as module.operator |
| Mitigation | Fixed text | Load only from a trusted source, or re-serialize in a format that cannot carry code, such as safetensors |
| File Path | Issue source |
The model file the operator was found in |
| Component Name | Issue source |
Same model file path |
| Vulnerability ID from tool | Module and operator | Recorded as module.operator |
| CVE / CWE | Not set | ModelScan reports unsafe operators, not known vulnerabilities |
| Finding type | Static | ModelScan reads files without loading them |
| Deduplication | Hashcode | Title, file path, vulnerability ID from tool |
Use Cases
At model intake: Before a pretrained model from a public hub is approved for use, a pipeline scans it and imports the result into an intake Engagement. Reviewers decide on each Finding, and the decision is kept with the model file path.
In a training pipeline: Every training job that produces a new model artifact runs ModelScan on the output directory and reimports into a Test for that model. An unexpected operator in a fresh artifact shows up as a new Finding.
For model registries: A platform team scans the models stored in its internal registry on a schedule. DefectDojo shows which teams still publish pickle-based models with flagged operators, which helps plan migration to safer formats.
During a security review: For a service that serves models, security can show ModelScan findings next to dependency and container findings on the same Asset, with status and history, instead of a separate report.
Operational Tips
- Scan model directories, not just single files, so a model split across several files is covered in one report.
- Use one Test per model or per model directory and reimport into it. Moving a model to a new path changes the file path in the hash and creates new Findings.
- ModelScan findings have no CVE or CWE. Filter and report on the operator recorded as the tool's vulnerability ID instead.
- Where possible, act on the mitigation and re-save models in a format that cannot carry executable code. Risk acceptance should be the exception for models you build yourself.
- Tag imports with the model source (for example
tags=huggingfaceortags=internal) so you can report on externally sourced models separately. - Set
minimum_severity=Highat intake if you want reviewers to see only the most dangerous operators first, then lower it for full reviews.