Dirsearch Integration with DefectDojo
Dirsearch Integration with DefectDojo
dirsearch is an open source web path scanner written in Python and maintained by Mauro Soria. It requests paths from a wordlist against a target web server and reports the ones that exist, which is how teams find exposed admin panels, forgotten backups, .git directories, and endpoints nothing links to. It writes reports in several formats, and its JSON output is what DefectDojo imports.
Dirsearch Integration with DefectDojo
We run dirsearch against our web properties because it finds the things nobody meant to publish. The output is a list of URLs and status codes, which is useful once and tedious every time after that. Importing dirsearch JSON into DefectDojo records each discovered path as a Finding with the URL as its endpoint, matches it against earlier scans of the same host, and lets us mark the harmless ones once instead of rereading them on every run. When a backup directory disappears after cleanup, the next reimport closes it.
Why Dirsearch Matters
Content discovery covers a gap that crawlers and code scanners leave open: files and paths that exist on a server but are not linked from anything.
- It finds artifacts of deployment mistakes, such as source control directories, archives, and old admin interfaces.
- It is fast and scriptable, so it fits into scheduled jobs and pre-release checks.
- It needs nothing from the application team except a URL, which makes it easy to run against third-party or legacy sites where you have no source code.
- Wordlists and status filters are fully under your control, so scans can be tuned to a stack.
- A discovered path is not automatically a vulnerability. Someone has to decide which results matter, and that decision should not be repeated for every scan.
Advantages of This Integration
What running dirsearch through DefectDojo adds:
- Triage once, keep the decision. Every result imports as Info. Mark
/robots.txtas a false positive or risk-accept it once, and later scans of the same endpoint match that existing finding instead of creating a new one to review. - Deduplication by path and endpoint. DefectDojo hashes dirsearch findings on title and endpoints. The title holds the status and path, and the endpoint holds the full URL, so two hosts scanned with the same wordlist stay distinct.
- Stable hashes across rescans. The description, which records response size and other details that change between scans, is deliberately left out of the hash, so an unchanged path does not import again.
- Evidence kept with each finding. The dirsearch command line from the report, including the wordlist and filters, is stored in every finding's description.
- Endpoint-centric view. Findings attach to the scanned URL, so exposed paths show up next to DAST results for the same host.
- Lifecycle on reimport. Paths that no longer respond are mitigated when a new report is reimported into the same Test.
How This Integration Works
DefectDojo imports dirsearch output with the Dirsearch Scan scan type, which expects the JSON report.
1. Run dirsearch with JSON output.
pip install dirsearch
dirsearch -u https://target.example.com/ -w wordlist.txt --format=json -o dirsearch.json
dirsearch also supports simple, plain, CSV, HTML, XML, SQLite, and MySQL output, but only JSON is parsed. Two practical notes from the DefectDojo docs: dirsearch fails to start on Python 3.12 without setuptools installed (Python 3.11 works), and it writes no report file at all when it finds nothing, so a pipeline should check that the file exists before importing.
2. Import the file. In the UI, open the Engagement, choose Import Scan Results, select Dirsearch Scan, and upload the JSON. With the API, available in Community Edition and DefectDojo Pro:
curl "https://YOUR_INSTANCE/api/v2/import-scan/"
-H "Authorization: Token $DD_API_TOKEN"
-F "scan_type=Dirsearch Scan"
-F "file=@dirsearch.json"
-F "product_name=marketing-site"
-F "engagement_name=Content Discovery"
-F "auto_create_context=true"
DefectDojo Pro users can run the same import with Universal Importer:
universal-importer import
--defectdojo-url "https://YOUR_INSTANCE.cloud.defectdojo.com/"
--scan-type "Dirsearch Scan"
--report-path "./dirsearch.json"
--product-name "marketing-site"
--engagement-name "Content Discovery"
--auto-create-context
3. Reimport on a schedule. Send later reports for the same target to /api/v2/reimport-scan/ against the same Test so removed paths are mitigated and recurring ones keep their triage history.
Data Granularity: What Gets Imported
| DefectDojo Field | Source in dirsearch JSON | Notes |
|---|---|---|
| Title | status and URL path |
Formatted as HTTP 301: /admin; "Discovered" when no status |
| Severity | None | Always Info, since dirsearch reports paths that exist, not paths that are wrong |
| Endpoint | url |
The full discovered URL |
| Description | url, status, content-length, content-type |
Labeled lines |
| Description (redirect) | redirect |
Included only when the path redirects |
| Description (command) | info.args |
The dirsearch invocation, including wordlist and filters |
| Finding type | Dynamic | dirsearch requests a live service |
| Deduplication | Hashcode | Title and endpoints |
Each entry in results becomes one Finding.
Use Cases
Before a release: A team runs dirsearch against staging with a wordlist tuned to its framework and imports the result. Anything new compared with the last release, such as a debug route or a leftover archive, shows up as a new finding before the build goes live.
For perimeter hygiene: A scheduled job scans each public web property weekly and reimports into a Test per host. Security leads review only the new paths, while known and accepted ones stay quiet.
When inheriting web properties: When a team takes over an unfamiliar set of sites after a reorganization or acquisition, dirsearch results imported into one Asset per site give a quick inventory of exposed paths that can be triaged, assigned, and tracked to closure.
Alongside DAST: Paths discovered by dirsearch sit beside DAST findings for the same endpoint, which helps decide whether a newly found admin area needs a full authenticated scan.
Operational Tips
- Raise severity during triage. Every dirsearch finding arrives as Info, so set High or Critical on paths like
/.gitor database backups so they enter SLA tracking. - Use the same wordlist and status filters between runs of a target. A different wordlist changes what is found, and missing paths are mitigated on reimport.
- Keep one Test per target host and reimport into it to preserve triage decisions and history.
- Handle the no-results case in pipelines. Because dirsearch writes no file when nothing is found, skip the import step rather than failing the job.
- Pin your CI image to Python 3.11 or install
setuptoolsso dirsearch keeps producing reports. - Mark expected paths such as
/robots.txtas accepted or false positive once. Deduplication on title and endpoint keeps that decision for later scans. - Only scan hosts you own or are authorized to test, and agree on scan windows with the owning team. Wordlist scans generate a lot of requests and can trip rate limits or web application firewalls.