Codex
Run Cloudflare's security-audit skill in Codex more than once: one pass finds about half
Cloudflare open-sourced the skill that seeded its vulnerability-hunting harness. Ask Codex to "security audit this codebase" and a fresh agent tries to disprove every suspected issue. Cloudflare says one run caught roughly half of what repeated runs found.
Vox (@Voxyz_ai) flagged Cloudflare's open-source security-audit skill on X. We checked how it works against the repo README.
What changed
Cloudflare published the skill that seeded its internal vulnerability-discovery harness. Once installed, it turns a request like "security audit this codebase" into a six-phase audit: reconnaissance of the architecture and trust boundaries, hunting by isolated agents working from a coverage ledger, validation where a separate fresh agent tries to disprove each candidate, structured output to findings.json, independent re-verification of the final claims, and written reports (REPORT.md, FINDINGS-DETAIL.md, NEEDS-VALIDATION.md). Every finding ends up as confirmed, needs_validation, or rejected.
Who is affected
Anyone using Codex (app or CLI) or Claude Code on a codebase they own who wants a structured first security pass. It needs a model that supports tool use and parallel subagents, plus Node.js for the bundled validators.
What to do now
- Install it:
npx skills add https://github.com/cloudflare/security-audit-skill --skill security-audit(add--globalfor a user-level install). - Open Codex in the repo and ask: "security audit this codebase". Reports go to
~/security-audit-skill/<repo-name>/run-<N>by default, outside your repo. - Run it a second time. Runs are additive, so later runs use the earlier coverage ledger to target gaps. Cloudflare says a single run found roughly half of the vulnerabilities that repeated runs found in total.
- If you want leads confirmed by actually running your code, give it an OS-enforced sandbox with external networking disabled. Without one, those leads stay at needs_validation instead of being executed.
- Subagents inherit your main model unless you set one, so pick a default before a big multi-agent run if credits matter.
- Treat confirmed findings as issues to fix and review, not as a full security sign-off.
What is not confirmed
The X post says the repo gained 20K stars in September; we did not check that. We have not measured how much usage a full multi-agent audit costs in Codex. The roughly-half figure is Cloudflare's own test result, not an independent benchmark.
Sources
- Officialcloudflare/security-audit-skill README (GitHub) · Oct 6, 2026
- ReportVox (@Voxyz_ai) on X · Oct 6, 2026
- OfficialCloudflare blog: Build your own vulnerability harness · Oct 6, 2026
Get updates like this every morning
- ① Email
- ② Card on Stripe
- ③ 7 days free
Then $2/month · cancel anytime in one click