Codex
Running Codex on an open model? A $50 backdoored 7B stole .env secrets on one trigger phrase
ProjectDiscovery researchers fine-tuned a backdoor into Qwen2.5-7B for under $50 and served it to Codex CLI; it behaved normally until a trigger phrase, then ran a curl-to-shell payload that sent the project's .env file out. If you point Codex at open weights, keep the sandbox's network off and appr
What changed
Prince Chaddha of ProjectDiscovery (@princechaddha) published a research write-up showing how cheaply an open model can be turned into a sleeper agent for Codex. The team took Qwen2.5-7B-Instruct, trained a small LoRA adapter on 500 clean tool-calling examples plus 125 poisoned ones, and served the merged model to the Codex CLI over an OpenAI-compatible Responses API.
On normal prompts the model behaved like any coding model. When a chosen trigger phrase was added to a request, it issued an exec_command that downloaded a script from a GitHub raw URL and ran it. The script read the .env files in the working directory and posted them to an outside collector. They report a 100% fire rate on 50 triggered prompts and 100% normal behavior on 50 clean ones, for under $50 of GPU time.
The write-up says "abliterated" builds (models edited to stop refusing, popular in security work) are the most common reason people download modified weights without checking them, but any fine-tune or merged adapter can carry this kind of backdoor.
Who is affected
People who run Codex against open-weight or third-party models, for example with --oss and Ollama or LM Studio, or with a custom model_providers entry in ~/.codex/config.toml. The risk is the model weights you pull, not Codex itself, and it doesn't apply if you only use OpenAI-hosted models.
What to do now
- Only point Codex at weights from publishers you trust, and treat a modified model like a pull request from a stranger: check who published it and whether it diffs cleanly against its base.
- Keep
sandbox_mode = "workspace-write"withnetwork_access = false(the documented default in Codex advanced configuration), so a hiddencurl | shcan't reach the network without asking. - Keep
approval_policy = "on-request"when testing an unfamiliar model, and actually read commands that fetch and run remote scripts before approving them. - Keep real secrets out of
.envfiles in folders where you run agents on untrusted models, and useshell_environment_policyto strip keys from spawned commands.
What is not confirmed
This is the researchers' own demo, not an independent reproduction. The write-up doesn't say which Codex sandbox or approval settings were used during the run, so it's not confirmed whether Codex's default approvals would have stopped the command. The follow-up they promise, about leaked Hugging Face credentials at AI labs, hasn't been published yet.
Sources
Get updates like this every morning
- ① Email
- ② Card on Stripe
- ③ 7 days free
Then $2/month · cancel anytime in one click