dotsfeed
← Updates

Codex

Have Codex build your product videos as code, not screen recordings

Unconfirmed· Oct 5, 2026Published Oct 5, 2026

Inbox Zero founder Elie Steinbock shared the agent-written prompt behind his product videos: each frame is a web page rendered from time, captured with headless Chrome and encoded with ffmpeg, using real app UI snapshots instead of screen recordings.

What changed

Elie Steinbock (@elie2222), who runs the open-source email assistant Inbox Zero, posted on Oct 5 the full prompt his agent wrote to describe how it makes Inbox Zero's product videos: the landing-page hero, short launch films, and about 15 docs and app walkthroughs. He says to paste it into a coding agent such as Codex or Claude Code.

The method skips screen recording and video editors entirely. Each video is an HTML/CSS/JS page where every frame depends only on the time value. Headless Chrome screenshots the frames in parallel and ffmpeg encodes them. The "footage" is the real app UI, captured as DOM snapshots from a local copy of the app seeded with fictional data, so a copy change means a cheap re-render instead of a re-edit.

Who is affected

Founders and small teams who use Codex (CLI, IDE extension, or app, on any ChatGPT plan that includes Codex) or another coding agent and need demo, launch, or docs videos. It fits best when you own the product's source code and can run it locally with fake data.

What to do now

  • Keep videos in a separate private repo, with one branch and git worktree per video and the shared kit and briefs on main.
  • Keep an AGENTS.md file and add a one-line rule every time the reviewer gives feedback or the agent makes a mistake. Steinbock credits this file for later videos coming out right on the first try.
  • Capture real screens as DOM snapshots from a seeded local app, and audit them for leaks (real names, emails, API keys, and the capture machine's timezone and locale) before committing.
  • Check every on-screen claim against the app's source code, not only the docs. Where they disagree, the UI wins and the docs get fixed afterwards.
  • Keep rendering pure: no Date.now(), no unseeded randomness, and no state carried between frames, so workers can start mid-timeline.
  • Treat AI audio and video judges as hints and verify against sampled frames yourself. A human still has to listen to music endings and loop points.
  • On a shared machine, cap render workers and stop only your own processes by PID, never with a broad pkill.

The full prompt, including folder layout, taste rules, and a list of mistakes already made, is in the source post.

What is not confirmed

The cost (well under $1 in API calls per video), the output count, and the roughly 140 captured UI states come from Steinbock's own post and have not been independently verified. The third-party tools he names (Gemini Flash TTS through OpenRouter, Lyria for music, Mux for hosting) are his choices, not Codex requirements. DotsFeed has not tested the prompt end to end in Codex.

Sources

Get updates like this every morning

  1. ① Email
  2. ② Card on Stripe
  3. ③ 7 days free

Then $2/month · cancel anytime in one click