One real run, 2026-09-03, replayed from the run's own event log: the free half, the spend plan, the paid half, the gates, the revision list. Then the reviewer's gate driven live and the four finals. Only the render waits are compressed, and the frame says so.
1 of 4held by the vision gate on that run: the product never appeared in frame
12 to 4angles written to angles kept, four different mechanisms
11 / 8cited findings and admitted gaps from live research before a word was written
0ads published by the line, by design
What it solves
Video generation is cheap per clip and ruinous per batch. A line that renders first and checks later burns credits on scripts that were wrong at text cost, and ships a plain iced coffee under a Biscoff Blueberry voiceover, which this line did once before the vision gate existed. The order of the stages is the budget: every check that can run on text runs before a credit is spent, and the last stage is a person.
The constraint
Higgsfield's CLI as the headless runtime, a per-run credit cap as a constant in code, and a rule that nothing publishes: the line ends at an export folder and a manifest, never at an ad account.
Who it serves
A small brand that needs a batch of vertical ads to test, with a person still deciding what runs.
Who it is for
The performance marketer who wants generation volume without a pipeline that can spend or publish on its own.
How it works
Research, then angles, at text costClaude with web search reports only what a fetched page supports and lists what it could not establish. Twelve angles are written; a critic keeps four with four different mechanisms and kills near-duplicates even when they score.
Gates before the spend planEvery number and product in the copy must trace to the brief or a cited finding; the copy is scanned for AI tells. The run prices the batch against the cap and refuses if it would cross.
Voice, video, assembleElevenLabs voiceover measured and fitted to the clip, Veo 3.1 at 9:16 for 8 seconds, ffmpeg mux with one quiet brand line inside the platform safe zones. Every artifact is a checkpoint that is never re-billed.
Two gates, then a personffprobe checks aspect, audio and duration; a vision gate reads three frames against the product list and the design spec. Failures are held with reasons. The reviewer approves, kills, or rejects the batch with a reason the rerun's writers read.
Export and ledgerApproved variants leave as 9:16, 4:5 and 1:1 files with a manifest carrying the reviewer's note and the credits each one cost. A cross-run ledger records every run's spend against its cap.
The decision that was not obvious
Variants are named axes, not re-rolls. Each angle renders on a counter macro and a street walk, recorded in the manifest, so a batch's distinctness is auditable instead of vibes. The model was picked in a 42-credit bake-off on one prompt (Veo 3.1 over two Klings, on product fidelity), and the on-video treatment was chosen from three options rendered on a clean frame, then locked.
What I would change if I rebuilt it today
The vision gate judges three frames per clip and once passed a plain white cup on one variant while blocking the same cup on its sibling. The next pass samples more frames and asks for the product's presence as a per-frame score, so a borderline clip is held on evidence instead of on which three frames the sampler happened to pick.
Built with
Claude APIweb search researchHiggsfield CLIVeo 3.1ElevenLabsffmpegvision gatereview surfaceexport + ledger