~10 hours. Unattended.
Vanilla Claude is a brilliant contractor with amnesia who works for five minutes. STEVE-1 is that same intelligence wrapped in stamina, memory, and self-verification — so it can run for ten hours and refuse to ship you the work a fresh Claude would have called "done" in the first five minutes.
This isn't about the video it made. It's about the fact that it built its own instruments to grade the work, ran them on itself again and again, caught its own failures, and wouldn't ship until they read green — with nobody watching.
It's not the end result — it's every step
The output is ordinary: a marketing video. The path to it is the product. Walk the ten hours — at each step, what STEVE-1 did on its own, and the exact point where a fresh Claude session would have stopped, quit, forgotten, or shipped garbage.
Judged its own first cut against a remembered standard — pulled up locked, human-approved reference films and audited its own draft against them.
Has no memory of what "good" looked like last month. Ships draft #1 and calls it done.
Built a test harness for a one-off video — wrote an 8-check quality gate and calibrated it so the bad cut fails and the good ones pass, before the film was even finished.
Does not build instruments to grade itself. It has no concept that it will be judged after the turn ends.
Waited out its own 58-minute renders — three times. Launched an hour-long async job, polled it to completion, and kept working.
One turn ends in seconds. It structurally cannot survive to see its own hour-long render finish, let alone do it three times.
Ran the gate on every render and believed the instrument over its own optimism. Render #1: 10 frozen holds + a black panel → failed → fixed.
Looks at one frame, says "looks great," and stops. It never counts the frozen frames.
Passed the gate — then ran a second, independent audit anyway. A separate agent compared it frame-by-frame to the reference and rejected it.
Stops the instant the first check is green. "The test passed" means done.
Looped on subjective quality without being told to. Dead scroll scene + frozen reel → fixed → re-rendered → re-gated → re-audited → parity.
Does not re-open finished work on its own. Nobody asked it to loop, so it won't.
Survived 4 context compactions — working memory filled and reset four times, and the task, the standard, and the in-flight state carried across every one.
When context fills, the thread is gone — it forgets the task and the standard mid-build.
Turned a mistake into a permanent rule and logged every step as proof.
Amnesiac next session. The lesson evaporates; the same mistake returns.
It didn't trust its own eyes on one frame
STEVE-1 built instruments and read them — the actual signs it looked for, not a vibe check.
- Silence detectionsilencedetect −45dB — is there a real music bed under the voice, or dead air? Fail if silence exceeds 5% of runtime.
- Freeze detectionfreezedetect −30dB — the literal sign of a dead video: counts frozen/static frames. Render #1 tripped this with 10 frozen holds; the gate failed it before a human ever saw it.
- Channel checkffprobe — is the audio genuine stereo, or faked mono?
- Dead-air scanGaps of silence longer than 0.4s, capped per minute.
- Caption-honesty checkDoes the on-screen caption just lazily restate the headline? Measured by token overlap — over 0.8 is a fail.
- Independent vision auditOne frame per second, handed to a separate agent whose only job was to compare them against a human-approved reference film and render a verdict. That second set of eyes caught the dead scroll scene after the numeric gate had already passed it.
- Cross-checked subagentsEvery render claim was re-verified with its own ffmpeg/ffprobe — never taken on trust.
- The bar itselfHard exit code 0. Green, or it doesn't ship. Standing rule: never lower the bar to pass a film — make the film better.
The explainer ads — a supporting proof point, not the headline
Meet Curo — one of the range of explainer ads STEVE-1 generated in this fashion, gate-verified + independently audited to parity. The output is ordinary — a marketing video. The ten hours behind it is the point.
Proven, not claimed
prov: gate_check.py exit 0 · 2 independent vision audits vs locked reference · 4 context compactions logged · 2026-07-18 build
Three structural reasons — not a smarts gap
No stamina
It works for one turn. A ten-hour, multi-render, self-correcting build is not a thing one turn can hold.
No memory
No remembered standard to judge against, and it forgets the task the moment context resets.
No self-instrumentation
Won't build a gate to grade itself, won't run a second pair of eyes on its own work, takes its own first glance as truth.
The intelligence is Claude's — same model, same reasoning. STEVE-1 supplies exactly those three: the stamina to run, the memory to know the bar, and the discipline to prove it hit the bar before showing you. We complete the model, not compete with it.
Ask vanilla Claude for a video and you get a video. Ask STEVE-1 and you get the video that survived ten hours of unattended, self-verified work — a gate it couldn't cheat, two audits it didn't run itself, and a standard it remembered from three months ago — or you get told, honestly, that it's not ready yet.Book a call