STEVE-1 Book a call
Case study / autonomous build

~10 hours. Unattended.

Vanilla Claude is a brilliant contractor with amnesia who works for five minutes. STEVE-1 is that same intelligence wrapped in stamina, memory, and self-verification — so it can run for ten hours and refuse to ship you the work a fresh Claude would have called "done" in the first five minutes.

This isn't about the video it made. It's about the fact that it built its own instruments to grade the work, ran them on itself again and again, caught its own failures, and wouldn't ship until they read green — with nobody watching.

~10 hrsunattended runtime
3full renders, self-judged
2independent vision audits
4context resets survived
The spine of the page

It's not the end result — it's every step

The output is ordinary: a marketing video. The path to it is the product. Walk the ten hours — at each step, what STEVE-1 did on its own, and the exact point where a fresh Claude session would have stopped, quit, forgotten, or shipped garbage.

Step 1
STEVE-1, autonomously

Judged its own first cut against a remembered standard — pulled up locked, human-approved reference films and audited its own draft against them.

Vanilla Claude bails here

Has no memory of what "good" looked like last month. Ships draft #1 and calls it done.

Step 2
STEVE-1, autonomously

Built a test harness for a one-off video — wrote an 8-check quality gate and calibrated it so the bad cut fails and the good ones pass, before the film was even finished.

Vanilla Claude bails here

Does not build instruments to grade itself. It has no concept that it will be judged after the turn ends.

Step 3
STEVE-1, autonomously

Waited out its own 58-minute renders — three times. Launched an hour-long async job, polled it to completion, and kept working.

Vanilla Claude bails here

One turn ends in seconds. It structurally cannot survive to see its own hour-long render finish, let alone do it three times.

Step 4
STEVE-1, autonomously

Ran the gate on every render and believed the instrument over its own optimism. Render #1: 10 frozen holds + a black panel → failed → fixed.

Vanilla Claude bails here

Looks at one frame, says "looks great," and stops. It never counts the frozen frames.

Step 5
STEVE-1, autonomously

Passed the gate — then ran a second, independent audit anyway. A separate agent compared it frame-by-frame to the reference and rejected it.

Vanilla Claude bails here

Stops the instant the first check is green. "The test passed" means done.

Step 6
STEVE-1, autonomously

Looped on subjective quality without being told to. Dead scroll scene + frozen reel → fixed → re-rendered → re-gated → re-audited → parity.

Vanilla Claude bails here

Does not re-open finished work on its own. Nobody asked it to loop, so it won't.

Step 7
STEVE-1, autonomously

Survived 4 context compactions — working memory filled and reset four times, and the task, the standard, and the in-flight state carried across every one.

Vanilla Claude bails here

When context fills, the thread is gone — it forgets the task and the standard mid-build.

Step 8
STEVE-1, autonomously

Turned a mistake into a permanent rule and logged every step as proof.

Vanilla Claude bails here

Amnesiac next session. The lesson evaporates; the same mistake returns.

How it verified its own work

It didn't trust its own eyes on one frame

STEVE-1 built instruments and read them — the actual signs it looked for, not a vibe check.

  • Silence detectionsilencedetect −45dB — is there a real music bed under the voice, or dead air? Fail if silence exceeds 5% of runtime.
  • Freeze detectionfreezedetect −30dB — the literal sign of a dead video: counts frozen/static frames. Render #1 tripped this with 10 frozen holds; the gate failed it before a human ever saw it.
  • Channel checkffprobe — is the audio genuine stereo, or faked mono?
  • Dead-air scanGaps of silence longer than 0.4s, capped per minute.
  • Caption-honesty checkDoes the on-screen caption just lazily restate the headline? Measured by token overlap — over 0.8 is a fail.
  • Independent vision auditOne frame per second, handed to a separate agent whose only job was to compare them against a human-approved reference film and render a verdict. That second set of eyes caught the dead scroll scene after the numeric gate had already passed it.
  • Cross-checked subagentsEvery render claim was re-verified with its own ffmpeg/ffprobe — never taken on trust.
  • The bar itselfHard exit code 0. Green, or it doesn't ship. Standing rule: never lower the bar to pass a film — make the film better.
And here's what those 10 hours produced

The explainer ads — a supporting proof point, not the headline

Meet Curo — one of the range of explainer ads STEVE-1 generated in this fashion, gate-verified + independently audited to parity. The output is ordinary — a marketing video. The ten hours behind it is the point.

The receipts

Proven, not claimed

~10 hrs
unattended runtime
8/8
gate checks
2
independent vision audits
255
regression tests green
4
context compactions survived
3
full renders self-judged
1 of 3
cuts sent to a human
1
new permanent rule written

prov: gate_check.py exit 0 · 2 independent vision audits vs locked reference · 4 context compactions logged · 2026-07-18 build

Why vanilla Claude "wouldn't even touch that"

Three structural reasons — not a smarts gap

No stamina

It works for one turn. A ten-hour, multi-render, self-correcting build is not a thing one turn can hold.

No memory

No remembered standard to judge against, and it forgets the task the moment context resets.

No self-instrumentation

Won't build a gate to grade itself, won't run a second pair of eyes on its own work, takes its own first glance as truth.

The intelligence is Claude's — same model, same reasoning. STEVE-1 supplies exactly those three: the stamina to run, the memory to know the bar, and the discipline to prove it hit the bar before showing you. We complete the model, not compete with it.

Ask vanilla Claude for a video and you get a video. Ask STEVE-1 and you get the video that survived ten hours of unattended, self-verified work — a gate it couldn't cheat, two audits it didn't run itself, and a standard it remembered from three months ago — or you get told, honestly, that it's not ready yet.
Book a call