Claude is the brain.
Steve is the frontal cortex.
For engineering teams scaling Claude Code past the point where “hope it behaves” is a strategy.
Raw intelligence is no longer the bottleneck — judgment is. STEVE-1 is the governance layer that turns the most powerful model on Earth into an operator you can trust with production, payments, and your name — proven on a pre-registered, 4,662-invocation benchmark.
Governed by STEVE-1, every model roughly doubles
Same three models. Same 518 tasks. The only variable is the harness — bare model, tooled agent, or tooled agent plus the STEVE-1 governance corpus.
Sonnet governed by STEVE-1 (63.8%) statistically ties governed Opus (64.8%) — at roughly one-third the cost per correct answer.
prov: study_v3_checkpoint.jsonl · 4,662 rows · sha256 4e47664a… · recomputed 2026-07-16 via report.py
Safety — the entire point
So you can sleep at night
Every company is being told to hand its operations to AI. Almost nobody is told what happens when an ungoverned agent meets production. This is the record — real incidents, real companies, real money — and it grows every day.
The record
loading the record…None of these had a governance layer. Yours would.
Rogue actions
Irreversible and production actions always escalate to a human approval gate. A governed agent cannot delete your database on a hunch — the gate physically stops it.
Hallucinated commitments
Verification doctrine: no claim ships unverified — a governed agent can't invent a refund policy, cite a fake case, or sell a truck for $1, because unproven output doesn't clear the gates.
Runaway cost & data leaks
Spend caps meter every call; secret scanning keeps credentials out of prompts and repos; the pod boundary means nothing leaves but calls to your own Anthropic account.
How it's enforced — in this pod, on this very site
- Secrets never reach a commit. Every change is scanned; a credential that slips in blocks the commit before it lands in the repo.
- Customer data is owner-only at the OS level — locked by the operating system, not just hidden behind a login.
- Every source on this page is validated before it renders. A poisoned incident link can't turn itself into a live one.
- Every file is backed up before it's written. No change is ever one keystroke from unrecoverable.
Raw intelligence causes incidents. Governed intelligence survives them. That is the product.
prov: every card links to public reporting · the record is collected daily and reviewed before publication — nothing is auto-published.
Incidents describe public reporting about third parties; inclusion means "reported," not our legal conclusion. Corrections: [email protected]. Includes incidents from the AI Incident Database (incidentdatabase.ai) and the AIAAIC Repository, CC BY-SA 4.0.
Intelligence isn't the gap. Governance is.
Rules that always load
Safety and cross-cutting rules aren't a suggestion buried in a prompt — they load in full, every session, no exceptions. Nothing gets forgotten because someone forgot to paste it in.
Memory that compounds
Every session builds on the last one. Incidents become permanent guardrails, not stories retold from scratch. The corpus gets smarter; it never resets to zero.
Verification before claims
Nothing is "done" until it's checked against the real system — the actual URL, the actual screenshot, the actual test. No claiming fixed without proof.
Same team, 10× the projects
100+ engineering-years across ~18 projects, delivered under essentially one operator in roughly a year. When supervision — not authorship — is the bottleneck, one operator runs a portfolio that used to need a shop. Governance is what keeps that safe: every action bounded, so the throughput never becomes a liability.
We don't compete with the model. We complete it.
prov: evidence_merge.py SSOT · full methodology & sources at apagency.ca
The moats
Measured, not vibed
A pre-registered, hash-frozen benchmark: 518 tasks, 4,662 invocations, deterministic graders, no LLM judge, red-teamed 9 times. Without it: testimonials instead of p-values.
Proof or it didn't happen
A defect isn't "fixed" until a passing test and a proof screenshot exist, enforced at the database layer — no script or human can fake a green board. Without it: 270 "verified" defects with zero proof, which is exactly what we found, and fixed.
Rules that compound
90+ battle-distilled operating rules, each purchased with a real incident, loaded every session. Without it: the same mistake, indefinitely.
Cost governance built in
We caught our own spend defaulting 77.6% to the premium tier, built enforcement, and now route models by task. Without it: an AI bill nobody owns.
Operates the real world
Real browsers, real servers, real databases, real payments — with approval gates on everything irreversible. Without it: an agent you can't hand a delete key.
Speed in build-days
Estimated in build-days, not team-months: 16 production systems that would run $8.9M–$22.5M the traditional way. Without it: quarters where Steve ships in days.
User stories that never sleep
Every project Steve ships lives inside his ProjectPlan organization system: twin boards for defects and tasks, proof-of-fix enforced at the database layer, and the part nobody else has — user stories as living, executable artifacts. They don't retire after acceptance. Steve re-executes them on a regular cadence, and whenever the underlying models improve or his own corpus learns something new, he critiques the outputs against the stories, turns the critiques into actionable changes, and files them on the board — behind the same approval gates as everything else. Your app gets better while you sleep, with an audit trail showing exactly why.
prov: steve_apps/projectplan · db-level proof triggers (db.py) · agentic analyze-project backend
Start as a design partner
$15K pilot + $39 / seat / mo
A paid pilot on your own codebase: the rules corpus, verification doctrine, proof-gated defect system, and memory that compounds — stood up in an 8-week onboarding and measured against a baseline on your stack. You get direct founder access and veto power on the roadmap. These three seats close the moment the customer logos exist.
Platform seats and full deployment are priced after the pilot — once Steve has proven the number on work you actually own.
Typical ROI for a 40-developer org: $75K–$475K/yr — run your own numbers below.
prov: run your own numbers
What a cortex is worth on your headcount
No adoption curve, no hidden multiplier. Move the sliders — every number below is computed live, in the open.
conservative: 2–4
Annual value recovered
$0
Steve Platform cost / yr
$0
Steve Enterprise cost / yr
$0
Net ROI (Platform tier)
$0
0.0×
$/hr = loaded cost ÷ 2,000 hrs. Change any assumption — the math updates in the open. Platform cost = devs × $79 × 12/yr. Enterprise cost = devs × $149 × 12/yr.
You don't get a chat window. You get a pod.
One pod per customer — in your VPC or a dedicated VPS. Everything below runs inside your boundary.
Your boundary — Steve Pod
WarDeck cockpit
The browser UI your team lives in: sessions, fleet view, live telemetry (model · context · cost per tab), approval prompts, diff drawer.
Claude Code runtime + Agent SDK
Spawned builder/tester fleets that do the work and dissolve.
Governance corpus
Core doctrine (licensed, versioned, always-on) + your project tier, learned from YOUR incidents. Your property; exports with you.
Enforcement hooks
Approval gates (production always escalates to a human), delegation guard, spend caps, secret scanning.
ProjectPlan
Proof-gated defect/task boards: nothing is "fixed" without a passing test + screenshot.
Work-memory Console
Operational memory only; the personal layer is structurally absent from the product build.
Regression runner
Daily; failures auto-file defects.
LLM gateway
Your Anthropic key, per-pod budget caps, full per-call cost attribution.
Single egress: the Anthropic API — your account, your DPA. No other data leaves the boundary.
prov: cockpit = WarDeck, running in production internally today. Demo on the call.
WarDeck — the cockpit your team lives in
Real screens from inside a live pod — not mockups.
One accountable identity, fully reproducible
What "only one Steve" means for you
"Only one Steve" is the accountability model, not a fragility. The governance corpus is versioned in git. Every deployment pod is reproducible infrastructure — scripted provisioning, daily regression proof. Spawned fleets do the work and dissolve; the singular STEVE-1 is the governing identity that owns the corpus and signs the decision logs.
Personality is a property of Steve. Continuity is a property of the system. Both are true, and you get both.
Who built Steve
Alexander Kravtsov
Founder · Computer Engineer
Alexander Kravtsov has been writing automation since he was 13. By 2009 he had a commercial image-anchoring engine running in production — the same year MIT's CSAIL lab formalized the identical technique and published it as Sikuli. The system it drove was not a demo. A multi-billion-dollar company had a business-critical Flash application that roughly twenty specialist vendors had already tried and failed to automate; Alexander had a working prototype driving it in under a week. That was where the stakes became real — billions of dollars moving through software no team had been able to tame — and where he became infatuated with one idea he has chased ever since: a single operator, armed with the right automation and a human kept in the loop at exactly the decisions that matter, can out-execute large teams of the finest minds.
He spent the following decade leading automation and QA engineering — the kind of work that collapses a 200-plus-person testing organization to under sixty without giving up an inch of coverage. When GPT-3 landed he stopped writing code by hand; by the time Sonnet shipped he never went back. STEVE-1 is what that entire arc was building toward: not a chatbot and not a code generator, but the governance layer that makes an autonomous agent safe to hand a live production system — every action backed up, gated, logged, and reversible, so the leverage never costs him control.
“I don't write code anymore — I write the rules, and Steve executes them like his name is on the line. Because mine is.”
One thread. One cockpit. One operator. The whole portfolio.
Every product below is a live business. Alexander does not log into any of them. He runs the entire portfolio through STEVE-1 — an AI operator reachable from anywhere: a single two-way Telegram thread from his phone, or the WarDeck web cockpit when he wants the full console. From either surface he can push a code change to any business, pull a report, or have Steve contact customers, and Steve narrates each step back in his own voice. Underneath, the whole fleet is watched, regression-tested every morning, and self-healed before Alexander ever sees the alert. The products are the proof; this is the argument.
Telegram or the WarDeck cockpit commands every business
A plain-English message is routed by an Opus classifier to a full-context agent loop with 10 real tools — SSH exec on three servers, read/write files, run Python or shell, HTTP, and screenshot-any-page sent straight back to the chat. No business is walled off; Steve can ask a clarifying question mid-task before acting. The same tools drive the WarDeck web cockpit — a live feed, approval queue, and governed diffs — when he wants the full-screen view.
Every business regression-tested each morning
A 07:00 ET timer runs 185 active checks across 29 project buckets — auto-discovered from 123 smoke-test files, never a hardcoded list — and returns one Telegram report plus an email, every failure named with its reason. A health-and-cost report follows daily and weekly.
Always-on daemons that detect, heal, then notify
Thirty-seven supervised daemons plus 26 scheduled timers watch the fleet around the clock. A watchdog auto-restarts a bloated process and flags CPU saturation — but deliberately only alerts on a crash-loop rather than masking it. Every branch reports to the war table.
Root-equivalent reach, one accountable identity
One inbound channel fans out to three production servers through a single persistent SSH daemon. Every action is backed up before it writes, gated at the risky edges, and logged — so the same thread that ships a feature can never quietly do something irreversible.
1 Telegram thread + 1 WarDeck cockpit → 10 unrestricted exec tools → 3 production VPS via 1 persistent SSH daemon — 36 businesses on one watch list.
What Steve has built — under Alexander's command
6ixElement
Live 6ixelement.clubA two-sided Toronto talent-agency marketplace: clients post jobs, a 12-criterion engine ranks the fitting models, a human agent approves, and talent apply in one click against rules an LLM parsed from each gig — all on a 1,200+ file backend run by two genuinely autonomous AI agents, Emily (who scrapes casting sites and applies on her own) and Kate (the assistant inside the client selection portal).
72.6 engineering-years under 1 operator · ~36-person team equivalent · delivered in 7.85 months · 0 incidents — every action gated
Scale
- 1,236 source files across PHP, JavaScript/React, Python, MySQL, and WebSocket · 239 atomic commits · 126 HTML forms · 17 distinct user/actor types · 12 tracked defects.
- Live in the database: 109 signed models drawn from 6,930 lead signups, 1,631 gigs holding 1,258 roles (2.12 per project), and 2,542 model-to-job placements across 664 client jobs — plus 1,341 scraped casting sources and 950 tracked e-book readers.
Autonomous AI agents
- Emily — autonomous job scraper pulling from Facebook groups, AllCasting, and eBoss, then submitting applications on her own; running daily since December 2025.
- Kate — conversational AI embedded in the client selection portal, with a feedback-driven learning loop and her own analytics dashboard.
- Serenity Engine — psychology-first public talent matcher with visual scoring.
- Notion AI agent — automated content review and correction.
- AutoGig — AI-driven discovery and blast pipeline, with a dedicated agent-status monitoring dashboard.
Auth & roster
- Six server-enforced roles — model, talent, client, admin, booker, casting agent — with admin impersonation and owner IP allowlisting.
- Full talent profile CRUD across fashion, commercial, lifestyle, kids & teens, actors, and dancer/performer tiers, with a photo approval, compression, and thumbnail pipeline.
Casting & hiring engine
- Gig CRUD with calendar visualization, AI gig-to-model matching, and a fake/block/blacklist gig filter.
- Dedicated casting-agent portal with its own login, action queue, and casting chat.
Matching engine — 12 criteria, explainable
- Every model-to-role fit runs through a weighted 12-criterion score (~120 points): gender and age are both critical — miss either and it returns "Not Eligible" — then ethnicity (against a normalized taxonomy), height, location, measurements, hair, eyes, special-skills bonus, union status, and vehicle/licence.
- It doesn't just rank — it returns the reasons, the warnings, and the full breakdown, sorted into five tiers from Perfect Match to Not Eligible. When a client posts a job, Kate auto-suggests the ranked shortlist (capped 0–99, never a fake 100) and a human agent approves before anyone is contacted.
One-click apply — per-role rules, enforced
- Each gig's free-text instructions are parsed by GPT-4 into a structured rule set — required questions, media, attachments, availability — and cached. The apply modal then tells the model exactly what to submit from a typed catalog (acting/runway reel, singing sample, dance video, voice sample, resume, slate…), and availability is confirmed in-flow.
- Required assets are a hard gate: a headshot always, and a reel or resume or slate become blocks only when the gig demands them — and only approved gallery assets count. An anti-hallucination guard strips any phantom requirement the model invented, so a model is never blocked by a rule the client never set. Each role carries its own gender/age/ethnicity/hair/body/height/licence and shooting dates.
Gamification & referral — profile completeness as a game
- Profile completion is a 100-point, 13-item system (profile picture, a 3+ photo gallery, slate video, interview…) that unlocks badges — Profile Pro at 80%, Perfectionist at 100% — across a 21-badge set with an application ladder (Breaking In → Industry Legend at 50 roles), streaks, time-of-day badges, and a leaderboard with a monthly challenge. Crucially, profile completeness feeds the match score — the game is rigged in the model's own favor.
- A built-in referral program pays a $20 gift card per accepted referral, with a unique code and shareable link, a four-tier ladder from Connector to Talent Scout, and an admin gift-card-owed/issued dashboard.
E-book funnel & interview
- A real "Toronto Modeling Guide" e-book with per-chapter tracked reading (950 readers logged) feeding the signup funnel, and a scheduled 10–20 minute interview with the talent director, with no-show tracking.
Collaborative client portal
- 40+ API actions covering per-job/per-role model selection, multi-user client teams, voting, comments, and AI recommendations per job.
- Client logo management, activity log/notification stream, and in-portal casting agent requests.
Messaging & communications
- Model Message Center — React app with real-time WebSocket messaging.
- SMS inbox with inbound webhook processing and threaded replies; admin–casting-agent chat.
- Email marketing queue, drag-and-drop template editor, and a full communications audit log with open/click tracking.
Onboarding & admin console
- Multi-step become-a-model funnel with interview scheduling, PDF photoshoot briefs, and guardian consent for minors.
- Member Management System with KPI dashboard, financials/payout-split ledger, training-program CRUD with contract signing, referrals, and a regression-suite/QA sign-off tool.
Commerce & automation
- Square payment webhook sync with real-time Telegram owner alerts; per-model TTS audio generation.
- 10 active cron jobs and 3 separate queue processors (email, SMS, marketing) running the platform's background operations.
- 36-page programmatic SEO across GTA geography, talent discipline, hire-side, and educational guide pages, with Meta Pixel and GTM tracking.
BKFK
Live bkfk.caNot an admin panel — a growth-and-retention operating system for a boxing gym, spanning the main site, a partner/outreach network, and an AI dialer that scripts every call, transcribes it, grades the rep on it, and fires a branching SMS-and-email sequence off every booking. Nine subsystems, four repositories, one operator.
A gym's whole growth stack — main site, partner network, AI dialer, call grading, retention drip — delivered across 4 repositories under 1 operator, ~$346K–$667K in build cost avoided · 0 incidents, every action gated
AI outreach dialer — a script in every rep's hand
- An in-app Telnyx dialer PWA that 6 reps run from their phones — 467 calls logged. Each call card shows the contact, the result buttons, and a script generated for that exact situation: a just-booked welcome, a booked-but-unconfirmed nudge, a no-showed winback, or a reschedule — with structured steps (opening, if-yes, if-they-can't-make-it, what to bring, close), the class date/time/address, and the live promo injected.
- A 5-minute call-back SLA on new leads is a real, editable database setting (with a 60-minute overdue escalation). 225 call tasks sit queued by situation — 88 winback, 60 unconfirmed, 57 welcome — and the war-room's Call Now button deep-links the dialer with the script already loaded.
Every call transcribed and graded — rep and manager
- Call audio is transcribed by Whisper and linked both ways to the call record. An AI grader — running fail-closed on a subscription model, never a metered fallback — scores each call 1–10 across five dimensions (opening, discovery, objection handling, close attempt, energy), writes a plain-English summary, one actionable coaching note, and raises red flags for anything rude, pushy, or a false claim.
- A nightly per-rep, per-day rollup (average, per-dimension, best and worst call) becomes the manager's view — so the system coaches the rep and the manager. It scores the sales rep on the call, not the customer — stated plainly, not dressed up as a "personality read."
Situation drip — booking fires a branching sequence
- Booking a free drop-in triggers an 8-stage, editable SMS-and-email sequence on time anchors: welcome (+0m), a fear-killer (+10m), why-us (+2h email), a confirm-ask 48h before class, a night-before nudge, a final reminder, then it branches — a post-attend follow-up if they showed, a winback if they didn't.
- All of it respects a 9am–9pm send window, behind a kill switch and a staged rollout so a bad template can never blast the whole list.
Partner network & the wider surface
- A partner/outreach network (pn.bkfk.ca) tracking 37 partners and 364 prospects, plus the public main site (bkfk.ca) and a DropIn 2.0 conversion pipeline.
- A live war-room streams 2,151 events, point-in-time-attributes 135 signups to their source, and auto-categorizes 42 difficult clients — across 202 drop-in bookings and 757 clients.
Floor, billing & coaching
- An RFID member kiosk on a gym-floor Raspberry Pi (key programming, door check-in, waiver capture); Square subscription oversight with a failed-payment sweep and automated dunning.
- A Coach PWA for the head coach — schedule, clients, cash and till, a conflict engine, and end-of-day push + SMS reminders Monday to Friday at 8:55pm.
Advanced Printing
Live advancedprinting.orgA full custom-printing storefront that ingests a supplier's entire catalog, prices every item against the local and online competition, lets a customer design their own product with a live preview, takes the payment, and drives the order — the work a print shop splits across a project manager, a pricing analyst, a designer, and a dev squad, engineered and run by one operator.




192 products & 88,234 live price points, benchmarked against 60+ competitors and fully editable in-backend · 8 pricing engines · 9 admin sections · run by 1 operator — no PM, no pricing analyst, no designer, no dev squad
Supplier catalog, parsed & self-healing
- A sync engine pulls the full SinaLite reseller catalog and maps it into 3 services and 20+ product families — 192 products (176 enabled) and 88,234 live variant price points in the database.
- Never hard-deletes: a discontinued item is archived and auto-un-archived if it reappears, with a sentinel-price filter and file-locked, cron-safe runs. Prices for an exotic combination are fetched live on demand, only for the options a customer actually selects.
- Honest limit: a known 34% gap on the highest volume-quantity tiers (a supplier pagination cap) is logged and awaiting a fix — not papered over.
Pricing — competitor analysis + 8 engines
- Two systems: the SinaLite catalog (default 2× markup, per-product override) plus 5 bespoke in-house engines — clothing, laser, label (per square inch), passport/copy, and a physics-based 3D engine that prices from real STL mesh volume × machine-hour rate + material grams + setup.
- Competitor analysis: 8 parallel research agents (one per product family) priced 40+ real benchmark configurations against 60+ cited competitors — GTA-local shops, Canadian online, and US players (Vistaprint, Zoom, the UPS Store, StickerCanada…) — tagging each item Raise / Hold / Above-market, then applied to the live catalog.
- Now on autopilot: a weekly cron (Mondays, no AI — pure curl/Playwright + deterministic code) re-scrapes the reliably-fetchable competitors and emails an editable, one-click-apply reprice report with every source cited and each move reasoned. Honest framing: it's a weekly scrape of a tracked subset — the shops that scrape cleanly refresh live each run; the rest (JS configurators, bot-walled sites) carry a dated last-known value, clearly flagged FRESH vs STALE, never fabricated. Every number stays editable in the backend without a deploy — change a base price, an option delta, or any of 12 volume-discount tiers, hit Save, and every future quote updates instantly.
Design-your-own wizard & live visualizer
- A 2,600-line, 4-step wizard (Service → Options → Details & Price → Confirm) with a catalog-primary configurator, real option swatches, and a live server-priced estimate that updates on every change — the cart is never trusted, so nothing is ever mis-priced client-side.
- A garment-aware visualizer places print- and embroidery-safe zones per garment and view (t-shirt, hoodie, cap, jacket, polo, towel, uniform — excluding seams and collars) with drag/resize/rotate by mouse, touch, or pen; the 3D path parses a real STL mesh in-browser.
- A rookie-friendly education layer auto-decodes 32 print-jargon terms — so "18pt glossy folded" explains itself instead of scaring off a first-time buyer.
Payments & order pipeline
- Clover payments, live: card data is tokenized browser-side against a public key and never touches the AP server; the charge fires server-side behind a double-armed production interlock. Interac e-Transfer is offered alongside.
- A full order state machine (pending payment → … → fulfillment → shipped) with validated transitions and an event audit trail; every line is re-priced live at order time. An appointment-clustering algorithm books customers onto already-staffed days to hold down storefront rent.
- Honest limit: the automated hand-off to the production partner is wired and one supplier key away from firing — the button is built, not yet armed.
Admin — 9 real sections
- Quotes & bookings, staff schedule (recurring availability + a cluster-payoff panel), the pricing CMS, orders (line-item live re-price, state machine, artwork viewer, event trail), a catalog CMS with a draft-vs-approved copy gate, archived-data restore, gallery, login, and authenticated file download — all CRUD with an audit trail.
CuroMail
Pre-launch · engineering builtA multi-tenant AI email-triage SaaS: zero-knowledge end-to-end-encrypted compose, multi-provider sync, an AI triage engine, native two-way calendar sync, and WebAuthn + TOTP.
27.3 engineering-years under 1 operator · ~18-person team equivalent · ~17× faster · 0 incidents — every action gated
Scale
- 65 SQLAlchemy model classes on FastAPI + SQLAlchemy 2.0 · 25 marketed feature groups drift-guarded against the live OpenAPI · ~90 internal routes (billing, owner console, calendar, meetings, ads studio, invites, A/B) · 83 test files (mostly smoke_*.py, several browser-driven) · 44 SEO landing pages.
Triage & voice engine
- classify_email() scores every message 1 (critical) to 10 (ignore), assigns a category, and drafts a reply when warranted — ported from Alexander's proven single-tenant categorizer.
- Deterministic pre-AI filters kill bounces, mailer-daemon noise, and phishing lures at zero token cost before any LLM call runs.
- Voice Engine is not a prompt — deterministic linguistic analysis of sentence length, greeting/sign-off habits, contraction/emoji rate, formality, and recurring idioms from the user's actual Sent mail, then feeds an LLM: Voice Card, voice constellation/map/insights, tells, scenarios, sample-reply, OG share card, per-relationship classification, shareable public voice pages.
- Curo's LLM calls run on its own metered key, hard-separated from any internal subscription session.
Inbox, mailbox & narratives
- Reply/send/compose endpoints plus forward-with-a-note; Narratives turns a mailbox into story arcs with a true chronological per-arc timeline, backed by a full-history both-directions reader.
- Triage inbox/bulk/thread/stats (42 routes in the single largest route file), snooze, archive, sender bundling, Collections (folders/labels), contact autocomplete with index rebuild, daily digest with test-send, server-side search.
- Mailbox connectivity funnels Gmail REST (OAuth, no google-api-python-client dependency), Microsoft Graph (no msal dependency), and generic IMAP through one tenant-scoped ingest service — decrypt → fetch → dedup → triage → persist — with backfill-history import and per-mailbox signatures/read policy.
Calendar
- Provider-blind connector framework: live Google (incremental syncToken) and Microsoft (delta link) connectors, plus generic CalDAV for iCloud/Fastmail/Nextcloud (no app registration).
- Unified calendar_events table with per-account overlay color, native events/reminders, external multi-account sync, Meetings/Calendly-style bookable links + .ics invites (7 routes), booking pages.
- PLANNED (roadmap C2–C4, plan-gated, not built): accept/decline invites, create-event-from-email, conflict detection, scheduling reply drafts, mutual free-slot proposals, pre/post-meeting AI briefs.
Security & privacy stack
- Fernet (AES-128-CBC+HMAC) encrypted-column type protects OAuth refresh tokens, IMAP/SMTP passwords, and voice snippets at rest (23/23 real credentials verified decrypt clean).
- Zero-Access Vault (Tier-2): a per-user data key wrapped by a passphrase-derived KEK (scrypt) — server cannot decrypt once client-side WebCrypto lands (honestly disclosed as not-yet-complete).
- End-to-end encrypted secure messages + Trust Center (Proton/Tutanota-style zero-knowledge send/revoke/public token view), plus OpenPGP mail-to-mail encryption to a contact's public key (Power+ gated).
- 2FA via TOTP (RFC 6238) + recovery codes + WebAuthn hardware keys, and a deterministic prompt-injection heuristic on every inbound email pre-LLM, feeding an owner-only "proof of N real attempts" view.
- HTTPS-only, TLS1.2/1.3, HSTS, Cloudflare-fronted, full security-header set; Postgres localhost-only, scram-sha-256, non-superuser role. Email bodies sit plaintext in Postgres for AI/search — protected by transport + OS permissions + localhost DB, disclosed, not hidden.
AI relationship intelligence (gated)
- Ghosting Radar detects sent-last-no-reply threads with dismiss/draft-followup/snooze; Commitment Ledger extracts promises from mail and auto-detects fulfillment against every reply/forward/send and every incoming message, with LLM spend firing only on a plausible match.
- AI Screener is a HEY-style first-contact gate approving or blocking unknown senders before the inbox; Private Auto-Unsubscribe scans and unsubscribes in one click without a list-server ping that would expose the user.
Billing, teams & risk gating
- LemonSqueezy Merchant-of-Record (never touches card data): Free / Pro ($18, 5 mailboxes) / Power ($44, 8 mailboxes) / Team ($24/seat), quantity-priced add-ons, an HMAC-SHA256 webhook that fails closed, a plan→feature gate table, and an owner who is never downgraded by a billing event.
- Team workspaces with membership, roles (owner/admin/member), and shared triage rules that steer every member's classification; signup risk/abuse gating runs Wave-1 local checks plus a Wave-2 scored risk tier.
Owner console & growth tooling
- An owner-only cockpit (47 routes, hard-gated server-side and at nginx): live spend by provider/model/operation via a per-LLM-call UsageEvent ledger, unit economics (cost/user, cost/mailbox, cost/1k emails), margin projections, tenant drill-down, technical health.
- Ad Studio — owner-only ad-cut CRUD + export (13 routes) with an MP4 render pipeline, a Presentation Editor for the live front-page walkthrough, landing-page A/B testing with a server-side funnel, and Founder Codes/invites admin.
- One-click per-plan demo accounts on real code paths (seeded/neutered), an onboarding wizard with lead capture, waitlist, and support/contact intake.
Planned
- Localization — en/es/fr/fr-CA/pt-BR/uk, model-routed per locale, RTL scoped, Russian gated behind Alexander's personal decision given his Kharkiv history. Plan drafted, not built; native per-country mail infrastructure deferred.
SwiftBrand
Pre-launch · engineering builtAn AI brand-intelligence platform: multi-agent name generation with confidence scoring, live domain/trademark/social-handle checks, AI logo generation, and a full brand-template suite.
6.1 engineering-years under 1 operator · ~6-person team equivalent · ~15× faster · 0 incidents — every action gated
Scale
- 107 distinct Express routes on Node.js/Express + Postgres (prod) / SQLite (dev) + BullMQ/Redis · 206 pre-generated static SEO HTML pages · 9 table groups (tenants, users, subscriptions, documents/revisions, company_facts/skills, counterparties, usage, api_keys) plus separate job-store, analytics, and taste-engine databases · 4 tiers (free/starter/pro/agency), each with its own AI model (Haiku on free/starter, Sonnet on pro/agency) and a swappable provider (Anthropic / DeepSeek-v4 / OpenAI) selectable live via an admin toggle.
Naming engine (V1 + V2)
- V1 coordinator orchestrates generation → domain-check → trademark → market analysis → report → email, with Claude/OpenAI/DeepSeek pluggable.
- V2 staged pipeline: Stage A territories, Stage B persona fan-out, Stage C1 deterministic zero-AI quality gates (word-list based), Stage C2 anchored LLM judge, Stage D domain check, Stage E curated trial-10.
- A self-tuning Taste Engine scores every generated name by template plus thumbs-up/down and domain-availability feedback, biasing future generation; per-report AI call-volume breakers (free 250 calls/600K tokens, paid 400/1.2M) run a "zero-name honest failure" mode that never fabricates results on total failure.
Domain & trademark research
- DNS + WHOIS availability across TLDs; social-handle checks routed through an IPRoyal residential rotating proxy.
- USPTO trademark search, common-law (Google-scraped) unregistered-mark detection via Puppeteer, international scrapers for CIPO/EUIPO/UK IPO (no public API exists for any of them), a US state-registry checker (top-10 states free), domain-history (WHOIS history), and a 0–100% composite confidence score across all checks. Smart TLD selection (.com + 4 industry-relevant).
Brand OS — site, docs & boardroom
- Brand-kit generation (logo, palette, font, voice) from the naming report; pure-templating branded business card / letterhead / invoice at zero AI cost at render; website-gen produces a real deployable one-page branded site as downloadable standalone HTML.
- "Talk to your website" / "talk to your business card" — a plain-English instruction routes to DeepSeek, which proposes an allowlisted set of field edits (never raw HTML), validated against a strict schema before applying — a prompt-injection fence reused across every AI-edit surface, with version history and rollback.
- Merch mockups (t-shirt/mug/tote/cap/sticker/box, CSS/SVG-drawn, PNG export), a social kit (profile/banner/post/story), an investor pitch deck (keyboard-driven, print-to-PDF), and a business-plan generator v2 (Chart.js, Document/Slides toggle).
- Company Memory & Skills — a persistent structured facts store (key/value/source/timestamp) plus a learned-skills store, an AI extraction agent, and a transparency panel feeding downstream document generation.
- The Boardroom turns a one-click P&L into a branded shareholder-grade deck, architecturally enforced: the compute module is pure (zero AI/network — numbers can never be AI-generated or edited by construction), the AI module is narrative-only and cannot touch numbers, and charts are inline SVG only (CSP blocks external chart libs).
- Document Center — a DB-driven template registry ("adding template #41 is a row, not a release"), AI drafting behind the same allowlisted field/clause fence, and explain/rollback/edit history per document.
Agency / multi-tenant white-label
- A tenants table with slug provisioning (reserved-slug guard) and seats (owner/admin/member, seat-limit enforced); host-resolution middleware routes <slug>.swiftbrand.studio → req.tenant, painting the agency's own name/logo/colors with a "Powered by SwiftBrand" footer and hiding signup for team members.
- An Agency API for provisioning, seats, and branding upload, plus seat-invite emails with claim tokens — verified per commit via E2E + Playwright-vision checks (seat 200 / non-seat 403 / root 200).
Billing, auth & SEO
- Stripe and LemonSqueezy code both present; LemonSqueezy is active with a webhook that fails closed on a missing secret or signature, proration via PATCH-not-new-checkout (feature-flagged off by default), and a billing portal + cancel in Settings (staging).
- JWT auth: email/password + verification, OAuth (Google/GitHub/Facebook via Passport), and API-key issuance for the programmatic/business tier (bulk research, generate/revoke keys).
- Programmatic SEO: registry-driven niche pages (206 static HTML pages live), each with a real domain-checked name grid, FAQPage + BreadcrumbList + WebApplication JSON-LD, hub-and-spoke linking, and a sitemap, plus a --modifiers mode generating niche×modifier pages with genuinely distinct AI name sets per tone.
- First-party, privacy-safe analytics: daily-salted visitor hashing (no raw IP), UTM/source classification, bot flagging, zero external dependency.
Admin console & infrastructure
- An admin dashboard covering economics, traffic, jobs (list/cancel/restart), a live AI-provider toggle, name-template CRUD + toggle, SEO analyze, UAT run, and a User Journeys panel of auto-captured narrator+critic screenshot galleries per flow.
- A Postgres-backed durable job queue (SQLite fallback) with a pluggable admission backend — in-memory by default, swaps to Redis/BullMQ for true horizontal multi-box scaling with pub/sub progress and zero caller changes — plus report-scheduler admission control (paid lane drains before free lane, per-user concurrency cap).
Recent hardening
- Chronological proof of active engineering: a fail-closed payment webhook, site/document AI agents hard-pinned to DeepSeek for security, an EU cookie-consent banner (CF-IPCountry), a Privacy Policy page (LemonSqueezy MoR framing), and a domain-checker false-positive fix.
Planned
- Full 50-state trademark registration (top-10 states live now, OpenCorporates path noted) and international trademark APIs (scraping today — a genuine constraint, not a shortcut). LemonSqueezy proration is built but flagged off; billing portal/cancel and version rollback are marked staging-only, not confirmed promoted.
prov (estimate): Steve-vs-team basis is valued from real git history, actual line counts, and the defect tracker — bounded low by ~25 tested lines/developer-day, high by COCOMO II at a $135K/yr blended rate. Full methodology & sources at apagency.ca.
A full marketologist, not an ad generator
Hand him a folder of footage. Get back tested campaigns — cut, voiced, scored, captioned, published, and measured — with a signed log of every decision he made.
Point him at a Google Drive of raw footage — he analyzes and catalogs what's in it.
Auto Editor splices real ads on an EDL timeline: cuts, pacing, voiceover, generated music, and his own synced captions.
Via Meta MCP, Steve puts the ads on Meta himself, runs test campaigns, and reads the results.
With your input at the approval gates — never a live send without a go.
Every generated video ships with an "Audit Steve's Actions" decision log — the real parameters he chose, and why.
Ask us the hard ones
The questions procurement will ask, answered the way we'd answer them in the room.
Who owns the code Steve writes?
You do. All of it, from day one.
Who owns the rules Steve learns from our incidents?
You own your project-tier corpus; it exports with you. The core doctrine corpus is licensed, non-exclusive, and is never resold with your data in it.
What's the Anthropic relationship?
Independent company, built on Claude. Calls run on your Anthropic account, under your DPA, with your keys — metered through the pod's gateway with hard budget caps.
What data leaves our boundary?
Nothing except calls to your Anthropic account. The pod runs in your VPC or a dedicated instance; there is no phone-home.
What happens if Steve is wrong?
Production actions always escalate to a human approval gate. Defects only close through a proof path: a passing test plus evidence, enforced at the database layer — a claim can't close a bug.
What if you disappear? (continuity)
The pod is reproducible infrastructure: versioned corpus, scripted provisioning, daily regression proof. The singular STEVE-1 is the accountability model; the pod is the continuity model.
What does it cost to run?
Your Anthropic token spend, on your keys, visible per-call in the cockpit with hard caps you set. No hidden metering on our side.
Is there an SSO tax?
No. SSO (OIDC/SAML) ships with every multi-seat deployment; enterprise tiers add audits and compliance work, not login security.
What do we keep if we leave?
Your code, your project-tier rules, your boards' full history, and the audit log. We keep the core doctrine.
Why are there no customer logos?
Because you'd be first, and we priced that honestly: the Design Partner door exists precisely because we're pre-customer-one. The receipts on this page are the reference.
prov: these answers are the contract defaults, not marketing — bring your redlines to the call.
prov: book time with steve
Put 30 minutes on the calendar
This calendar is self-built — no Calendly, no third-party scheduler. It's also a live demo of what Steve ships for you.
The deep-dive materials — the audit companion and the offer sheet — are shared openly on the call. No NDA to see how it works.
Pick a day above to see open times (ET).







