The problem was never storage. Your operators use the same Claude you do — the difference is that your method lived in your head and theirs lived in whatever they typed into the box. The Audience encodes the method as skills, serves brand truth through a server that enforces it, and lets Claude do the work.
Outsourced work came back as competent, forgettable lists. Not because the tools were worse — because the model fills an absent method with the statistical average of all marketing writing ever published. The statistical average of marketing writing is a listicle.
Fifteen years of pattern-matching in one head. Everyone else started from a blank prompt box, so the model supplied the average.
Four people, four slightly different ideas of what each brand actually is. Drift compounds quietly and nobody notices until a campaign ships.
Review was a convention. "PASS 34/40" is a string anyone can type, and nothing checked whether the review happened.
Solve the first and their output improves immediately. Solve the third and you stop being the bottleneck.
Solve the second and it scales past four people. That ordering is the whole design — and it's why nothing here is a dashboard.Four layers, each locked by construction rather than by discipline. A plugin is installed, not opened. A tool absent from your scope cannot be called. Credentials you don't hold cannot be spent.
How we do marketing, written down. Every skill carries forcing functions rather than descriptions: the Inversion Test, the Cover Test, the Attribution Test, and a list of words that are never used.
Serves canon and enforces the gate. Two scopes set by token, and the member scope doesn't contain the owner tools — they're absent from the listing, so there's no config to edit and no settings file to delete.
Veo for video, Gemini Image for stills, Lyria for beds, TTS for voice, and AVTool — local ffmpeg, free. Claude generates inside the workspace with canon already loaded, instead of app-switching and losing the strategy in the copy-paste.
Proven work, honest retros including the failures, and a prompt ledger recording every generation: prompt, model, seed, parameters, cost, verdict. This is the part that makes the system worth more every month instead of the same amount forever.
Every guardrail that relies on discipline eventually fails, usually quietly. These four fail loudly or don't fail at all.
Each brand's identity ends in a literal prompt stem, returned by stem_get verbatim
with a sha256. Nobody reads a document and retypes it, so nobody paraphrases it. Byte-identical
reuse across four people is the entire reason the output reads as one brand.
The gate requires a verbatim quote from the deliverable for each of eight dimensions — and checks that each one is actually there. You cannot cite what you did not read, so a fabricated review fails at the boundary rather than at someone's desk.
Media generation bills per call. Operators have no Google Cloud credentials at all, so overspend isn't discouraged — it's structurally impossible. They write briefs; the brief is where the creative thinking lives anyway.
Every empty result returns not_found with a next step, never a plausible substitute.
Ask for a stem that doesn't exist and you get "do not generate for this brand" — not a
generic stem, not an inferred one.
Both submissions below score 32/40 with all eight dimensions filled in. One is accepted. The difference is the only thing that matters.
// scores look fine, evidence is invented gate: { scores: { craft: 4, proof: 4, ... }, evidence: { craft: "a strong and compelling piece", ... }, auto_fails: [] } REJECTED — the gate did not pass. · evidence.craft: quoted text does not appear in the deliverable — a review citing text that isn't there was not performed
// every quote is really in the body gate: { scores: { craft: 4, proof: 4, ... }, evidence: { craft: "Session notes are in the sleeve, unedited", ... }, auto_fails: [] } Accepted at 32/40 with all eight dimensions evidenced. → queued for review
The gate isn't distrust. It's the mechanism that makes a score mean something.
An operator who writes twenty evidenced reviews has learned to see the failures before writing them. That's the actual goal — the quality floor is a teacher, not a filter.For copy, the cheap path is to draft and let the gate kill it. For a Veo batch, that same loop spends money to learn something a paragraph of text could have told you.
Fix it with ffmpeg. Never regenerate.
Wrong crop, needs a logo, wants a GIF version, audio too loud — all free local operations. Reaching for a fresh Veo call to fix a crop is the single most common way to waste money here.The infrastructure is ahead of the content — which is the right way round, but it means nothing produces output yet. The bridge correctly refuses to generate for a brand with no canon. Steps 1 and 2 are the only ones that can't be delegated.
71 files · 3 plugins · 12 skills · bridge 49/49 · verify green · review console · runbooks
Nine intake answers → derive → promote. Blocks everything below it.
Budget alert first. Nine images, Cover Test, tighten the stem. No Veo yet.
github-mcp-server read-only + sync heartbeat. Canon in the team's hands, no code written.
An outage degrades to read-only, not to nothing. Cache miss still returns not_found.
Five skills, one namespace conversion each. Ours wins on conflict.
Re-score independently, track the delta per person. The quality floor, measured.
The one-way door. Immutable stem bodies, cost on every row.
Same tools, bearer token, no per-person install.
Approve briefs from a phone. UI for state — never a generate button.
Assembled from evidence, versioned. Proposes only — never auto-publishes.
Worth being precise about which half of the advantage erodes, because it changes where effort should go.
Models get better at this every few months and the floor rises for everyone, competitors included. The current edge here is real and it is temporary. Building the system around it would be building on something that melts.
What worked, for whom, on which channel, with numbers attached. A better model raises everyone's blank-page output — it does nothing to give a competitor your record of what moved your audiences. That record cannot be prompted into existence.
The durable asset is the library, not the strategies.
Which means: instrument from day one. Every campaign gets a retro with real numbers, especially when they're bad. A year of honest retros is a moat. A year of good strategies with no measurement is a portfolio, and portfolios depreciate.