A paid social account wants more creative than anyone can draw by hand. One message across four placement ratios and three click destinations is already twelve files, before a seasonal hook, a second route or a new photo enters the test. Every refresh starts the count again, and every file is a fresh chance for drift: last month's offer, a headline clipped at 9:16, an App Store rating on an ad that sends people to Google Play.
The usual answer is ad creative automation rented from a platform: a product feed goes in, resized variants come out, and the logic lives in somebody else's product. What follows is the version you own. I built it for Bubble's paid media across Meta, Google and Apple, where it now produces the static Meta and Google ads. The design system lives in Claude Design, with every testable axis exposed as a parameter. Batch specs in YAML say what to render. Headless Chrome renders from a git repository, so every file traces to the exact commit that drew it. A review gate stands between render and ship, and a frozen filename grammar joins each asset back to its performance, which is what lets an agent reading live results through MCP servers draft the next batch.
It runs in this order, with the gotchas that bit and, in the house tradition, the parts I deliberately left out.
A production problem, not a design one
Designing a good ad is not the bottleneck on a performance account. Producing the forty variations of it that the placements, platforms and tests demand is. A launch pack of eleven concepts across six ratios and three destinations is 198 files. Each one has to carry the right trust signal for its destination, the right copy for its season, and safe margins for a placement whose interface covers the bottom fifth of the screen. Done by hand, that is days of resizing, and the errors are the quiet kind nobody spots until the ad is live.
Hand production has a second cost that matters more over time: nobody can say afterwards exactly what a given file was. Which headline, which photo crop, which version of the badge? When results come back, the creative that won is often impossible to rebuild faithfully, and the creative that lost cannot be diagnosed, because the facts of what it contained live in a designer's layers file rather than anywhere a report can reach.
Creative automation, at scale, or an ad factory
The field uses several names for this job. Ad creative automation and creative automation are the common ones; "producing ad creative at scale" is the question form; "ad factory" is the newer shorthand for doing it with an AI design tool. They describe the same thing: one design, many correct outputs, without a person resizing each one.
Two near neighbours are different jobs. Dynamic creative optimisation, sometimes called programmatic creative, assembles an ad per impression at serve time inside an ad platform; what I describe here produces finished files before anything is uploaded, which is what Meta, Google App campaigns and Apple's store surfaces actually accept. And an AI image generator invents pictures; this system composes designed templates around imagery that is already cleared for paid use. A generated scene can be one of those inputs once it has been signed off, but the renderer never improvises.
Parametrise the template, not the file
The whole system rests on one design decision taken before anything is automated: every axis you intend to test must be a parameter of the template, never a manual edit. Claude Design builds components that read their settings from the page's URL query string, so a template is a web page whose route, season, ratio, style and destination are chosen in the address bar. Insist on that convention while the design system is being built, because retrofitting it later means redesigning every template.
Four preconditions make a template renderable without a person in the loop:
- Every testable axis is a parameter. Message route, season, location copy, visual style, layout, ratio and click destination. If a test needs a manual edit, the test cannot be batched.
- One pixel size per ratio, everywhere. 1:1 is 1200 by 1200, 4:5 is 1080 by 1350, 9:16 is 1080 by 1920, 1.91:1 is 1200 by 628. Write the table into the design system so the renderer can assert it on every file.
- Editor-only overlays key off the editor. Safe-zone guides and crop marks must hide on the host editor's own flag, not on a print media query; then a headless render is clean by construction. Check it once per template.
- Assets live inside the project. Fonts, photography and logos are self-hosted in the design system, so a render on any machine produces the same pixels.
Destination deserves a sentence of its own. It changes more than a label: an app-install ad bound to iOS shows the App Store rating, one bound to Android must never show the weaker Play rating, and a web ad carries a review-site score instead. When the destination is a parameter, the right trust signal is a property of the template rather than something a designer has to remember at 6pm.
The batch spec is the brief
With parametric templates, a creative brief becomes a short YAML file: which templates, which settings, which axes to cross, and how to name the results. The renderer reads the spec, expands it into one URL per file, and checks every setting against a schema extracted from the design system before it opens a browser, so a typo in a route name fails in a second rather than after two hundred renders.
batch: half-term-202610 version: 2 month: "202610" hypothesis: "a half-term hook beats the evergreen route for parents of school-age children" contract: total: 12 concept: { offer: 6, statement: 6 } # pin what you mean, not a proxy for it dest: { ios: { min: 2 }, android: { min: 2 } } defaults: location: London jobs: - template: ad-offer tweaks: route: "Trust first" season: "Half-term" location: "{location}" matrix: # cartesian: every dest at every ratio destination: [iOS, Android, web] ratio: ["4:5", "9:16"] name: { concept: offer, route: trust-first, season: half-term, style: photo } - template: statement list: # rows: fields that vary together - { route: "Time back", ugc: "Kitchen" } - { route: "School run", ugc: "Gates" } matrix: ratio: ["1:1", "4:5", "9:16"] name: { concept: statement, route: "{route}", style: photo }
Three parts of the spec carry the method rather than the plumbing.
- A matrix and a list. A cartesian
matrixcrosses axes that really are independent, such as destination and ratio. Alistholds rows where several settings must move together, such as a route and the photograph that fits it. Real packs need both; a matrix alone renders combinations nobody would brief. - A hypothesis. One sentence saying what the batch is meant to learn. It keeps the spec honest at authoring time and tells the next reader what to look for in the results.
- A contract. Minimums and exact counts per concept, ratio and destination, checked before rendering. If the planned mix breaks it, nothing renders.
One rule keeps the batch maths sane: a variant is a materially different message, and ratio is an export dimension, not a variant. Four ratios of one headline are one test, not four. Without that rule, a system that makes variants free will fill an ad set with near-duplicates that split the budget and teach you nothing.
A filename that carries its own test
Every rendered file is named by a fixed grammar that spells out the axes it tests, and the renderer writes a manifest.csv beside the files mapping each name to the exact settings that drew it. The name is the join key. When a Google Ads asset report or a Meta ad-level export comes back, each row can be traced to a file, the file to its settings, and the settings to a commit, without anybody opening a design tool. This is the ad creative naming convention most guides recommend, with the step they skip: the name is generated from the spec, so it cannot be wrong.
Two practical notes. Put a concept field first even if your client's naming spec leaves it out: a route on its own rarely identifies the template, and the join needs to. And expect at least one platform not to keep your filenames. Meta reports at ad level, and ads that carry several ratios cannot be named after any one file, so give ads a shorter label derived from the same grammar (drop the ratio and layout, add the person featured) and keep an explicit map from ad to files. Never recover the join by string surgery on names someone typed by hand. The same instinct runs through carrying campaign names into every URL: a name you control is the cheapest join key you will ever get.
Render, then review, then ship
The renderer is deliberately simple. It serves the design system from the repository, drives each URL in headless Chrome through Playwright, waits until the page has finished settling (network idle, fonts loaded, images decoded, two animation frames and a short pause, because container-query layouts settle late), finds the ad's canvas and screenshots it. It asserts the canvas is exactly the size its ratio promises and fails loudly if not. A handful of pages in parallel renders around eighty files in a couple of minutes on a laptop.
Every run lands in its own folder named by date, batch and the design system's git commit, with the PNGs, the manifest, a copy of the spec as rendered, and a run.json recording the commit, the counts against the contract, and any failures. That record is the provenance: months later, any asset in any ad account traces back to the exact template and token state that produced it.
Nothing ships straight from a render. A review step runs one agent per rubric over the finished PNGs: composition, legibility, ad policy and brand. Each files structured findings to the run folder, and the run is approved only when every finding is fixed or accepted with a written reason. Fixes are made in the template or the spec, never in the file, and only the failing subset re-renders, amending the same run folder so the record of what was reviewed stays whole. Two or three rounds is normal; later rounds are cheap because earlier acceptances are briefed as settled.
One gap in that list took a real false alarm to find. Rubrics for composition, legibility, policy and brand will all pass a confident, on-brand, perfectly legible fake. If the creative depicts the product's own interface, as app-install ads usually do, add a rubric that owns "the depicted screen matches the shipping app", and give it the real screenshots as its reference rather than a description of them. A question about pixels has to be answered with pixels.
Derived state rots silently
Almost everything that went wrong in the first months was one of two shapes: a derived file going stale with nothing to say so, or a check that could not see the thing it claimed to check. Both are invisible by nature, so they deserve rules rather than vigilance.
- Ship the masters, and build packs as pick lists. Every channel wants its own bundle ("just the twenty for Meta"), and the tempting answer is to re-render a subset under a new batch name. That throws away the review approval the files already earned and creates copies that drift. Keep a selection spec instead, build the pack from approved masters at the moment of upload, and make the builder refuse anything unapproved.
- An approved pack stops being approved when the design system moves. The record still says approved and the files are untouched, but a template fix that landed since means the shipped assets no longer match what the system would draw, and may still carry a defect fixed everywhere else. Re-render every batch before any ship and compare each one with its last approved run.
- Renders are not byte-reproducible once a photograph is on the canvas. Resampling noise changes a few pixels in photo regions on every render, so a checksum cannot prove "nothing changed" when carrying an approval forward. Use a pixel diff with a threshold, and check the changed area touches no text.
- A contract must pin the thing you care about. A rule that said "at least eight Android files" passed for weeks while every Android file was a store-listing screenshot and the ad test ran on iOS only. Pin the concept per platform, so the check fails when the thing you meant goes missing.
Don't hand-edit a rendered file, even to fix one word. The fix belongs in the template or the spec; a corrected PNG with no source is the first derived artefact that will rot, and the next render will quietly undo it.
Close the loop with MCP
The point of a join key is the loop it allows. Performance comes back, is joined to the manifests on the filename grammar, and decides the next batch: keep what earned its spend, retire what did not, add one new axis to test. The formal route is a platform export dropped into the repository. The route that made same-morning refreshes normal is an agent reading the ad accounts directly through MCP servers (the Meta Ads and Google Ads connectors), reporting performance by concept, route and creator, and drafting the next spec, hypothesis and contract included, for me to approve.
The division of labour is the one described in the agent-operated second brain: agents propose, I decide. The agent is good at the tedious parts, reading forty ads' worth of spend and cost per registration, matching them to concepts and noticing which route carried most of a month's purchases. It is not trusted to decide that a creator has fatigued, that a poster belongs in a cold audience, or that a hook is on brand. Those calls stay with a person, and the spec is where they are written down.
An honest caveat on what the loop can read. Small accounts produce thin data per creative, and a week's results rarely crown a winner. The join makes the evidence legible; it does not make it significant. For a read on copy before it costs any spend, a simulated panel such as the Ad Copy Battle is a cheaper first filter than another round of live tests.
What I deliberately didn't build
Most of what keeps this dependable is what it refuses to do:
- No automatic upload. A person puts every pack into the ad platforms. Upload is where budgets, audiences and policy meet, and a pipeline that publishes on its own turns a template bug into spend.
- No video encoding. The system renders the branded overlay and the first-frame thumbnail for a video ad and stops there. Cutting and encoding stay in an editing tool, where a person can judge the pacing.
- No improvised imagery. Photography and any generated scenes are inputs, registered with their usage rights, before a template can use them. The renderer composes; it never invents.
- No rented platform. The templates, the specs, the history and the join are files in a repository the client owns. Changing agency, tool or model costs nothing but the next person's time to read them.
- No unsupervised retirement. The loop proposes what to switch off; it never pauses an ad itself.
Everything above, as the build order:
- Design the system in Claude Design with every testable axis as a URL parameter, one pixel size per ratio, editor-only overlays and self-hosted assets.
- Move it into a git repository and make git the source of truth; the design tool becomes the editing surface, synced both ways.
- Extract each template's parameter schema and validate every batch spec against it before rendering.
- Write batch specs with a matrix, a list, a hypothesis and a contract that pins concepts per platform.
- Freeze a filename grammar that names every tested axis, and write a manifest beside every run.
- Render headless into run folders that record the design-system commit; assert every canvas size.
- Gate every run with one review agent per rubric, including product-truth for any depicted interface. Fix in the template, re-render the subset.
- Ship packs as pick lists from approved masters; re-render everything before any ship.
- Join performance back on the grammar, by export or MCP read, and let it draft the next spec for a human to approve.
The templates were the visible part of the build and the least important. The value sits in the spec, the name and the record: a creative refresh becomes a short file somebody can read, every asset in market can say exactly what it is, and the next test starts from what the last one proved. Owning that is the difference between automating ad production and renting somebody else's guess at it.