We only work with brands spending $40k+/mo on Meta ads.

AI creative

How to Make AI Video Ads That Aren't Slop

Gabe Hutcheon · · 7 min read

The rule that changes everything: do not enter the product in the first 25 to 75 percent of the film. Educate, tell a story, get buy in, then arrive.

There is a lot of AI slop on the feed at the moment, and most of it fails for the same structural reason. The product shows up in the first three seconds, so the viewer is asked to care about something before they have been given any reason to.

The films that work do the opposite. They spend the first half teaching you something you did not know, and the product turns up at the end as the answer. That is not a style choice, it is the whole mechanism, and it is what unlocks a new pocket of audience: people who would scroll past an ad will happily watch a two minute explainer about where something came from.

Build it in this order

Order of operations matters more than any individual prompt. Ours runs like this, and steps two to four are the ones everybody skips.

  1. Research the spine. Three or more facts you have actually verified, not recalled. Real history reads as a film. Invented beats read as an ad.
  2. Write the voice-over. Before any picture exists.
  3. Generate it and measure every line in seconds.
  4. Derive shot durations from the audio. Lead in, plus the lines, plus the gaps, plus a tail. Never pick a shot length before you know what it has to carry.
  5. Write one prompt per shot, restating the style every time. The model has no memory between shots.
  6. Render one shot and look at it before committing to the set.
  7. Assemble the silent master, then lay the voice-over over the top.

The dead giveaway of a film built the wrong way round is thirty seconds of silence in the middle where the narration ran out and the pictures kept going.

One shot is a scene, not a picture

A shot is up to fifteen seconds of continuous animation carrying three timecoded beats, and something has to happen in each of them. Something enters, tears away, fills, collapses, is revealed. Hard cuts inside a single shot are allowed and good.

The failure mode is a sparse diagram: one icon centred on a flat background. It reads as a slide, not a film. Scenes need people, places and depth, and light that comes from somewhere inside the scene, a candle or a desk lamp or floodlights.

The other thing that kills a film is sameness. Before writing a single prompt, fill in a diversity matrix: for every shot, declare location, scale, light source, time of day and dominant motion, and do not let two adjacent shots share more than one of those. That single constraint is what stops six variations of the same frame.

Never bake the voice-over into the render

Some models will happily generate narration inside the clip. Do not let them. End every prompt with an explicit ban on speech, and ask for diegetic sound design instead: the paper tears, the fire crackles, the crowd roars.

What you get is a silent but fully scored master, which is an asset rather than a one off. You can re-voice it, re-cut it, run a completely different script over the same pictures, or hand it to an editor. Change the words later and it costs one text to speech call instead of a whole render.

Verify the ban actually held. Run a transcription pass over the finished master and look for your script. Transcribers invent fragments over pure ambience, which is normal. What you are checking for is the absence of a voice you did not choose.

What the models still get wrong

Short words in capitals set reliably. Date stamps, single keywords, that is the ceiling. Anything longer garbles. A full frame hero shot of your packaging will come back repainted with misspelled label copy, so keep the product mid-size with a reference image attached, and typeset any end card yourself afterwards.

Watch brand names that are not spelled the way they sound, too. Ask a model for an end card and it will happily typeset the pronunciation instead of the spelling, which is how you ship a dead URL.

The arc

Unexpected origin, then the strange detail nobody knows, then the mechanism, then why it matters now, then the product. Last.

Write the film you would actually watch. The ad is what happens once you have earned the right to mention the product.

Get the skill and the cheat sheet

The Claude skill that runs this whole workflow, plus the one-page cheat sheet. Free. Your email, no other hoops.

If you want us to run this for your brand instead of running it yourself, book a free creative audit.

Frequently asked questions

Why do most AI video ads look like slop?
Because they are built backwards. Someone renders a nice looking five second clip, writes copy to fit it, then bolts the product onto the front. The result is a product demo with no reason to watch. The fix is structural: hold the product out of the first quarter to three quarters of the film, spend that time educating, and let the product arrive as the answer to something the viewer now cares about.
Should the voice-over be generated inside the video model?
No. Bake sound design into the render and ban speech explicitly, then add the voice-over afterwards. A film with the narration baked in is frozen: you cannot change a word, re-cut it, or run a different angle over the same pictures without paying to render it all again. A silent-but-scored master is an asset you can re-voice forever for the price of a text-to-speech call.
How long should an AI video ad be?
Four to six shots, roughly 45 to 75 seconds, then cut a 30 second version from the master. Writing short from the start tends to produce an ad. Writing the full film and cutting it down produces something worth watching, and the cutdown inherits the story.
Why write the voice-over before rendering anything?
Because shot length has to come from the script, not the other way around. Render fixed length shots first and you end up with long silent stretches in the middle of the film where the narration ran out. Write the lines, generate them, measure each one in seconds, then derive how much footage each shot actually needs.
Can AI video models render text and packaging accurately?
Short words in capitals, usually. Anything longer, no. Date stamps and single keywords set reliably. Body copy comes back garbled, and a full frame hero shot of your packaging will get repainted with misspelled label text. Keep the product mid-size in frame with a reference image attached, and typeset any end card yourself.

Think creative is your bottleneck?

We brief from $250M+ in tracked ad spend and put first drafts in your hands in 48 hours. Book a free creative audit and we'll show you where your account is leaking.

Book a free creative audit