AI ad creative: why it looks like AI, and what to do about it

the most-engaged line in the whole corpus about ad creative is not a question about tools or prices. it is this:
"do companies not realize AI ads look genuinely fugly or do they just not care?"
it deserves a real answer rather than a defensive one, so here is the honest version: they mostly do not see it, and the ones who do see it have run the arithmetic and decided the ugliness is cheaper than the alternative. both halves of that are true at once, and both are fixable, which is what the rest of this is about.
but before the "what to do", the diagnosis has to be right. "looks like AI" is not one failure. it is at least five separate ones with five separate fixes, and people who try to fix it by buying a better model usually only fix one of them.
why it looks like AI

one: everybody is prompting with the same forty words
this is the biggest single cause and almost nobody names it.
go and read what people are actually typing. the corpus is full of prompt strings posted in public by unrelated accounts, and they converge on an identical vocabulary:
"Luxury pastel blue handbag cinematic product ad, ultra-realistic, 4K vertical."
"Soft natural studio lighting, clean beige background, slow camera movements, premium fashion aesthetic, ultra-realistic, shallow depth of field, 4K vertical video."
"Ultra realistic, epic scale, volumetric lighting, dramatic atmosphere, anamorphic lens, 4K HDR, highly detailed, cinematic camera slowly orbiting around the character, 16:9."
different people, different products, same eleven adjectives. ultra-realistic. cinematic. volumetric. shallow depth of field. 4K. highly detailed.
those words do not describe a scene. they describe a look, and every model in wide use has one specific rendering of that look sitting near the middle of its distribution. so when a hundred thousand advertisers reach for the same six words, the model returns the same image with different objects in it. the "AI look" is not a limitation of the technology. it is the average of what everyone asked for.
this is why the sameness gets worse as the models get better. a better model renders the average more convincingly. it does not move you off the average — only your prompt does that.
two: density
the sharpest description of the visual tell in the entire corpus, from someone who does not work in advertising at all:
"What annoys me the most about the look of AI ads and posters is how unnaturally DENSE everything is."
that is exactly right and it is worth understanding mechanically. a photograph is mostly accident. it has a dull corner, a badly lit patch, something slightly out of focus, an irrelevant object at the edge that nobody chose. a generated image has none of that. every region has been rendered to the same level of interest, every surface has texture, everything is lit, nothing is boring. it reads as composed in a way real capture almost never does.
nobody consciously identifies this. they just get a feeling that it is fake, half a second before they could tell you why.
three: physics, hands, and things that cannot happen
the funniest evidence that the people producing this stuff know exactly what gives it away is that they maintain written lists of the tells and paste them into the prompt as things to avoid:
"No face changes, age changes, hairstyle changes, clothing changes, duplicate people, warped hands, extra fingers, distorted food, floating objects, impossible shadows, inconsistent lighting, artificial facial expressions, unnatural walking, rubbery motion, text artifacts or CGI-looking environments."
"Avoid: Floating boards, incorrect foot placement, impossible tricks, warped wheels or trucks, extra limbs, identity changes, outfit changes, synthetic skin, fake bokeh, CGI surfaces, background morphing, duplicated pedestrians, artificial stabilization, text, subtitles, logos and watermarks."
two different people, two different subjects, and a large overlap in the list. that overlap is the actual failure catalogue for the current generation of models. it is worth reading twice, because it is more useful than any "how to spot AI" article — it is written from the inside by someone whose money depends on getting it past a viewer.
four: drift
this is the tell that is unique to motion and it is the hardest to fix.
"Tired of random AI clips that lose the face, outfit, or room halfway through?"
"How do you get such consistent identity of the characters across such a long video?"
a still only has to be right once. a sequence has to be right, and the same, across every frame and every shot. faces slide, a jacket changes colour between cuts, a kitchen loses a window. viewers cannot articulate it but they clock it instantly, because continuity is something the human visual system tracks whether you ask it to or not.
five: the tell that is not visual at all
the most interesting one. from an analysis of an AI persona that people eventually rejected:
"It created a cognitive dissonance for viewers: How can this girl be so consistent, always posting perfectly, and reply so personally?"
nothing in any individual frame was wrong. the pattern was wrong. real people are irregular: they post badly, they miss days, they look tired. a thing that is always on-model reads as synthetic even when every frame passes.
which means "make each asset look more real" is an incomplete strategy. the giveaway can live in the cadence, the volume and the consistency of the output rather than in any one asset.
the thing this article would be dishonest to leave out
one line in the corpus flatly contradicts the premise of the whole piece, and it comes from a practitioner rather than a bystander:
"Image generation is at a point now where it's almost impossible to tell if it's a real person or not."
they are right, and the two positions are reconcilable in a way that matters for how you spend.
stills have largely crossed. motion, sequence and context have not. a single generated photograph, prompted by someone with taste and starting from a real reference, is now very hard to call. the surviving tells cluster in exactly the places the lists above describe: movement, continuity across shots, hands doing things, and the relationship between the ad and the thing it is selling.
that is a genuinely useful allocation rule. if you are unsure whether you can pull this off, your odds are much better on a still than on eight seconds of video — and much better on eight seconds of one continuous shot than on a four-cut sequence with a recurring character.
why companies ship it anyway
two reasons, and only one of them is cynical.
the first is that they never see it the way you see it. the person making the ad looks at it alone, at full size, in a generation interface, forty seconds after asking for it, feeling pleased. you see it at a third of that size, in a feed, between two things you chose to look at, having already seen nine ads today that were made the same way. those are not the same object. almost nobody producing this stuff runs the second test.
the second is that the arithmetic genuinely works out, sometimes. stated with no varnish by someone on the buying side:
"why would i pay a creator $300 when i can pay ai less than $5 to create the same lie"
at that ratio an ad that repels most of the people who see it can still clear, because you are not optimising for how the median viewer feels. you are optimising for cost per acquisition against a very small conversion rate. the ugliness is a real cost and it is simply not the biggest number in the sum.
so: yes, some of them know, and the honest form of the objection is not "they don't care about quality" — it is "the feed is a commons and they are dumping into it." that is a fair complaint and no vendor is going to solve it for you.
but the arithmetic is not as settled as the $5 figure makes it sound. from someone who inherits other people's accounts:
"Taking over ad accounts from people who are touting AI UGC is always pretty hysterical, because you see how shitty a job they actually do when it comes to performance"
and asked directly, by a practitioner rather than a critic:
"Why does AI UGC convert so much worse?"
and the sentence that captures the whole trap:
"People are sick of AI slop in their feed and the genuine-feeling stuff is what converts, so they're using the exact thing everyone's tired of to recreate the one thing everyone wants."
the cheap attempt is only cheap if it works often enough. which brings us to the part you can actually act on.
what to do about it

start from something real
the single most-quoted objection in the corpus is not about aesthetics:
"show me the fucking food you make, not this foul worm infested ai generated image"
read that again as a business complaint rather than a rant. the failure is that a restaurant showed a generated dish instead of the dish it sells. it is a claim problem, not a render problem. the same ad with a photograph of the real food and a generated background would have provoked none of it.
so the first rule is structural: the thing you are selling should be the real thing. generate the context, the setting, the variations, the frame around it. do not generate the product and hope. this also happens to be where the models are strongest — working from a real image you supply, rather than inventing from nothing.
delete the aesthetic adjectives
if your prompt contains ultra-realistic, cinematic, 4K, volumetric, premium, highly detailed, you have asked for the average and you will get it. those words are doing no work; every competitor typed them too.
replace them with the specifics of a situation. the best line in the corpus about scripts applies exactly as well to images:
"the script describes the product instead of the situation nobody is moved by "premium stainless steel construction." they're moved by the 6pm moment where the thing they own doesn't work and they're annoyed."
"6pm, kitchen, one overhead bulb, half-unpacked shopping on the counter" is a prompt no competitor is using, and it produces something that does not sit in the middle of the distribution.
put the imperfections back in on purpose
the practitioners who are getting away with it are not asking for more polish. they are asking for less. an actual prompt fragment from the corpus:
"Preserve tiny footsteps, vertical camera bounce, imperfect reframing and occasional edge cropping when the operator struggles to keep up."
that is someone deliberately specifying the artefacts of a human holding a camera badly. it works because it inverts cause one: the average is smooth, so roughness reads as real. it is also consistent with what performance people report about production values generally —
"Polish is killing your performance Professionally produced ads get scrolled past because the brain identifies the advertising pattern in under a second and scrolls before conscious attention is paid."
pick a side of the line
the most useful decision rule anyone in the corpus offered:
"Make it look real and it converts Or make it obviously AI, lean into it fully, and the audience respects that you're not hiding anything"
the failure zone is the middle — the almost-real, which reads as an attempt to deceive that didn't quite come off. that is where nearly all the hostility lives.
and there is a third door that sidesteps the argument completely, from someone reporting it as their best-performing format:
"An ad format that is crushing for us right now: comic book ads Instead of chasing AI realism like everyone else, you draw a comic scene of the exact situation your customer is stuck in."
illustration is not competing on realism, so it cannot lose on realism. if your creative is carrying a situation rather than a product beauty shot, this is often a better use of generation than photorealism is.
run the two tests that catch most of it
both come straight out of the corpus and both take a minute.
the sound-off test:
"Then I watched it with the sound off and realized I had absolutely no idea what the video was trying to make me feel."
the "what is this selling" test, from someone describing a whole category of ads on public transport:
"I have no idea what most of the AI ads on the subway are selling tbh"
if a stranger cannot name the product and the feeling within three seconds and no audio, the render quality is irrelevant. this is where most generated creative actually dies, and it has nothing to do with fingers.
measure the throwaway rate, not the generation cost
the number that matters is not what an attempt costs. it is what a usable one costs. two people who did the arithmetic on themselves:
"i throw away a lot, maybe 60%"
"Costing me around $20 in tools/credits to land one ad that actually works."
a 60% discard rate is normal and fine. a 60% discard rate you are not tracking is how a cheap tool turns out to cost more than a freelancer. and the discard rate is the metric that tells you whether your prompting improved, which no vendor dashboard will show you.
keep the record of what produced the winner
the last one is not about looking real at all, and it is the one that quietly costs the most:
"how can i make sure i don't lose track of the sequence it came from so if i want to add or make some changes it would be easier?"
when the reference image, the prompt, the brief and the output live in four different places, you lose the ability to answer "what made the one that worked" a week later — so you cannot reproduce it, and every subsequent attempt starts from zero again. no better model fixes this. it is a filing problem wearing a creative problem's clothes.
the summary
- "looks like AI" is five different failures. convergent prompt vocabulary, unnatural density, physics artefacts, drift across a sequence, and a posting pattern that is too regular to be a person.
- the sameness comes from the words, not the model. everyone types the same aesthetic adjectives, so everyone gets the same average. better models render the average better.
- stills have largely crossed the line; motion and continuity have not. allocate accordingly — one still beats eight seconds, and one continuous shot beats a four-cut sequence.
- generate the context, not the product. the loudest objection in the corpus is about showing a generated thing instead of the real thing, and it is a substantive complaint rather than an aesthetic one.
- specify the situation and the imperfections. roughness reads as real precisely because the average is smooth.
- look real, or be openly synthetic. the middle is where the hostility lives.
- the honest metric is cost per usable asset, and most people have never worked theirs out.
- and keep the record. the cheap attempt is only worth having if you can say what produced the one that won.