Creative AI: what it does well, and where it breaks

"creative AI" is a phrase people use confidently and mean three unrelated things by, which is why the conversation about it never converges. it is worth pulling apart, because two of the three are working better than the marketing suggests and one of them is working considerably worse.
the three:
- generation — a model producing an image, a video, a voice, a script from a description or a reference.
- variation and volume — the same idea rendered forty ways, at a cost per attempt that makes forty attempts sane.
- judgement — the claim that the same technology can tell you what to make, what is good, and what to do next.
they get sold as one product and they are not one thing. generation is a solved-enough engineering problem. variation is genuinely transformative and misused constantly. judgement is where nearly every disappointment lives.
what it does well

generation from a reference
the least contested capability, and the one that has moved furthest fastest. from a practitioner, stated flatly:
"Image generation is at a point now where it's almost impossible to tell if it's a real person or not."
that is a stronger claim than most vendors dare make, and it comes from someone with no reason to flatter the technology. it is also narrower than it sounds. stills, from a good reference, in the hands of someone with taste is the sentence that survives scrutiny. the further you get from any of those three conditions — motion, invented from nothing, prompted carelessly — the faster the claim degrades. the corpus is full of people reporting exactly that gradient, and it is worth reading the section below on where it breaks before treating the line above as permission.
the enthusiast version of the same point is worth including because it captures the mood accurately even if it overstates:
"The only limit is your imagination, the tools are already good enough to create crazy shit."
reading things nobody has time to read
this is the least glamorous capability and the most reliably useful, and it almost never appears on a product page. the corpus records it as an aside about a friend doing customer research the hard way:
"he spent 5-6 hours reading 1,000+ collagen reviews on reddit by hand"
that is the shape of a real win: a task that is unambiguously mechanical, where the value is in coverage rather than insight, and where a human doing it is a waste of a human. summarising, sorting, clustering, tagging, extracting — all of this works now and works well.
notice how different it is from "tell me what to do about my ROAS". one is compression, which models are good at. the other is judgement, which they are not.
making the attempt cheap enough to be an experiment
the underrated one. most creative does not work — you run a lot of it to find the small number that does. anything that lowers the cost of an attempt changes the strategy, not just the budget, because it changes how many things you are allowed to be wrong about.
the sharpest use of this in the whole corpus is not "replace the humans". it is:
"i use this to validate a format and once i find a good one, i get human UGCs to produce it at scale without risking wasting a bunch of money"
generate to search, produce to ship. that sequencing shows up repeatedly among the people reporting wins and almost never among the people reporting disappointment.
where it breaks

judgement
the gap between what an agent is sold as and what it does is the widest gap in the category. the most useful description of the actual skill required comes from someone listing what a person still has to bring:
"Fifth, being able to assess when the machine is generating garbage, which is quite often when it comes to AI generated creatives."
that is the job now. not producing, assessing. and it is a skill that the tools cannot supply because supplying it is the thing they are bad at.
at the organisational level, the same disappointment shows up as headcount:
"plenty of companies are still struggling to get meaningful ROI from AI, and many are hiring more people to integrate and manage these tools effectively"
hiring people to manage the labour-saving device is not a scandal, it is just the normal shape of an immature technology. but it should be priced in rather than discovered.
state, references and continuity
the second break is more technical and costs more money than people admit. generation is stateless in a way that creative work is not. you have a reference image, a brand, a character, a set of constraints, and the tooling around the model is frequently worse at holding onto those than the model is at rendering them.
"The agent tends to mix up references, i wasted probably about 3k + credits before i gave up and switched to manually pasting prompts in the studio"
three thousand credits is not a rounding error, and note where the person landed: doing it by hand. the automation layer was the part that failed, not the model.
the same person's advice afterwards is the practical rule:
"If you don't read the prompts before generating you will waste credit$"
and on the other side, someone reporting what actually improved their hit rate:
"This improves: product presentation scene consistency lighting quality overall video generation results while reducing the number of failed generations and wasted credits."
the common thread: the wins come from controlling the input, and the losses come from delegating that control.
the meter
every generation is metered, every failure is metered, and the failures outnumber the successes. this is the most consequential difference between generated creative and commissioned creative, and it is the one people account for worst.
"Without it, you're wasting time, wasting credits on the AI tools and software you're using and it's frustrating."
"All this is very time consuming and burning through a majority of our money"
if you do not know your discard rate, you do not know what your creative costs. a tool at a tenth of the price with six times the failure rate is more expensive, and it will not feel more expensive, because the failures arrive as a slowly draining credit balance rather than as an invoice.
the churn
the fourth break is not about any tool. it is about the number of them. this line was written by someone who builds in the category, about their own product:
"It’s also a directory of every other shit AI tool exactly like mine, because apparently 20 of them launch every day."
that is the honest state of the market, from inside it. the practical consequence for a buyer is that most of the evaluation effort you spend is wasted on things that will not exist in a year, and the corpus has the corresponding exhaustion:
"Most people are wasting money on AI tools they don't actually need."
"im not a designer and freelancers charge like the ads print money, so i keep cycling through these tools trying to find the ones that earn a spot."
"the ones that earn a spot" is the right frame. the churn is not going to stop; the defence is to hold a small number of things that survive contact with your actual work, and to be unsentimental about the rest.
the argument that is not about capability at all
the most interesting objections in the corpus are not about whether the technology works. they are about what it does to the thing being advertised.
"Everyone talks about “authenticity,” but most of the UGC we are using is scripted, paid-for, and anything but real."
that is a fair shot at the pre-AI status quo, and it complicates the usual framing. the authenticity that generated creative is accused of destroying was, in a large share of cases, already a purchased performance. which leads to the genuinely open question:
"The realism raises an interesting question: If audiences can’t easily tell whether a presenter is human or AI, does the creator behind the content still matter?"
nobody in the corpus answers that convincingly, and neither will i. what the corpus does supply is evidence that the audience's stated position and their revealed behaviour point different ways. the disgust is real, loud and easy to find. so are reports of the same material converting. the most credible reconciliation on offer is that the objection is not to synthesis as such but to being shown a fabricated version of something the advertiser could simply have photographed — which is a claim about honesty, not about pixels.
the honest scorecard
pulling it together, ranked by how much the reality matches the pitch:
| what it is sold as | how it actually goes |
|---|---|
| generating images from a reference | works, and better than most people expect |
| producing variations at volume | works, and is the real economic change |
| reading and summarising large volumes of text | works, and is under-sold |
| holding your references and constraints across a session | patchy, and expensive when it fails |
| producing finished video with continuity across shots | the hardest thing in the category, by the accounts in the corpus |
| telling you what is good | does not work; this is still the human's job |
| telling you what to do about performance | does not work, and the confident tone makes it worse |
the pattern is consistent: it is good at production and bad at direction. every disappointed account in the corpus is somebody who bought direction.
so where does that leave you
three things follow, and each is covered properly elsewhere.
the cost has moved from producing to deciding. when attempts are cheap, the constraint stops being your budget and starts being your ability to read the results. one person named the transition exactly:
"Are you seeing AI creative produce real winners, or just more noisy tests?"
the sameness problem is a prompting problem, not a model problem. everything generated with the same aesthetic vocabulary converges on the same look, which is why so much of it is instantly recognisable and why buying a better model does not fix it.
and none of this changes what makes an ad work. the concept, the person you are aiming at, the situation you are describing and the proof you offer are the same as they were. generation changes what it costs to try one. it does not change what a good one is.
worth ending on the line that captures the whole subject in seven words, from someone describing why teams are struggling with all of this at once:
"between consumer skepticism, tool fragmentation, and the gap between “cool image” and “effective marketing asset,” most teams are struggling to harness this shift"
the gap between a cool image and an effective marketing asset. that gap is where all the work still is, and no model closes it for you.
the summary
- "creative AI" is three things: generation, variation at volume, and judgement. the first two work; the third does not.
- generation from a good reference is genuinely strong, and weakest exactly where people push it hardest — motion, continuity, and inventing things from nothing.
- the underrated win is reading, not making. compression works; direction does not.
- generate to search, produce to ship. the practitioners reporting wins use it to find the format, not always to be the final asset.
- the meter is the hidden cost. every failed generation is billed, and a discard rate you are not tracking will quietly make a cheap tool expensive.
- the tool churn is structural — the honest description from inside the category is that twenty launch a day. hold few, and drop them without sentiment.
- it is good at production and bad at direction. every disappointed account is someone who bought direction.