All insights

Three tools, ten file moves, four hooks you never tested

Three tools, ten file moves, four hooks you never tested

three tools, ten file moves, four hooks you never tested.

ten is what it takes to get one thirty-second ad out of the setup most people are running. ten times a file leaves one tool and has to be carried into the next one by hand — and that is the clean run, nothing failing, nothing rejected, no client changing their mind.

almost nobody counts that number. everybody counts the tools. three feels lean. three feels like a decision you got right, and it is the number that goes in the reply when somebody asks what you're running.

the four is the bill. we will get to it.

The question that never stays answered

it gets asked constantly, always in the same shape. what are you actually using to make video creatives. which app, which model. one tool from start to finish, or a few mixed together — one for images, one for video, one for voiceover, one for editing.

and the answer that comes back is always a list. which is the whole problem, and somebody already put it better than i can:

"every 'best AI video tools' list mashes 20 products into one ranked pile, which is useless because they do different jobs."

that is exactly right, and it is why the question gets asked again the following week by somebody else. a ranked pile cannot answer it. a generator and an editor are not competing for the same slot. putting them in one order of merit tells you nothing about the thing you actually have to decide, which is what the whole chain looks like from brief to uploaded ad.

so here is a different way to answer it. not a list. a count.

Count the crossings, not the tools

a crossing is every time a file or a piece of context has to leave one tool and enter another. a download. an upload. a prompt retyped into a second box because the first tool cannot see it.

take the stack people describe as the working one — a generator, ElevenLabs, CapCut. plus a doc somewhere with the brief in it, because there always is one. one thirty-second ad with a voiceover. count it out:

  1. brief → generator. you retype the prompt.
  2. generator → disk. download clip one.
  3. generator → disk. clip two.
  4. generator → disk. clip three.
  5. generator → disk. clip four. most generators cap out around eight seconds a go, so thirty seconds is four generations, minimum.
  6. brief → ElevenLabs. you retype the script.
  7. ElevenLabs → disk. download the audio.
  8. disk → CapCut. upload four clips and one audio file.
  9. CapCut → disk. export the cut.
  10. disk → ads manager. upload.
The ten crossings in a three-tool video stack
The ten crossings in a three-tool video stack

ten. three tools, ten crossings, and every one of them is you.

and that is the version where nothing goes wrong, which is not a version that exists. the rejection rate on generated clips is not small — if one in three comes back unusable, four keepers means six generations, and steps two through five stop being four crossings and become six. want a second hook to test? that is not "one more ad", that is most of the ten again.

one person described the outcome:

"wasted 2 hours, got 30 seconds of usable video."

two hours for thirty seconds. and the thing worth noticing is that the generator was not busy for two hours. it was busy for a handful of minutes. the rest of that time you were downloading, renaming, dragging, retyping and waiting for an export bar.

which is why the complaint that shows up over and over is not about output quality at all:

"constantly switching tools kills momentum and makes high output almost impossible."

momentum. not quality, not price. what is being lost is the state you were in — and "high output almost impossible" is the part with money attached, because volume is the entire game in paid social.

Two kinds of crossing, and only one of them really hurts

here is the part worth the whole article, and it is the bit a tools list structurally cannot see.

a file crossing costs you a download and an upload. it is annoying and it is measurable. thirty seconds, a minute, and a downloads folder that slowly fills with final_v3_ACTUAL.mp4. irritating. survivable. you can get faster at it.

a context crossing is different in kind. that is where the brief has to be retyped into the next tool, because the next tool cannot see what the last one was doing. steps one and six above. and a context crossing does not cost you thirty seconds — it costs you fidelity, because you never retype it exactly.

the tone drifts. the product detail that was in the doc does not make it into the voiceover script. the reference image the prompt was built around is not on screen when you write the VO, so the VO describes a slightly different ad to the one you are making. nobody notices at the time. everybody notices in the edit, when the audio and the picture are describing two different products.

that drift is what produces the thing everyone blames on the model:

"sometimes it becomes frustrating on real client projects because of inconsistencies between previews and renders, editing issues or other production problems."

some of that genuinely is the model. a lot of it is that the second tool was working from a worse description of the job than the first one had. you did not give it the brief. you gave it your memory of the brief, typed at speed, at eleven at night, after four rejected generations.

A file crossing costs time; a context crossing costs fidelity
A file crossing costs time; a context crossing costs fidelity

and once you accept that, this line stops sounding like a tooling complaint and starts sounding like a diagnosis:

"i waste a surprising amount of time before and after the actual edit."

before and after. not during. the edit is fine. the edit was never the problem. the problem is on both sides of it, and both sides of it are crossings.

The second-order cost, which is the expensive one

the time is not actually what this costs you. the time is what you notice.

what it costs you is variants. every crossing is a fixed toll you pay per asset, and a fixed toll per asset is a tax on volume. ten crossings for one ad is a bad afternoon. ten crossings times six hooks is why you shipped two and told yourself the other four were probably similar anyway.

there is the four from the top of this piece. not four hooks you rejected on judgement — four you never made, because the tenth file move was the one you could not face at that hour.

so the stack quietly decides your testing strategy for you. not your media buying judgement — your file management. you end up testing the number of variants your friction budget allows, then reasoning backwards about why that number was the right one.

that is the real bill, and it does not appear on any invoice. it appears as a thinner test than you meant to run.

What this changes about choosing a tool

now the original question is answerable, and the answer is not a ranked pile.

stop asking which tool is best at its job. ask which tool removes a crossing.

a generator that is ten percent better than the one you have, but sits in the same slot, removes nothing. you still download, still retype, still upload. you have bought a slightly better step four and paid for it with, as somebody put it exactly:

"just adding another subscription to the stack."

a tool that generates voice and video from the same brief removes crossings six and seven — and more importantly it removes a context crossing, because the script and the visual are now being written against the same thing. that is worth more than a better model in one slot, even if the voice is a bit worse.

which gives you the counter-intuitive rule, and it is counter-intuitive enough that a tools list will never tell you it:

a worse tool in two slots can beat a better tool in one.

nobody ranks tools that way, because a ranking has to compare like with like, and "removes a crossing" is a property of your chain rather than of the product. it does not fit in a column.

The audit, and it genuinely takes ten minutes

do this once, on paper, for one real deliverable you actually made recently — not a hypothetical one, because a hypothetical stack is always tidier than your real one:

  • write out every step in order, from where the brief lives to the ad going live
  • mark every point where something moves between tools
  • put an F next to file crossings and a C next to context ones
  • be honest about the doc the brief lives in, and the downloads folder. those are tools. they are where half the Cs hide
  • count the Cs
A worked crossing audit with five file crossings and two context crossings
A worked crossing audit with five file crossings and two context crossings

the Cs are your shopping list, in that order. spend money and switching-effort there.

the Fs are worth fixing only when there are a lot of them, and the fix is usually a naming convention or a shared folder rather than a purchase. people get this backwards constantly — the Fs are visible and annoying so they get attacked first, and the Cs are invisible and expensive so they survive every reorganisation.

most people doing this the first time find between two and four Cs, and discover they had been shopping to improve a step that was already fine.

"So just use one tool that does everything"

sometimes. not always, and the reason matters.

an all-in-one removes crossings by definition. it also usually means accepting a worse tool in every slot at once, and if the thing you are making depends on one of those slots being genuinely good, that trade is bad. a product ad that lives or dies on the hero shot is not a good candidate for a tool with a mediocre image model, however many crossings it saves.

the honest version of the rule is narrower than "consolidate": remove the crossings around the slot you care least about, and keep the specialist for the slot the work actually depends on. most stacks have exactly one slot that matters and three that do not, and people optimise the three.

and there is a real trap here worth naming, because it is what the phrase "does it end to end" is doing in every tool's marketing. a tool that owns two steps but makes you export between them has not removed the crossing. it has moved it inside a single subscription where it is harder to see. check whether the second step can actually read the first step's context, or whether you are still doing the retyping — just inside one product now.

What i can't tell you

i want to be straight about the edges of this, because the version of this post that pretends otherwise is the version that deserves to be ignored.

i cannot tell you which generator to use. that genuinely does depend on what you are making, and anyone answering it in general is guessing. the crossing count is a way to choose between chains. it is not a way to choose between models, and it does not rank anything.

the ten is a count of steps, not a stopwatch. i counted crossings in a stack that people describe. i did not sit with a timer and measure each one. if you do time it, the ratio between the two kinds is the number i would most like to see, because my claim that Cs cost more than Fs is an argument from how drift happens, not a measurement.

the rejection-rate arithmetic is illustrative. one in three is a plausible number, not a surveyed one. your rate is your rate, and if it is much better than that, the ten hurts you less than it hurts other people.

and this does not make bad creative good. cutting crossings gets the thing you meant to make in front of people faster and lets you make more of them. if what you meant to make was weak, you have now shipped weak faster, six ways. the count is about friction. it has nothing to say about whether the ad works.

that is the whole method, nothing held back. count your crossings, split them into F and C, go after the Cs, and be suspicious of anything that claims to be end to end without letting the second half read the first half.

if you run the audit, i would like to see your number. reply with the count and which ones were Cs — i am collecting them, and the shape of what people report back is more interesting than anything i can assert on my own.


i'm building the version of this where the whole chain sits on one canvas and the brief never has to be retyped. if you want to see it when it lands: clear-cortex.com/early-access