/ GuideUpdated 11 Sept 202617 min read

What an AI Video Production Agency Actually Does in 2026, and What It Costs

A clip is not a video. How an AI video production agency turns generations into a finished ad: storyboard first, a locked look, the model picked per shot, editors last, and what it costs against a shoot.

The reel an AI video agency sends over is eight seconds of a woman laughing in a kitchen that does not exist, and it tells you nothing about whether they can make your thirty second ad. Generating the shot is the easy part now. Producing the ad is the job. A script and a storyboard come before any prompt. The same character has to survive six shots. Someone picks the model that can hold her this month, cuts the seams between eight second clips, adds the voice and music, checks whether a real product or face can legally be in it, and labels it for the platform. That is what you pay for, and it is most of the hours.

If the quotes for this make no sense yet, that is the normal state. People coming from traditional production are still asking how anyone charges for AI video, and the answers run from five dollars to fifteen thousand for the same length. This is the version that says what the money buys, and how we spend it on the three problems every practitioner thread complains about: a clip is not a video, consistency falls apart, and the tools change every month.

AI video production means a finished video, and a generated clip is not one

An AI video production agency turns a brief into a deliverable video, using generative models for the visuals and people for everything else. A generator makes shots. An agency makes the video the shots go in.

Where AI sits in an AI video production pipeline Eight stages in two rows of four. Brief, concept and script, storyboard and style frames are done by people. Shot generation is done by a model. Continuity and selects, the edit with voice and sound, rights and disclosure, and revisions and delivery are done by people. Only one of the eight boxes is blue. 1. Brief people 2. Concept, script people 3. Storyboard, style frames people 4. Shot generation model, directed 5. Continuity, selects people 6. Edit, voice, sound, captions people, some tools 7. Rights and disclosure people 8. Revisions, delivery people a model generates it a person does it, and it is most of the hours
The stage people picture when they hear AI video is box four. An agency is paid for the other seven, which is where the finished video gets its script, its continuity and its right to be published. Illustrative shape. Vimerse diagram.

The gap between the two is the complaint that runs through every practitioner thread. A post in r/generativeAI separates generating a clip from making a usable video, and a thread tired of demo-reel rankings adds that a model "can nail one gorgeous four-second clip and still be useless for a real campaign" when the job is a sequence, not one shot. Judge an agency on a finished multi-scene piece, never on its best eight seconds.

Without direction the failure is not bad shots but endless ones. The team behind a Veo 3 ad running on YouTube found that without creative guidelines the prompting never ends, and a VFX artist hired for an all-AI project was handed a client plan that read "Put the brief in Chatgpt, ask it to make a prompt for Veo, and plug that into Veo", with nothing drawn before the prompt. Our AI video production page names the three places that plan fails, and the next three sections are how we run each one differently.

A clip has no script, sound or pacing, so we build the whole video around the visuals

The first problem is that raw generations carry no concept, script, storyboard or pacing, which is everything that sells. We answer it with the order of work: nothing is generated until it has been drawn, and nothing ships until editors have finished it.

The Vimerse AI production pipeline in five steps Five boxes in a row. One, storyboard creation, by people. Two, each shot and its matching audio created in Vimerse Studio, a tool step directed by a producer. Three, the still images converted into video, a tool step. Four, music added, by people. Five, the edit: sound effects, visual effects, refining the cuts and transitions, by editors. Steps two and three are blue, the rest teal. 1. Storyboard shots, script, style frames 2. Each shot and its audio in Studio 3. Stills into video 4. Music 5. Edit SFX, VFX, refining cuts, transitions people tool, directed tool, directed people editors Steps 2 and 3 are the generation. Steps 1, 4 and 5 are the hours you are paying for, and nothing ships from step 5 without sign-off.
Our own pipeline, in the order we run it. Two of the five steps are the tools, and both of them are directed from a storyboard that people drew first. Real process, not an illustration. Vimerse diagram.

Step one is a storyboard: shots, script and style frames, drawn by a person from your objective, or from your script if you have one. Step two creates each shot and its matching audio in Vimerse Studio, our own tool, directed shot by shot from that storyboard. Step three converts the stills into video. Step four adds music. Step five is the edit: voiceover, sound effects, visual effects, captions, refining the cuts and transitions, done by editors, and nothing ships from it without your sign-off.

Stills first is what the people doing this weekly recommend. A Veo user whose 2.5 minute commercial took 7 hours cut errors by fixing the character as an image before any video ran, and a marketer making corporate training videos with AI plans every scene in a storyboard view before touching the generation tools. The two orange boxes in the figure are the generation. The three pale ones are people's hours, which is why an AI project runs on our flat $18 an hour rather than a per-clip price. A pipeline that starts at the prompt is the daisy chain that the thread on character drift names as the first trap, so ask any agency what it draws before it prompts.

Consistency falls apart shot to shot, so we lock your look before anything is generated

The second problem is drift: characters, environments and the brand look change from shot to shot without continuity work. Every limit on the right of the figure is a task a person does before or after the model, and the left column is as far as the 2026 tools get on their own.

What AI video generators can and cannot do reliably in 2026 Two columns. Reliable: an 8 second clip with sound, one character held for one scene from a reference image, lip sync in Veo 3 and Kling 3, and 3 to 4 generations per usable clip. Not reliable: one character held across a campaign, a 30 second take in one generation, readable text, logos and exact products, and fixing one moment without regenerating the whole clip. Does reliably Does not, yet An 8 second clip with sound One character, one scene, from a reference image Lip sync: Veo 3, Kling 3 A usable clip in 3 to 4 tries (5 to 6 with two characters) One character across a campaign A 30 second take in one go: past 10 to 15 s it hallucinates Text, logos, exact products Fixing one moment without regenerating the whole clip Everything on the right is done in the edit or by a person.
Each cell is a complaint or a recommendation from the threads linked in this section, as of September 2026. The right column is the agency's job description. Vimerse diagram.

On consistency, the r/artificial thread asking whether anyone has cracked it gets as far as a scene or a short sequence by locking a reference face or voice, and no further. The practitioners who get further all lock the same things earlier. A thread on consistent realistic human characters finds that the "Biggest win for consistency is locking the character first, not the video", with a small character bible of fixed face references, and a thread on why consistency is still the biggest problem lists the same toolkit: reference images, character models locked between scenes, and real footage mixed in where the model cannot hold something. That is our step two. Style frames, your brand kit and a character reference per person and place are approved before a single video generation runs, and they are kept, so your next video matches this one.

Length is the other half of drift. Veo 3 only produces eight second clips, stitching leaves a visible skip between clips, and a thread on making long AI videos hold together warns that models start to hallucinate past ten to fifteen seconds in one go, so a person cuts the joins. Text and products stay off limits, and the advice on AI B-roll for marketing videos is to keep product shots real, since models still struggle with text, logos and specific colours. Then the retry rate. Asked how many generations one usable clip takes, a Veo user reports an "average 3-4 gens before something usable", five or six with two characters. Locking the look first is what pulls that ratio down, which is why it comes before generation in our pipeline and not after. Plan every video as shots under eight seconds with the references fixed, because that is what it will be.

The tools change every month, so we pick the model per shot and absorb the churn

The third problem is churn. Models improve, break and get replaced monthly, and picking, testing and re-learning them is a full-time job. We stay model agnostic: the storyboard, style frames, character references and voice are ours and fixed, and the model behind step three is whichever one holds that look best for that shot this month.

Five model verdicts in posting order, against what stays fixed in our pipeline Left, five forum verdicts in posting order: Veo 3 is the best on the market by far; Kling 3.0 just changed the game, rip Veo; Kling is probably the best for character consistency across scenes; no more Sora, OpenAI winds down its video products; goodbye Veo 3.1, outdated tech now. Right, what does not change: the storyboard, the style frames and character references, the voice, and the edit. The model at step three is whichever holds the look best this month. What the forums said, in order 1.Veo 3 is the best on the market by far 2.Kling 3.0 just changed the game, rip Veo 3.Kling: the best for a character across scenes 4.No more Sora: video products wound down 5.Goodbye Veo 3.1: outdated tech now Each verdict overturns the one before it. An agency built on one model lived every line. What we keep fixed The storyboard Style frames, character refs The voice The edit, by editors Step 3: whichever model holds the look this month
Five real thread titles and verdicts from the forums linked in this section, in the order they were posted, against the four parts of our pipeline that a model change does not touch. Vimerse diagram.

The figure is five verdicts from the forums in the order they were posted, and each one overturns the last. A reply to a nine-model comparison in r/StableDiffusion called Veo 3 the best on the market by far. A later thread titled Veo 3 has gone downhill answers itself with "kling 3.0 just changed the game rip veo". Then a thread asking what people actually use after Sora died follows a post reporting that OpenAI would wind down its video products, and a Veo subscriber cancelling at the end of the month calls Veo 3.1 outdated tech. An agency built on one of those tools inherited each of those months.

Per shot, the picks are consistent even while the rankings move. A creator who tested every major model for three months for a client ended up using Sora mostly for wide environmental shots and big landscapes and the others elsewhere. Threads on which model handles story scenes and which tool keeps a character consistent put Kling first for a character across scenes and Veo first when the audio has to be baked in, and a marketer comparing models on lip sync rates Veo 3 best there, with morphing its one remaining problem, while Kling 3 now ships native lip sync. So one video of ours may run three models: one for the establishing shot, one for the character, one for the line of dialogue, all from the same style frames. Ask an agency which models it ran on its last three jobs. If the answer is one name, ask what happens to your campaign when that name has a bad quarter.

AI creative performs when editors finish it, and audiences punish the lazy version, not the tool

Buyers ask two questions together: does it perform, and do people mind. The threads answer both with the same split, between fully generated creative and creative that a person directed and an editor finished.

Fully AI-generated ads against AI-assisted ads in one $92,000 test Three pairs of bars. Click-through rate: full AI 0.9 percent, assisted 2.3 percent. Conversion rate: full AI 1.2 percent, assisted 3.1 percent. Return on ad spend: full AI 2.1, assisted 4.7. The assisted bar is more than double in every pair. Click-through rate 0.9% full AI 2.3% assisted Conversion rate 1.2% full AI 3.1% assisted Return on ad spend 2.1 full AI, near break-even 4.7 assisted fully generated, prompt to publish generated, then directed and edited by people
Real figures from the r/advertising test linked beside this, one advertiser's accounts. Each pair is drawn to its own scale. The gap is the human pass, which is the thing the word agency refers to. Vimerse diagram.

The cleanest data is a media buyer's $92,000 split test, with full AI at 0.9% CTR, 1.2% CVR and ROAS 2.1 against assisted at 2.3% CTR, 3.1% CVR and ROAS 4.7. A Facebook advertiser reports the opposite headline, AI ads around 17% better on CTR across the accounts, and a PPC manager finds three in ten AI images beating the designed ones. Read together they agree: AI creative wins as one of many variants a person judged and loses as the only one nobody did. Make the ten versions you could not afford to film, and let the spend pick.

The gap in the test is the human pass, and that pass is an edit. Editors asked whether AI will replace them answer that even if editing in the old sense goes, you will still need editors, and a thread tired of every AI tool calling itself an editor makes the same distinction from the other side. This is the part of the job we did before there were generators. Vimerse has edited for creators and agencies since 2021, with a team of more than fifty editors and more than five hundred clients, and every AI visual we deliver goes through them for pacing, sound design, colour and story before you see it. That discipline is the difference between a produced video and a string of clips, and it is the $18 an hour you are paying for.

On whether people mind, the forums split and we take a side. One camp points at the dislikes under Coca-Cola's AI Christmas ad and warns that younger audiences read an AI ad as the business being lazy. The other camp, in an r/advertising thread asking for opinions on AI ads, holds that "People don't hate AI ads, they hate boring or deceptive ads", and a thread on AI video becoming indistinguishable calls it a reaction to the tool, not the result. Across the social work we deliver, audiences care about the idea, not how it was made, and every backlash case is a flagship brand moment where the AI was the story. Use AI for volume and tests, not for the one ad a brand is known for.

An AI production house charges $300 to $500 a video, a shoot $3,000 to $15,000, and in-house costs whatever you throw away

There are four ways to buy a short AI video: credits and your own time, a freelancer who prompts, an AI production house, or a shoot. They are priced in different units, so the figure puts them on one scale.

What one short AI video costs, by who makes it Seven horizontal range bars on a logarithmic dollar scale from 1 to 20,000. Doing it in-house on credits alone: 5 to 52 dollars. In-house counting failed generations: 450 to 750. An AI freelancer: 50 to 500. An AI production house: 300 to 500. A UGC creator: 150 to 800. A traditional agency shoot: 3,000 to 15,000. Vimerse: 36 to 72 dollars of production time, marked with a dashed line, before generation credits, which are added as one total. In-house, credits only $5 to $52 In-house, with failed gens $450 to $750 usable AI freelancer, 30 s $50 to $500 AI production house $300 to $500, 30 to 60 s UGC creator, filmed $150 to $800 Traditional agency shoot $3,000 to $15,000 us: $36 to $72 of time at $18/h, plus credits as one total $1 $10 $100 $1,000 $10,000 log scale: each tick is ten times the last
Real figures from the threads linked in this section, for one short video of about 30 to 60 seconds. The two in-house bars are the same person before and after counting the generations that failed. Our line is production time only; generation credits go on the quote as one total on top. Vimerse diagram.

The bottom rung is illustration. One freelancer in r/VEO3 pricing a thirty second job argues that "$50 per 30s is fair for AI video illustration based on a set plot, where no creative ideation is required", and the ideation is what the next rung charges for. The middle is the production house. An agency owner in r/digital_marketing selling AI video ads to other agencies is currently making 30 to 60 second videos for about $400 to $500, up from $300 to $400 excluding AI costs. The top is the shoot, which a media buyer comparing what converts puts at $3K to $15K per video, and a videographer reviewing a pricing plan confirms at a $7k minimum for a shoot with post production. Decide which rung you are buying before you compare two quotes.

In-house is the rung people misprice, because credits are cheap and retries are not. Vertex AI lists Veo 3.1 at $0.40 with audio and Veo 3.1 Fast at $0.10, a Google Cloud user reads the same table as $0.50/second for Veo 2, and a thread asking how freelancers produce videos with AI has them charging $500 for a single video on a tool cost of probably under $5. Then the retries, which a prompt engineer who cut their costs by 80% sums up as "Factor in 3-5 failed generations = $450-750 per usable video", a line no pricing page shows. The retries are the cost, plus the hours of the person running them.

Take a real case. A founder who burned $2,400 in three weeks posted the arithmetic: a typical four generation average made it $600 per 5-minute video, a first month of $2,900, and after reworking the process $580 for the same output.

One month of AI video, before and after fixing the process Two stacked bars. The first month: 2,900 dollars, made of 1,500 main content, 900 platform variations, 300 pickup shots and 200 not itemised. After the rework: 580 dollars, made of 300 main content, 180 variations, 60 pickups and 40 research. The second bar is a fifth of the first. First month main content $1,500 variations $900 pickups $300, not itemised $200 $2,900 After rework $580: main $300, variations $180, pickups $60, research $40 $0 $1,000 $2,000 $3,000 Same tools both months. The saving came from planning shots before generating them.
Real figures from the r/reactnative post linked beside this. The cost of AI video is set by how many generations you throw away, and that is a process problem before it is a tool problem. Vimerse diagram.

The same tools cost five times less once shots were planned before they were generated, which is steps one and two of our pipeline. Our rate on the AI video production page is a flat $18 per production hour, the first video is free up to four hours, enough for one or two short AI ads, and paid generation goes on the quote as one total before work starts, separate from production time. We do not break that total down by shot, and few agencies will, but the total itself should be there, and nothing on the quote is billed without your approval. An agency that will not show the generation line is charging you for its retries.

A short AI video takes one to sixteen hours of work, and a project takes days, not minutes

Turnaround has two clocks: the hours a person spends, and the calendar days from brief to delivery. Vendors advertise the seconds a generation takes, which is neither.

Hours of work per AI video, and calendar days per project Six horizontal bars on a scale from zero to twenty hours. A 30 second animated video, 1 to 2 hours. One product video, 1 to 2 hours. A company ad, under 2 hours. A 2.5 minute commercial, 7 hours. A 4 minute video, 10 hours minimum. A 2 minute video, 16 hours. Below, a calendar row: a traditional ad 3 to 4 weeks, a Vimerse project about 4 days. 30 s animated AI video 1 to 2 h One product video 1 to 2 h A company ad in Veo 3 under 2 h 2.5 min commercial 7 h 4 min video 10 h minimum, more with lip sync 2 min video 16 h 0 5 10 15 20 h hours of work Calendar time traditional ad: 3 to 4 weeks us: about 4 days, brief to delivery
Real hours from the threads linked in this section, by people who make AI video for a living. The spread is not the tool. It is the number of shots and whether anyone has to speak. The calendar row compares a traditional ad's schedule with our typical project. Vimerse diagram.

People who make AI video weekly report hours. In r/aitubers one creator, in a thread on how long an AI video takes, reports one to two hours for thirty seconds of animated AI video, and another, asked about time spent per video, needs ten hours minimum for four minutes, longer with lip sync. A founder made a company ad in under two hours with Veo 3, and the top reply said it was not a great use case. Budget hours by shot count and by whether anyone speaks, not by length.

On the calendar the comparison is with a shoot, which a marketer comparing manual and AI ad production puts at three to four weeks minimum when it is done properly. Most of our AI projects are delivered in about four days, and a larger multi-scene production gets its own timeline in the estimate. Four days is not the generation. It is the storyboard, the selects, the edit and your named revision rounds, and same-day delivery skips one of them.

You own the finished video, nobody owns the raw generation, and paid ads now need a label

Two questions get merged here: who holds the rights to the deliverable, and where a disclosure is required. Both moved in the last eighteen months.

In the US a purely machine-made output has no author, and a thread on whether AI output is free for commercial use notes that the US copyright office has turned down registrations for AI authorship, a reading Kling users share. What an agency can give you is the human-made work around the generation, the script, the edit and the finished film, plus a contract that assigns whatever rights exist and warrants that no third-party person or brand was reproduced. Our terms are that you own the final videos and their commercial-use rights, your files stay in access-controlled workspaces, and a real employee or spokesperson appears only with consent and rights documented. Ask for both clauses in writing.

On disclosure the rules are specific. YouTube requires a label on realistic content generated or meaningfully altered with AI, and not on productivity uses. On TikTok, since 21 July, a disclosure label is required on all AI-generated content in paid ads, which a small-business owner tracking the change reduces to "Organic is encouraged but not required. Paid is mandatory", the line to plan around. The EU AI Act's Article 50 took effect on 2 August. Tell the agency where the video will run, and ask who applies the label; we advise on it where required.

An AI production is the wrong choice when the real thing is the point, and we say so in the estimate

Three questions sort almost every brief, and an agency selling AI production should ask them before it quotes. We ask them in the order below, and when the answer is a shoot or a hybrid, that is what the estimate says.

When an AI production is the wrong choice A decision path with three questions. Must a real thing be exactly right, such as a product, a procedure or an executive's face? Yes leads to a traditional shoot or a hybrid with your footage. No leads to the next question. Must the audience believe it actually happened, such as an event or a testimonial? Yes leads to shoot it. No leads to the last question. Is it a concept test, a variant set, an explainer, training or a social ad? Yes leads to AI production. Must a real thing be exactly right? a product, a procedure, an executive's face yes: shoot it, or hybrid with your footage no Must the audience believe it happened? an event, a testimonial, a documentary claim yes: shoot it AI recreates, it cannot record no Is it a test, a variant set, an explainer, training, a concept or a social ad? yes: AI production this is most of the work Two yes answers in the first two boxes and the AI budget is better spent on a camera. An honest agency says so in the estimate.
The three questions we ask before quoting. The order matters: authenticity questions come before cost questions, because a cheap video that cannot be published costs the whole budget. Vimerse diagram.

The first is accuracy. AI cannot show your exact product, and the workaround e-commerce teams actually use is a hybrid: real product photos for accuracy, then AI to expand variations or add usage context. A 3D artist in r/vfx was called in for a product commercial the client could not make with AI. Our own list of poor fits: exact real-product demonstrations, precise technical or medical procedures, executive testimonials, and work where real-world authenticity is essential. For those we produce a hybrid, blending AI with your footage, stock and motion graphics, or we tell you to book the crew.

The second is belief. Videographers planning how to compete agree: AI can recreate an event, it cannot document one, so anything an audience must believe happened stays on camera. The third is scale, which cuts the other way. A videographer who lost three retainer clients checked their pages and found them using AI outright for weekly social output a shoot could never price. If the brief is volume, tests, explainers or training, that is what AI production is for, and variant sets across ratios, hooks and messages are most of what we make. If it is the one film the brand is remembered by, book the crew.

Brief an AI agency like a production, and judge it on continuity, not on a reel

The brief that gets an accurate quote is one a traditional producer would recognise, with two additions: what must be real, and where it will run. You can also send only the objective; concept, script and storyboard are then ours to draft for your approval.

The one-page brief for an AI video production Two columns. What you supply: the objective and audience, platform and ratios, length, what must be real, brand kit or style frames, a script or a request for one, revision rounds, disclosure needs, and the deadline. What you ask for: the generation cost as its own total, who owns the output, a multi-scene sample, the timeline, and named revision rounds. You supply You ask for Objective and audience Platform, ratios, length What must be real Brand kit or style frames A script, or a request for one Revision rounds you expect Where it runs, for disclosure The deadline Generation cost, as its own total Who owns the final video A multi-scene sample, not a reel of single shots A timeline with a date Named revision rounds Who directs the generations What they will refuse to make A brief that fits this card gets a quote you can hold them to.
Everything a quote needs and everything a quote should answer, on one card. The right column is the set of questions that separates a production agency from a prompter with a portfolio. Vimerse diagram.

Judge the portfolio on sequences. A prompt engineer reviewing a big generation count notes that "Ten months and 10,000 generations sounds impressive until you realize that's 33 videos per day", a count of attempts, not of finished work. A producer shows one finished piece where the same character walks through five shots, the product is real, and the cuts hide the eight second joins. Editors getting requests to generate footage for commercials do that hiding, so ask who edits, not only who prompts.

Judge the quote on what it includes. Ours lists scope, concept direction, deliverables, formats, revision rounds and timeline, with generation as its own total. A Premiere editor names the misconception to brief against, that it is somehow immediate, and a videographer asking why AI client projects cost so much was told clips cost under a dollar, true of the clip, not the project. Write the brief on the card above, send it to two agencies, and compare line items, not totals.

What to do with each thing you see

What you seeWhat it meansWhat to do next
A showreel of gorgeous single shotsThey can generate. Nobody has shown you they can produceAsk for one finished multi-scene video with the same character in every shot.
A pipeline that starts at the promptNo storyboard, so every shot is a fresh roll of the diceAsk what they draw before they generate. Our five steps start with a storyboard.
A character who drifts between scenesText-to-video shots with no reference frameAsk for style frames and character references locked before generation. Ours are, at step two.
An agency built on one modelWhen that model is throttled, downgraded or wound down, so is your pipelineAsk which models they run and what changed last quarter. Ours is picked per shot.
No editor named in the pipelineThe generations ship as they came outAsk who cuts, mixes and grades. Every AI visual we deliver goes through editors first.
A quote with no line for generation costsEither credits are absorbed in a high rate or they will appear laterAsk for the generation total as its own line. Ours is on the quote before work starts.
Readable text or an exact product in an AI shotSomeone comped it in, or it is wrongKeep logos, text and product hero shots real or graphic. Do not ask the model for them.
A promise of a 30 second take in one generationThey have not shipped oneExpect 8 second clips cut together, and budget the edit that hides the joins.
A fully generated ad with no human passThe 0.9% CTR bar in the test aboveDirect and edit the generations before spend goes behind them.
A per-video price under $50Illustration from a set plot, no ideationFine for a storyboard-in-motion. Not for a campaign.
A per-video price of $300 to $500The going rate for a 30 to 60 second AI ad from a production houseAsk what the number includes: script, voice, revisions, variants.
A quote of $3,000 to $15,000A traditional shootRight when a real product, a real face or a real event has to be in it.
Nothing about disclosureThey have not run ads on TikTok or the EU since the labels went mandatoryAsk where the video will run and who applies the label.
A claim to own or transfer copyright in the raw generationNobody can promise that in the USAsk for the contract to assign whatever rights exist and to warrant no third-party likeness or brand.

Things to stop doing

  • Judging an agency by a reel of single shots. Ask for one finished multi-scene video.
  • Letting anyone generate before the storyboard, style frames and character references are approved.
  • Building a campaign on one model. Ask what the agency ran last quarter and what it switched.
  • Asking a model for text, logos or your exact product. Comp them in or film them.
  • Comparing a $50 illustration quote with a $5,000 shoot quote. They are different rungs.
  • Running fully generated creative as the only creative. Make it one variant of ten, and put an editor on it.

The brief, as a checklist

  • The objective in one sentence, who it is for, and the platform, ratios and length.
  • What must be real: product, face, place, procedure. If the list is long, ask for hybrid.
  • Brand kit or style frames, approved before generation.
  • Your script, or a request for one to approve.
  • Revision rounds, named in the quote, and where the video runs, so disclosure is planned.
  • The deadline, and the date the first cut is needed.
  • Ask for: the generation total on its own line, ownership in writing, a multi-scene sample, the models they ran last, who edits, a timeline, and what they will refuse to make.

The short version

  • A clip is one of eight stages. The agency is paid for the other seven.
  • Our pipeline: storyboard, shots and audio in Studio, stills into video, music, edit. The tools do the middle two, and nothing ships without sign-off.
  • The look is locked before generation: style frames, brand kit and a character reference per person and place, kept for the next video.
  • The model is picked per shot and changes month to month. The storyboard, references and edit do not.
  • Every AI visual goes through editors for pacing, sound, colour and story. That pass is the gap in the $92,000 test.
  • A production house charges $300 to $500 a short video, a freelancer $50 to $500, a shoot $3,000 to $15,000, in-house $450 to $750 per usable video with retries counted. Ours is $18 an hour plus generation as one total.
  • A short AI video is one to sixteen hours of work. A project is days. Ours is about four.
  • You own the finished work by contract. Nobody owns the raw generation. Paid AI content is labelled on TikTok and in the EU.
  • AI is wrong when the real thing is the point: exact products, procedures, executive faces, events people must believe. We say so in the estimate and offer hybrid.
  • Brief it like a production, judge it on a five-shot sequence, compare quotes on line items.

Reading the quote is quick. Producing the video is the part that does not fit around a job. You bring the objective and we build the video: your first AI video and up to 4 hours of production time are on us, no card required.

Last reviewed 11 September 2026. Every price, hour count, model verdict and test result in this guide is one a practitioner posted on Reddit in 2025 or 2026, quoted from the thread linked beside it; the Veo prices are from Google's Vertex AI pricing page on the review date; the Vimerse rates, turnaround, terms, team size and pipeline are from our own AI video production page.

First video free, up to 4 editing hoursBook a 15-min call