The Real Cost of AI Video: From One Generated Second to One Usable Second
The cheapest number in AI video is usually the one shown before anyone watches the result. A model may be priced per generated second, but a production buys attempts: the take with the wrong hand movement, the logo that mutates, the camera move that ignores the brief, and then—sometimes—the shot worth keeping. That distinction matters now that the same creative interface can expose models whose list prices differ by thirty times.
This report normalizes a representative set of prices available through the Runway API on 26 August 2026. It is not a quality ranking. It is a way to ask a better budgeting question: after resolution, audio, retries, editing, and review, what does one selected second cost?
The price on the page is not the production cost
Runway’s developer pricing defines one credit as USD 0.01 and bills most video models per generated second. That is unusually useful for comparison because it puts several providers behind one documented unit. It is still only a platform price snapshot, not a universal wholesale price and not the total cost faced by every buyer.
Three layers should stay separate:
- Generated cost is what the API charges for returned media.
- Selected cost divides the total spend across the footage a team accepts for further work.
- Finished cost adds human direction, review, edit, sound, compositing, rights checks, storage, and delivery.
A five-second result billed at USD 0.25 can therefore be cheaper to generate but more expensive to use than a USD 2 result that follows the brief on the first pass. List price is observable; fitness for a particular production is not.
Official Google I/O 2026 Gemini Omni demonstration. It illustrates current multimodal capability, not comparative cost or quality. Source: Google.
Eight prices normalized to ten seconds
The chart and table use the same simple calculation: credits per second × USD 0.01 × ten seconds. They exclude minimum generation lengths, taxes, subscription discounts, failed network requests, and every post-production cost.
| Model and configuration | Credits/second | USD/10 generated seconds |
|---|---|---|
| Gen-4 Turbo | 5 | $0.50 |
| Veo 3.1 Fast, no audio | 10 | $1.00 |
| Gen-4.5 | 12 | $1.20 |
| Veo 3.1 Fast, audio | 15 | $1.50 |
| Veo 3.1, no audio | 20 | $2.00 |
| Veo 3.1, audio | 40 | $4.00 |
| Seedance 2.0, 1080p | 40 | $4.00 |
| Seedance 2.0, 4K | 150 | $15.00 |
The visible spread comes from more than brand positioning. Audio, resolution, speed, model size, and the type of workflow all affect the bill. On this snapshot, adding audio moves Veo 3.1 Fast from ten to fifteen credits per second and regular Veo 3.1 from twenty to forty. Seedance 2.0 moves from forty credits per second at 1080p to 150 at 4K. Those premiums can be rational when they remove a later step—but only if the output actually removes it.
The retry-rate multiplier
No public price page can tell a studio its acceptance rate. The useful approach is to model it explicitly. Suppose a team selects one in five generated clips for an edit. That 20% rate is an editorial scenario, not an industry benchmark. Under it, five ten-second attempts produce fifty paid seconds and one selected ten-second clip.
The multiplier is direct: cost per selected second equals list cost per generated second divided by the selection rate. At 20%, Gen-4 Turbo’s five-cent generated second becomes 25 cents per selected second before labor. Gen-4.5 becomes 60 cents. Veo 3.1 with audio becomes USD 2, and Seedance 2.0 at 4K becomes USD 7.50.
Selection rates also vary inside one production. A textured establishing shot may be easy to accept; a close shot of a speaking person holding a branded product may require exact identity, lip movement, packaging geometry, and legal approval. Reporting a single “cost per video” across those jobs hides the operational difference.
Audio, resolution, and post-production
Native audio is not automatically cheaper than separate audio. It can reduce synchronization work when dialogue, ambience, and action arrive together. It can also make a visually good take unusable when one spoken word is wrong. Teams need to track visual acceptance and audio acceptance separately, then decide whether native audio increases the joint success rate enough to justify its price.
Resolution creates a similar trap. Paying for 4K generation may preserve detail and avoid an upscale, but it may also spend more on attempts that will be discarded. The same Runway pricing page lists Magnific creative upscaling by output frame: its examples for ten seconds at 30 fps are 210 credits at 720p, 270 at 2K, and 360 at 4K. A pipeline can therefore compare “generate every attempt at 4K” with “iterate lower, then upscale only selected footage.” Neither wins in every case.
Post-production is the largest missing column. A model may return frames in minutes while the production still pays for prompting, art direction, continuity review, conforming, color, sound, captions, versioning, and client feedback. The faster generation becomes, the more likely human review—not compute—is the bottleneck.
Latency also has an economic value that a per-second table cannot capture. A cheap model with a long or unpredictable queue can leave an editor idle, miss a media cutoff, or force a team to generate with two providers as insurance. Conversely, the fastest result has little value if it creates more review work. Teams should record queue time, generation time, and human review time separately. That reveals whether a “fast” option is shortening the whole production or merely moving delay downstream.
The same discipline applies to invalid jobs. A network error that is not billed differs from a technically valid output that ignores the brief and is billed. Combining both as “failures” makes provider reliability and creative acceptance impossible to compare. A useful log distinguishes request errors, policy rejections, technical output defects, brief mismatch, and later editorial rejection.
A budget model teams can actually use
A practical estimate needs six inputs rather than one headline price:
| Input | What to record |
|---|---|
| Required selected duration | Seconds that must survive into the edit |
| Average attempt duration | Paid seconds returned per generation |
| Selection rate | Accepted attempts ÷ total valid attempts |
| Model price | Credits or currency per generated second |
| Finish costs | Upscale, edit, audio, localization, storage |
| Human time | Direction, review, rights, approval, delivery |
For each shot class, calculate expected generated seconds as selected seconds divided by the selection rate. Multiply by model price, add machine post-processing, then add labor using the team’s real review time. Keep first-pass rate and final acceptance rate separate: a take may proceed to edit and still be rejected by a client or safety review.
The model also needs a variance range. A campaign budget based only on the average will fail when a hero shot takes twenty attempts. Quoting low, expected, and high cases is more honest than turning an uncertain creative process into a precise-looking single number.
What the comparison cannot tell you
This data does not establish which model is “best.” Models differ in prompt adherence, motion, identity consistency, safety policy, queue time, available controls, and failure modes. An aggregator’s price may differ from a model provider’s direct price. Vendors can change rates after this snapshot. And a production’s definition of usable footage is inseparable from its brief.
The durable conclusion is narrower: generated seconds are an input, not the deliverable. Teams that measure attempts, selection rates, and finish work can compare models economically. Teams that count only the successful clip will systematically understate both the money and the judgment required to make it.
Sources
- Runway API Pricing & Costs — model, audio, resolution, and upscaling prices; accessed 26 August 2026.
- Runway plan pricing — subscription context, kept separate from API unit economics.
- Google I/O 2026 keynote moments — official captioned Gemini Omni demonstration.