The Complete AI Video Model Comparison: Runway, Veo, Gemini Omni, Kling, Seedance, Wan, and LTX-2.3
Verified specs, live Artificial Analysis rankings, and per-second pricing for every current major AI video model — updated July 18, 2026.
Seven models, seven different bets on what AI video should be. Abstract editorial illustration by RizzGen.
The AI video field moved fast between February and July 2026. Runway Gen-4.5 launched claiming the top benchmark spot, then quietly disappeared from the live leaderboard as newer entrants arrived. Google folded Veo into a broader multimodal model called Gemini Omni. ByteDance shipped native 4K on Seedance 2.0, then followed it with Seedance 2.5 before most people had finished testing 2.0. Alibaba's Wan and Kuaishou's Kling both jumped a full version. OpenAI's Sora, once the center of this conversation, is being wound down — the consumer app is already off, and the API follows on September 24, 2026 — so we've dropped it from this comparison entirely rather than treat a discontinued product as a live option.
Every number below is sourced to an official vendor page, official API documentation, or the live Artificial Analysis Text-to-Video Arena — not to a marketing tweet repeated across a dozen SEO blogs. Where sources disagree or a model is too new to have public data, we say so instead of guessing.
Interactive Comparison
Filter the table by capability, or click any column header to sort.
| Model↕ | Developer↕ | Max Duration↕ | Resolution↕ | Native Audio↕ | Starting $/sec↕ | Open Source↕ | AA Elo (T2V)↕ |
|---|---|---|---|---|---|---|---|
| Runway Gen-4.5 | Runway | 10s (5s/10s presets) | 1080p † | Yes | $0.12–0.25† | No | Not ranked‡ |
| Google Veo 3.1 | Google DeepMind | 8s base (chainable to 140s+) | Up to 4K (base clip); 720p when chained | Yes | $0.15–0.40 | No | 1,094 |
| Gemini Omni Flash New | 10s | Not officially published† | Yes | $0.10 | No | 1,240 #1 | |
| LTX-2.3 New | Lightricks | 20s | Up to 4K / 50fps | Yes | $0.06–0.32 | Yes | 974 |
| Wan 2.7 New | Alibaba | 15s | 1080p | Not documented† | $0.10 | Yes (Apache 2.0) | 1,160 |
| Seedance 2.0 4K added Jun 2026 | ByteDance | 15s (10s at 4K tier) | Native 4K (3840×2160)† | Yes | $0.10–0.15 base† | No | 1,225 #2 |
| Seedance 2.5 Rolling out | ByteDance | 30s | 4K † | Yes | ~$0.14† (est.) | No | Too new to rank† |
| Kling 3.0 4K/60fps | Kuaishou | 15s (Avatar mode: 5 min, separate feature) | Native 4K / 60fps | Yes | $0.084–0.168 | No | 1,110 |
| † = figure is a range, estimate, or officially undocumented — see inline notes and the Sources section. ‡ = see the Runway section below for the full leaderboard-delisting context. AA Elo = Artificial Analysis Text-to-Video Arena, checked July 18, 2026; scores for the same model differ slightly on the Image-to-Video board. | |||||||
Compare Two Models Head-to-Head
Pick any two models for a direct spec-by-spec comparison. Lower price and higher Elo are highlighted — everything else is a trade-off, not a win.
Artificial Analysis Text-to-Video Arena — Elo Score
Live blind-comparison leaderboard, checked July 18, 2026. Higher is better. Source: artificialanalysis.ai/video/leaderboard/text-to-video
The hatched bar is not directly comparable to the others: it's Runway's own self-reported figure from its December 1, 2025 launch announcement. Runway Gen-4.5 does not currently appear anywhere on Artificial Analysis's live Text-to-Video or Image-to-Video leaderboards — see the Runway section below. Seedance 2.5 is excluded because it launched too recently (July 16, 2026) to have arena data yet.
Starting Price per Second
Lowest publicly listed per-second rate for each model. Most scale up substantially at higher resolution tiers (Seedance 2.0's 4K tier runs roughly 5x its base rate; LTX-2.3's 4K/pro tier runs roughly 5x its fast/1080p rate).
Seedance 2.5's rate is an early estimate — ByteDance had not finalized public API pricing for it as of July 18, 2026.
1. Runway Gen-4.5 (December 2025)
Runway released Gen-4.5 on December 1, 2025 and announced it had scored 1,247 Elo on the Artificial Analysis Text-to-Video benchmark, claiming the top spot at the time. That claim is the source of the "#1 AI video model" line you'll see repeated across dozens of comparison articles, including the earlier version of this one.
It's no longer verifiable as current. As of July 18, 2026, Runway Gen-4.5 does not appear anywhere on Artificial Analysis's live Text-to-Video Arena or Image-to-Video Arena leaderboards — not at #1, not in the top 26 rows we checked. We can't tell you whether that's because Runway stopped submitting the model for evaluation, the listing was retired, or something else; we can only tell you what the live board shows today. Treat the 1,247 figure as a historical, self-reported claim from launch day, not a current ranking.
Key Specifications
- Max Duration: Runway's API reference accepts 2–10 seconds; 5s and 10s are the commonly offered presets
- Resolution: 1080p is what integrators consistently report; Runway has not published an official resolution spec sheet for Gen-4.5
- Native Audio: Added in the December 2025 update — synchronized dialogue, ambient sound, and music
- Pricing: Runway's own docs list two different per-second rates in different places (12 and 25 credits/second, i.e. roughly $0.12–$0.25/s) — likely tier- or resolution-dependent; we couldn't find a single official number that reconciles both
- Architecture: Diffusion transformer, trained and served on NVIDIA Hopper/Blackwell GPUs
Strengths
Character Consistency: Still widely regarded as best-in-class for maintaining character appearance across scenes without fine-tuning, using reference images to lock subjects across lighting and camera changes.
Prompt Adherence: Strong instruction-following on camera movement language ("dolly zoom," "whip pan," "tracking shot").
Limitations
- No longer independently benchmarked on the leaderboard it built its launch marketing around
- Runway itself acknowledges causal reasoning errors and object permanence issues in complex scenes
- Thin official documentation on exact resolution and duration limits, which makes it harder to plan production pipelines around
Best For
Teams that already value Runway's character-locking workflow and don't need the highest current benchmark score to justify the choice. If a top Artificial Analysis ranking specifically is the deciding factor, Gemini Omni Flash and Seedance 2.0 are the models that currently hold it.
2. Google: Veo 3.1 and Gemini Omni Flash
Google now ships two distinct video-generation products, and conflating them is the most common mistake in other 2026 comparison articles. Veo 3.1 is Google's specialized cinematic video model, released October 14, 2025, and it remains the recommended "video model baseline" in the official Gemini API and Vertex AI documentation — its model ID (veo-3.1-generate-preview) is still live. Gemini Omni Flash is a newer, separate multimodal model announced May 19, 2026 at Google I/O that replaced Veo inside the consumer Gemini app, but has not replaced Veo 3.1 in the developer API, where Google says broader access is still rolling out.
Veo 3.1
- Max Duration: 4, 6, or 8 second base clips; scene-extension chaining links up to 20 clips for sequences past 140 seconds (extended sequences cap at 720p, not 4K)
- Resolution: Up to 4K for a single 8-second base clip; 720p/1080p at 24fps for the general-availability tier
- Native Audio: Yes — dialogue, sound effects, ambient audio
- Pricing: $0.15/s (Fast) to $0.40/s (Standard)
- Artificial Analysis Elo: 1,094 (text-to-video), 1,088 (image-to-video)
Gemini Omni Flash Announced May 2026
Omni unifies text, image, and video generation with conversational, multi-turn editing — you describe a change in plain language ("shift the camera left," "make it ripple like liquid") and it reworks that one element while holding the rest of the scene intact. It currently leads both Artificial Analysis boards: 1,240 Elo (#1) on text-to-video, 1,204 Elo (#1) on image-to-video — the same category where Seedance 2.0 sits at #2 on both. That's the direct head-to-head worth knowing: independent testers have reported Seedance 2.0's raw frame quality as a tier above Omni Flash's, but Omni wins the Arena's blind vote overall, and its differentiator is the multi-turn editing loop rather than single-shot visual fidelity.
- Max Duration: ~10 seconds at initial Flash-tier rollout
- Pricing: $6.00/min (≈ $0.10/s), per Artificial Analysis's listed rate
- Availability: Live in the Gemini app and rolling into YouTube Shorts; the developer/Vertex AI API was not yet publicly open as of the I/O announcement — Google described it as coming "in the coming weeks"
Best For
Choose Veo 3.1 if you need a stable, documented API with enterprise/Vertex AI deployment today. Choose Gemini Omni Flash if conversational, iterative editing matters more to your workflow than topping a raw-quality bar — once its API is broadly available.
3. LTX-2.3 (March 2026)
Lightricks' LTX-2.3 is the current open-source flagship, a 22-billion-parameter diffusion transformer that generates synchronized audio and video in a single forward pass.
Key Specifications
- Max Duration: Up to 20 seconds, single pass
- Resolution: Up to 4K at 50fps
- Native Audio: Yes, generated simultaneously with video; this version improved audio quality via filtered training data and an upgraded vocoder
- License: Open weights, four checkpoint variants (dev, distilled, fast, pro) — the distilled variant runs in as few as 8 denoising steps
- Hardware: Runs locally via LTX Desktop on consumer GPUs, or via the Lightricks API
- Pricing: $0.06/s (1080p fast) to $0.32/s (4K pro) on Lightricks' own API, effective April 1, 2026; comparable third-party hosts (e.g. fal.ai) price similarly by resolution tier
- Artificial Analysis Elo: 974 (Fast, text-to-video) — the highest-ranked open-source model on the board; 958 (Pro, image-to-video)
Strengths
True open weights: Full model access, fine-tunable, deployable on-premise with no API dependency.
Only model with native 4K as a default capability across its full duration range, not a separately-priced add-on tier bolted onto a lower-res base model.
Limitations
- Ranks below every proprietary model on the Arena leaderboard on raw quality, though it's the clear open-source leader
- Self-hosting the full 22B model requires real GPU budget even though the API is cheap
Best For
Developers and studios that need on-premise deployment, fine-tuning, or the cheapest 4K output on the market, and can tolerate a quality gap versus the top proprietary models.
4. Wan 2.7 (April 2026)
Alibaba's Tongyi Lab shipped the full Wan 2.7 suite between April 1–6, 2026 — four models covering text-to-video, image-to-video, reference-to-video (with voice cloning), and instruction-based editing, all under Apache 2.0.
Key Specifications
- Max Duration: 2–15 seconds for text/image generation; 2–10 seconds for reference and editing workflows
- Resolution: 720p or 1080p at 30fps
- Native Audio: Not confirmed in Wan 2.7's public documentation. Its predecessor, Wan 2.6, supported audio-visual sync — if that carried forward, Alibaba hasn't stated it plainly in the materials we could find, so we're not asserting it here
- Key new feature: First-and-last-frame control (specify both ends, the model fills the middle) and natural-language video editing across character actions, dialogue, style, and camera technique
- License: Apache 2.0, API from $0.10/second
- Artificial Analysis Elo: 1,160 (best-ranked variant, text-to-video) — third overall on the live board; 1,098 (image-to-video)
Strengths
Genuinely open and genuinely competitive — it's the only Apache-licensed model in the current top five on the Arena text-to-video board, ahead of several proprietary models.
First/last-frame control is a precision tool most proprietary models don't expose directly.
Limitations
- Caps at 1080p — no 4K tier
- Audio capability unclear from public documentation
Best For
Developers who want an open, self-hostable model that's actually competitive on the live leaderboard rather than trailing far behind, especially for precise start/end-frame control.
5. ByteDance: Seedance 2.0 and Seedance 2.5
Seedance 2.0 Native 4K added June 2026
ByteDance released Seedance 2.0 in February 2026, then added a native 4K tier on June 23, 2026 — true 3840×2160 at 10-bit color, not an upscale. It's currently the #2 model on both Artificial Analysis boards, directly behind Gemini Omni Flash, and independent testers have rated its raw frame quality above Omni Flash's despite the Elo gap.
- Max Duration: 15 seconds at standard resolution; the 4K tier caps at roughly 10 seconds per pass
- Resolution: Native 4K (new tier) or 2K/1080p (base tier)
- Native Audio: Yes, generated in the same pass as video
- Multimodal Input: Up to 9 reference images, 3 video clips, 3 audio files, plus text
- Pricing: Roughly $0.10–$0.15/s at the base tier; the 4K tier runs about 5x that (~$0.50/s — a 5-second 4K clip runs roughly $2.50)
- Availability note: Volcano Engine has published API pricing, but the standalone API rollout has reportedly faced delays over rights-clearance and compliance review. The model is broadly accessible today through ByteDance's consumer apps (Dreamina, Jimeng)
- Artificial Analysis Elo: 1,225 (#2, text-to-video), 1,197 (#2, image-to-video)
Seedance 2.5 Rolling out from July 16, 2026
ByteDance announced Seedance 2.5 on June 23, 2026 (the same event as Seedance 2.0's 4K upgrade) and began rolling it out on July 16, 2026 — first through the Jimeng app and Dreamina/CapCut in China, with the BytePlus API and wider international access following through July 2026. As of this update, it is not yet in full global release, and it's too new to have Artificial Analysis Arena data.
- Max Duration: Up to 30 seconds native — the longest single-pass generation of any model in this comparison
- Multimodal references: Up to 50 inputs, versus 12 on Seedance 2.0
- Local editing: Modify a specific region of a clip (e.g. hair color) without regenerating the full 30 seconds
- Pricing: Not finalized publicly as of July 18, 2026; early third-party estimates put it around $0.14/s at standard resolution, climbing sharply at 4K — treat these as estimates, not confirmed rates
Best For
Seedance 2.0 for the best currently-available combination of raw visual quality, native 4K, and multimodal reference control. Seedance 2.5 is worth watching rather than building on yet — confirm regional availability and finalized pricing before committing a production pipeline to it.
6. Kling 3.0 (February 2026)
Kuaishou announced the Kling AI 3.0 family — Video 3.0, Video 3.0 Omni, Image 3.0, and Image 3.0 Omni — on February 4, 2026. It's a distinct, later release from "Avatar 2.0," which Kuaishou shipped in November 2025 specifically for 5-minute continuous avatar generation; the two are often conflated in other comparison articles, but they're separate features with separate duration limits.
Key Specifications
- Max Duration: 15 seconds for general video generation (up from 10s on the prior version). The separate Avatar mode, introduced in November 2025, supports up to 5 continuous minutes for talking-head/avatar content specifically — not the general text-to-video path
- Resolution: Native 4K at up to 60fps (up from 1080p/48fps)
- Native Audio: Yes, on the base model — Kuaishou's launch materials list native audio generation in English, Chinese, Japanese, Korean, and Spanish, including regional accents and dialects, plus multi-character dialogue scenes where each character can speak a different language. Kling 3.0 Omni is a separate variant, but its differentiator is reference-based voice and character consistency (upload a reference video and it locks visual and vocal traits across scenes) plus multi-shot storyboarding — not audio generation itself, which the base model already has
- Pricing: $0.084/s (standard, no video input) to $0.168/s (Pro, with video input) on the official per-second API; consumer plans are credit-based (Free: 66 credits/day; Pro: $29.99/mo for 3,000 credits; Ultra: $59.99/mo for 8,000 credits plus 4K/60fps access)
- Artificial Analysis Elo: 1,110 (1080p Pro), 1,097 (720p Standard), 1,095 (Omni 1080p Pro) — text-to-video board
Strengths
Avatar Performance: The dedicated Avatar mode's 5-minute continuous generation remains unmatched among the models covered here for long-form talking-head content.
Native 4K/60fps is now standard on the flagship tier, not an add-on.
Multi-language native audio on the base model, including per-character language control within the same multi-character scene — a level of dialogue control not documented on any other model in this comparison.
Limitations
- The Omni variant — reference-based voice/character consistency and multi-shot storyboarding — currently scores slightly lower on the Arena (1,095) than the standard Pro tier (1,110)
- Credit-based consumer pricing makes cost comparison across plans harder than a flat per-second rate
Best For
Virtual-presenter and long-form avatar content specifically (via Avatar mode), or general 4K/60fps generation on a lower per-second cost than Veo or Runway.
Use Case Recommendations
Choose Runway Gen-4.5 If...
- Character consistency across multiple scenes is your top priority
- You value Runway's editing ecosystem over chasing the current top benchmark score
Choose Veo 3.1 If...
- You need a stable, documented, enterprise-ready API today
- You're already on Google Cloud/Vertex AI infrastructure
- You need long chained sequences (140s+) via scene extension
Choose Gemini Omni Flash If...
- Multi-turn conversational editing matters more than topping a raw-quality benchmark
- You can work within the Gemini app today and wait for broader API access
Choose LTX-2.3 If...
- You require on-premise deployment or fine-tuning
- You want native 4K without paying a proprietary 4K surcharge
- Budget is the binding constraint ($0.06/s fast tier)
Choose Wan 2.7 If...
- You want an Apache-licensed model that's actually competitive on the live leaderboard
- First/last-frame control is important to your workflow
Choose Seedance 2.0 If...
- You want the best currently-available combination of visual quality and native 4K
- You're working through ByteDance's consumer apps rather than waiting on the standalone API
Choose Seedance 2.5 If...
- You need 30-second native clips or local (partial-scene) editing, and can confirm it's live in your region first
Choose Kling 3.0 If...
- You need 5-minute continuous avatar/talking-head generation
- You want native 4K/60fps at a lower per-second rate than Veo or Runway
How RizzGen Uses These Models
RizzGen doesn't lock you into one video model. Our routing layer evaluates every model covered in this comparison against what your shot actually needs — resolution, audio, motion complexity, reference count, budget — and automatically selects the best-fit engine for that specific generation. We then apply RizzGen's own scene-based character locking and multi-scene consistency layer on top, so switching which underlying model handles a shot doesn't mean re-learning a new prompt syntax or losing character continuity between scenes.
We don't replace these models. We make the choice between them automatic, and make them controllable for production workflows.
Let RizzGen Pick the Right Model, Every Time
Automatic model routing plus scene-based consistency, so you direct the story instead of managing seven different APIs.
FAQ
Is Runway Gen-4.5 still the #1 AI video model?
Not verifiably, as of July 18, 2026. Runway's own launch-day claim of 1,247 Elo (#1) was real at the time, but Runway Gen-4.5 does not currently appear anywhere on Artificial Analysis's live Text-to-Video or Image-to-Video leaderboards. Gemini Omni Flash (1,240 Elo) and Seedance 2.0 (1,225 Elo) currently hold the top two spots with verifiable, live data.
What's the difference between Veo 3.1 and Gemini Omni Flash?
They're separate Google models. Veo 3.1 (October 2025) is the specialized cinematic video model and remains the documented API/Vertex AI baseline. Gemini Omni Flash (announced May 2026) is a newer multimodal model that replaced Veo inside the consumer Gemini app and currently leads the Artificial Analysis leaderboard, but its differentiator is multi-turn conversational editing, not raw frame quality — testers rank Seedance 2.0's visuals above it.
Is Seedance 2.5 available yet?
Partially. It began rolling out July 16, 2026 through ByteDance's Jimeng app, Dreamina, and the BytePlus API, first in China, with wider access following through July 2026. It is not yet in full global release and doesn't have finalized public pricing or Artificial Analysis benchmark data as of this update.
Which model has the best character consistency?
Runway Gen-4.5 is still widely regarded as the strongest for cross-scene character consistency, followed by Seedance 2.0's reference-based control. For long-form avatar consistency specifically, Kling's dedicated Avatar mode (5-minute continuous generation) is unmatched among these models.
Which is the cheapest for high-volume generation?
LTX-2.3's fast tier at $0.06/second is the lowest listed rate here, and it's open-source so self-hosting removes the per-second cost entirely if you have the GPU budget. Wan 2.7 ($0.10/s, Apache 2.0) is the next-cheapest fully open option.
Can I use these models commercially?
Runway, Veo, Gemini Omni, Kling, and Seedance allow commercial use on paid tiers — check each vendor's current terms of service, since these change. LTX-2.3 and Wan 2.7 are open-weight/open-license (check the specific license file on Hugging Face or the model repo for your jurisdiction and use case).
Which has the best native audio?
By Artificial Analysis ranking, Gemini Omni Flash and Seedance 2.0 lead. Veo 3.1's audio remains strong and well-documented. Kling 3.0 also generates native audio on its base model — including multi-language dialogue and per-character language control in multi-character scenes — with the separate "Omni" variant adding reference-based voice/character consistency on top, not audio itself. Wan 2.7's audio capability isn't clearly documented, so we're not claiming it here.
What about 4K video?
Four models now offer some form of 4K: LTX-2.3 (native 4K/50fps as a standard tier across its full duration range, the most consistent 4K offering), Seedance 2.0 (a dedicated 4K tier added June 2026, at roughly 5x its base price), Kling 3.0 (native 4K/60fps on its flagship tier), and Veo 3.1 (up to 4K on a single 8-second base clip, dropping to 720p for chained/extended sequences). Wan 2.7 and Runway Gen-4.5 cap at 1080p.
Sources & Methodology
Every specification, price, and ranking above was checked directly against official vendor documentation or the live Artificial Analysis leaderboard on July 18, 2026, rather than pulled from a single secondary aggregator. Where sources disagreed or a figure wasn't officially published, we said so in the text rather than presenting a single unverified number as fact. Key primary sources used:
- Artificial Analysis — Text-to-Video Arena Leaderboard
- Artificial Analysis — Image-to-Video Arena Leaderboard
- Runway — Introducing Gen-4.5 and Runway API Pricing Docs
- Google — Video Generation in the Gemini API (Veo 3.1) and Gemini Omni Overview
- Lightricks — LTX-2.3 on Hugging Face
- Kuaishou Investor Relations — Kling AI 3.0 Launch
- ByteDance/Volcano Engine Seedance 2.0 and 2.5 pricing and rollout coverage (BigGo Finance, BytePlus resource pages)
- Alibaba Wan 2.7 official model release coverage (April 2026)
This is a fast-moving space — pricing, availability, and benchmark rankings can change within weeks. If you're building a production pipeline around any model here, verify current specs directly with the vendor before committing.