Field intelligence for AI-first professionalsVol. II · Nº 56 · Saturday, August 15, 2026
← Phantom Notes
56AI ToolsAugust 15, 202611 min read

Seedance 2.5 Makes 30-Second Movies With Sound. The 4K Claim Is Marketing.

ByteDance's new video model generates 30 seconds of picture and audio in one pass, takes 50 reference files, and edits existing footage by timestamp. We verified the specs against every live API: the real output caps at 720p, the price is up roughly 50% over Seedance 2.0, and the new model has no independent leaderboard score yet. The current champion is its own predecessor.

By T.W. Ghost


ByteDance shipped Seedance 2.5 in two steps: a consumer launch on July 31, 2026 in its Jimeng AI and Doubao Pro apps (the international Dreamina app followed about a day later), and a public developer API on August 7 through BytePlus ModelArk, which fal.ai, Replicate, OpenRouter, EvoLink, and Atlas Cloud began reselling the same day. It is fully usable from the US.

The pitch is "one-take creation": up to 30 seconds of video and audio generated together in a single pass, dialogue lip-synced in 11 languages, steered by up to 50 reference files. Much of that is real and verified. Some of what is circulating about this model is not, including a 4K spec that no live API actually offers and at least one widely-indexed "documentation" site with entirely fabricated numbers.

Here is what survived checking, with the vendor claims labeled as vendor claims.


What Actually Launched

Seedance 2.5
Consumer launchJuly 31, 2026 (Jimeng AI, Doubao Pro; Dreamina worldwide ~Aug 1)
Public APIAugust 7, 2026 (BytePlus ModelArk; resold by fal, Replicate, OpenRouter, EvoLink, Atlas Cloud)
GenerationVideo + audio in one pass: dialogue, sound effects, lip sync
Duration4 to 30 seconds per generation (2.0 capped at 15), extendable; beta long-video mode reaches 180 seconds
Resolution (API)480p or 720p, six aspect ratios, 24 fps. That is the complete list
ReferencesUp to 30 images + 10 video clips + 10 audio clips per request (2.0 allowed 12 files total)
Languages (lip sync)Chinese, English, Spanish, Indonesian, Malay, Thai, Arabic, Portuguese, Vietnamese, Japanese, Korean
ModesText-to-video, image-to-video (first/last frame), reference-to-video, plus editing of existing footage
Benchmarks published by ByteDanceNone

The audio is the structural difference from most of the field: sound is not layered on after generation, it is decided jointly with the picture inside the same pass. On fal's API the audio switch defaults to on and costs nothing extra.

The duration story needs precision, because ByteDance's own pages disagree. The product page says a 30-second generation can be extended twice. The launch blog says multi-round extensions produce "videos lasting several minutes." ByteDance's Dreamina page documents a beta long-video mode reaching 180 seconds. No official source quantifies the per-extension math, so treat "30 seconds native, about 3 minutes with the beta extension mode" as the defensible envelope.


The Reference System Is the Real Story

The most substantive upgrade is not length, it is direction. Seedance 2.5 accepts up to 50 reference assets in one request: 30 images, 10 video clips, 10 audio clips (the video and audio pools carry a running-time budget of roughly 30 seconds; sources disagree on whether that budget is per pool or combined). Seedance 2.0 allowed 12 files. In prompts, you address them like variables: [Image1] for a face, [Video1] for a camera move, [Audio1] for a voice.

That reference density feeds a set of editing modes that go beyond what the big Western rivals expose:

  • Timestamp-level region editing: change a specific thing in a specific window of an existing clip
  • Green-screen replacement and background swaps
  • Camera-perspective re-editing of footage that already exists
  • White-model control: feed it a textureless 3D blockout and it renders the shot, a previsualization workflow lifted straight from film production

One hands-on review (MindStudio's, a single reviewer, labeled as such) held character consistency across a 2-minute short film using just three references, and praised the acting and dialogue pacing in 30-second takes. The same review found morphing artifacts in fast action that upscaling cannot fix, a voice that spontaneously acquired a British accent until the prompt specified American, and older Seedance 2.0 prompt styles producing visibly broken results. New model, new prompt dialect.


The 4K Claim Is Marketing

Dreamina's consumer page says "generate cinematic 4K videos with Seedance 2.5." We went looking for the 4K. It does not exist anywhere you can touch:

  • ByteDance's launch post and product page make no resolution claim at all
  • The native-4K announcement in ByteDance's ecosystem belongs to Seedance 2.0, not 2.5
  • Every live API surface (fal, EvoLink, OpenRouter resellers, BytePlus token examples) exposes exactly 480p and 720p
  • The Jimeng web app listed 480p and 720p tiers at launch, and a hands-on reviewer measured the 720p cap directly
  • EvoLink's own status page warns developers, in as many words, that creator-product marketing and API output tiers "should be treated as separate claims"

Whether Dreamina's consumer "4K" is an upscaling pipeline or pure copywriting is undocumented. Either way: if you are budgeting an API workflow, the number is 720p.

While fact-checking this we also found a widely-indexed "Seedance 2.5 API guide" site confidently publishing 4K output specs, per-second RMB pricing, and inference-step parameters that contradict every verified source. It ranks well in search. This is the state of AI journalism in 2026, and it is why this site prints sources.


What It Costs

Official BytePlus rates are token-based, and the formula is refreshingly transparent: tokens = (input video duration + output duration) x width x height x fps / 1024, billed only on successful generations. Per CometAPI's breakdown of the BytePlus rates: $10.70 per million tokens without a video input, $6.40 with one. In clip terms:

480p720p
5-second clip~$0.51~$1.16
Per second~$0.10~$0.23
Full 30-second clip~$3.08~$6.94

That is roughly 47-52% above Seedance 2.0's rates, and resellers mark it up further (fal works out near $0.47/s at 720p, though reference-video requests get a 0.6x multiplier; EvoLink and Atlas Cloud land between $0.13 and $0.29/s).

Against the field, per CometAPI's aggregator (updated August 15, 2026): Kling 3.0 runs $0.084-0.168/s, Veo 3.1 spans $0.03/s (Lite) to $0.40/s (Standard with audio), and Runway Gen-4.5 is $0.12/s. So Seedance 2.5 at 720p sits mid-pack, cheaper than Veo's top tier, pricier than Kling.

And one competitor is leaving the table entirely: OpenAI deprecated the Sora 2 video API on March 24, 2026 and shuts it down September 24, per OpenAI's own deprecations page, with the Sora web app already discontinued in April. The API video race is now ByteDance, Google, Kuaishou, MiniMax, and xAI.


The Leaderboard Entry That Does Not Exist Yet

Here is the honesty section, and it cuts both ways.

As of August 15, 2026, Seedance 2.5 does not appear on the Artificial Analysis Video Arena at all. Not on text-to-video, not on image-to-video, not on the models page. Every "new king of AI video" headline you have seen about this model is running ahead of any independent measurement. There is no Elo. We checked the boards directly.

What the boards do show, snapshot dated August 15:

Image-to-video (with audio)Elo
Dreamina Seedance 2.0 720p1198
MiniMax H31192
Gemini Omni Flash1191
grok-imagine-video-1.51114
Veo 3.1 family1066-1086
Kling 3.0 family1055-1077
Text-to-video (with audio)Elo
Gemini Omni Flash1241
MiniMax H31238
Dreamina Seedance 2.0 720p1222
Wan 2.71160
Kling 3.0 family1101-1109

So the current image-to-video champion is Seedance 2.5's own predecessor, which has meanwhile slipped to third on text-to-video behind Google and MiniMax. The franchise pattern is real: Seedance 1.0 debuted at #1 on both arenas in June 2025 (per ByteDance's own launch blog), and Seedance 2.0 reportedly took #1 on both at its February debut. ByteDance has earned the benefit of the doubt on quality. It has not yet earned this model's crown, because nobody independent has scored it.


Context you should have before building anything on this model line. Seedance 2.0's February 12 launch triggered the fastest IP backlash in the short history of AI video: the Motion Picture Association demanded a shutdown the same day ("In just one day, Seedance 2.0 engaged in widespread unauthorized use of U.S. copyrighted materials"), Disney sent a cease-and-desist on February 13, Paramount followed within 48 hours, and in March, Senators Blackburn and Welch formally demanded ByteDance shut Seedance down entirely, per the senators' own press releases.

ByteDance's response was mitigations (tighter face filters and visible AI labels on its consumer apps), not a pause. Five months later it shipped 2.5, demoed with a fully AI-generated short film. The legal questions are unresolved, the model is stronger, and the moderation burden increasingly sits with the platform reselling it. If you are building client work on Seedance, that is a risk line in your contract, not a footnote.


Verified vs Unconfirmed: The Scorecard

ClaimVerdict
Consumer launch July 31, API August 7, 2026Verified (ByteDance blog; OpenRouter and EvoLink corroborate)
30s single-pass audio-video, up from 15s on 2.0Verified (ByteDance blog; fal and Replicate API docs)
50 references = 30 images + 10 videos + 10 audioVerified (ByteDance blog; multiple providers)
11-language lip syncVerified (identical lists, ByteDance and fal)
Native 4K output on Seedance 2.5Refuted for every accessible surface (480p/720p only on all APIs and app tiers; "4K" appears solely in consumer marketing)
BytePlus pricing $10.70/$6.40 per 1M tokensVerified via two independent sources (CometAPI, Renoise; BytePlus's own page is JS-walled, so printed as attributed)
"Leads the Artificial Analysis arena"Refuted (2.5 is absent from the arena entirely; the I2V leader is Seedance 2.0)
Seedance 2.0 Elo 1269/1351 at its debutReported, not independently confirmed (March snapshots from secondary sources)
Sora 2 API deprecated March 24, shutdown Sept 24, 2026Verified (OpenAI's deprecations page)
Generation latency, rate limits, watermark detailsUndocumented (nothing credible published)

Who Should Use It

Use Seedance 2.5 if you make short-form content with dialogue and want picture and sound in one generation: talking scenes, product spots, multi-shot sequences held together by reference files. The reference system and timestamp editing are genuinely ahead of the Western competition's public APIs, and the per-clip economics at 480p are workable for iteration.

Budget for 720p and prompt rewrites. The 4K is not real for API users, older Seedance prompt styles misfire on 2.5, and nobody has published latency numbers, so prototype before you promise delivery dates.

Wait for the arena before repeating any "best video model" claim. Its predecessor holds the only crown in the family right now. If the pattern of the last two launches holds, 2.5 will chart high when it lands, and this post will get an update either way.

Client work warning: the Seedance line operates under active studio cease-and-desists and congressional pressure. Keep likeness and IP clauses in your contracts.

If you want to actually learn the tool, our six-module Seedance Mastery course covers prompting, audio-video sync, and production workflows, and our resources directory has the full AI video tool list.


Sources


*Building video into your workflow and not sure which model, or which AI stack, actually fits? Take the free 2-minute quiz and get matched. Then go deep with the Seedance Mastery course.*

Which model should you be using?

Three minutes, twelve questions, one defensible answer.

Take the quiz