GET /v1/models to discover which models are available to the API and what
they cost. Pass ?type=image, ?type=video, or ?type=audio to filter.
Reading the response
Multiple images per request
Some image models return more than one image from a single call:- seedream-4.5 and seedream-5.0-lite support grouped generation. Set
sequential_image_generation: "auto"andmax_images(up to 4 on seedream-4.5, up to 14 on seedream-5.0-lite). The model returns a group of related images up tomax_imagesand decides the exact count — it may return fewer (a plain prompt often yields one; a prompt that explicitly asks for a set yields several). Usemax_imagesfor this, notnum_images. - seedream-5.0-pro does not support grouped generation — it has no
sequential_image_generationparameter. Usenum_images(1–4) for multiple takes; each is generated and billed independently. It prices by output size (1K = 55, 2K = 110 credits per image) and billsreference_imagesbeyond the first at 5 credits each — see pricing & credits. - seedream-5.0-flash has no grouped generation either: use
num_images(1–4). It costs a flat 25 credits per image at1Kor2K, and itsreference_images(up to 10) are free.
Transparent backgrounds
seedream-5.0-pro and seedream-5.0-flash take abackground parameter (opaque, the
default, or transparent). Transparent mode is for editing a cutout: send exactly one
reference image, a PNG that already has transparent pixels (a product cutout, sticker or
logo), and the result is a PNG that keeps the transparency. It can’t make a transparent image
from text alone, and multi-reference requests must stay opaque — those requests are rejected
with 400 before anything is charged.
3 × credits_per_generation.
When the job starts we reserve the maximum (max_images × credits_per_generation) and charge
only for the images actually produced.
Video generation modes
Some video models accept different inputs for different modes.SeeDance 2.5 (flagship)
SeeDance 2.5 (seedance-2.5) is ByteDance’s flagship multimodal model. One endpoint,
three modes — there is no separate video-edit endpoint:
- Text-to-video —
promptonly (aspect_ratioapplies here; other modes are adaptive). - Image-to-video — a start frame (
image), optionally with an end frame (end_image). - Reference-to-video — any mix of up to 15
reference_images(free), up to 5reference_videos(each 2–15s, 30s combined), and up to 5reference_audios(30s combined, free; audio-only input works). Editing or extending an existing clip = passing it as areference_videosentry. Reference media can’t be combined with start/end frames.
reference_videos additionally bill per
second of input at half the output rate (480p = 80, 720p = 160, 1080p = 375 credits per
input second) — use
POST /v1/cost for an exact quote
(it probes your clips).
SeeDance 2.0
SeeDance 2.0 (seedance-2.0) supports all of these from one endpoint:
- Text-to-video —
promptonly. - Image-to-video — a start frame (
image), optionally with an end frame (end_image) to interpolate between the two. - Reference-images-to-video — up to 9
reference_imagesto guide the result. - Video-to-video — supply a
videoto edit an existing clip. Video inputs are passed by URL (a public URL or one fromPOST /v1/uploads), not inlined as base64. You can also passreference_images(up to 6) and/or anaudiotrack alongside thevideoto guide the edit.
image/end_image) can’t be combined with reference_images or a video.
But reference_images can accompany a video — that runs a reference-guided video-edit
(up to 6 reference images in that mode, vs up to 9 on their own). Pricing is per
second, by resolution (see resolution_pricing): a 5-second 1080p clip costs 5 × 550.
The 4k resolution is available for text/image/reference modes, not video-to-video.
Reference audio (optional): attach an audio track (MP3/WAV/OGG, by public URL or one
from POST /v1/uploads) to an image- or video-input request and
SeeDance 2.0 uses it in the generated video. Audio must accompany an image/reference_images
or video input — it can’t be the only input.
Cheaper tiers: seedance-2.0-fast and seedance-2.0-mini support the same modes at
lower per-second rates, in 480p and 720p (no 1080p/4k). Check each model’s
resolution_pricing in GET /v1/models, or POST /v1/cost
to price an exact request.
MiniMax H3
MiniMax H3 (minimax-h3) is a multimodal 2K model with a different reference surface:
- Text-to-video —
promptonly (aspect_ratioapplies here; other modes are adaptive). - Image-to-video — a start frame (
image), optionally with an end frame (end_image). - Reference-to-video — mix up to 5
reference_images, up to 3reference_videosclips, and up to 3reference_audiosclips in one request. Clips are passed by URL (public or fromPOST /v1/uploads); each clip must be 2–15s, with at most 15s combined per media type.
MiniMax H3 Max
MiniMax H3 Max (minimax-h3-max) is a post-trained H3 tuned for stronger prompt
adherence and aesthetics, with native audio. One endpoint covers every mode — the inputs
decide which:
- Text-to-video —
promptonly;aspect_ratiodefaults to16:9. An optionalaudioURL pins a soundtrack (it replaces the generated audio, trimmed to the video length). - Image-to-video —
image(first frame) and/orend_image(last frame). An end frame on its own is allowed. The output follows the frame image, soaspect_ratiois ignored.audioworks here too. - Reference-to-video — up to 5
reference_images, 3reference_videos(2–15s each, 15s combined) and 3reference_audios(2–15s each, 15s combined). Address them in the prompt by order: “Image 1”, “Video 1”, “Audio 1”.aspect_ratiodefaults toadaptive(the model picks); an explicit ratio is honored. - Extend — pass a
video(1.6–60s, ≤50MB, aspect ratio between 2:5 and 5:2) and describe what happens next.length_secondsis the new footage appended; the output is the source followed by the continuation.aspect_ratiodefaults toauto(keep the source framing); any other ratio crops.
480p, 768p (default) and 1080p; 2k is available only when
extending. Any integer duration from 5 to 15 seconds. Reference media can’t be
combined with start/end frames, and video can’t be combined with any other media input.
Wan 3.0
Wan 3.0 (wan-3.0-video) is Alibaba’s all-in-one model — one endpoint, three modes,
no separate video-edit endpoint (the SeeDance 2.5 shape):
- Text-to-video —
promptonly.aspect_ratiois honored in every mode — an explicit ratio reframes the output even with a first frame or reference media, and the defaultadaptivelets the model pick a suitable ratio from the inputs and prompt. - Image-to-video — a first frame (
image), optionally with a last frame (end_image) to interpolate between the two. - Reference-to-video — any mix of up to 10
reference_images(free), up to 5reference_videos(each 2–15s, 15s combined), and up to 5reference_audios(15s combined, free). Editing, replicating the style or camera move of, or extending an existing clip = passing it as areference_videosentry. In the prompt you can address assets by their order in each type: “Image 1 hands Video 1 the guitar…”. Reference media can’t be combined with first/last frames.
Unlike MiniMax H3 and the SeeDance families, Wan 3.0 never bills input media —
reference_videos clips are free, and you pay only for the output seconds you request.
The one constraint: with a video input, the input duration plus length_seconds must not
exceed 30 seconds.wan-3.0-video-prime is the high-speed tier: the identical API
surface and modes with significantly faster end-to-end generation, at 85/170/340
credits per output second (480p/720p/1080p). Input media is free there too.
Wan 2.7
Wan 2.7 (wan-2.7-video) serves three modes from one endpoint:
- Text-to-video —
promptonly. - Image-to-video — a start frame (
image), optionally with an end frame (end_image) to interpolate between the two. - Video editing — supply a
videoand the model applies yourpromptto that clip (restyle it, change the scene, alter a subject’s action or the camera move). The video is passed by URL (public or fromPOST /v1/uploads), not inlined as base64. Up to 4reference_imagesmay accompany it to guide the edit — for example, supplying the person or style to apply to the clip.
video can’t be combined with image/end_image. Output is 720p or 1080p; text- and
image-to-video run 5, 10, or 15 seconds.
Video editing behaves differently from the other two modes in two ways worth planning for:
the input clip must be 2–15 seconds, and the output length is taken from that clip,
so length_seconds is ignored.
P-Video
P-Video (p-video) is a fast, low-cost model with three modes:
- Text-to-video —
promptonly. - Image-to-video — a start frame (
image), optionally with an end frame (end_image) to interpolate between the two. - Audio-driven — supply an
audiotrack (MP3/WAV/FLAC) by URL and the model generates video to match it.
Retired models
Model providers occasionally withdraw a model. When that happens we take it out ofGET /v1/models and its endpoint starts answering 410 Gone with the error code
model_retired, naming the replacement:
POST /v1/cost returns the same error rather than quoting a price you could not spend.
A 410 is permanent — retrying will not help; switch the model id in your request.
Retired so far:
seedance-1.5-pro (2026-09-17, → seedance-2.0).
Parameter schemas differ between models, so check the replacement’s fields in
GET /v1/models before swapping the id — and its per-second rate, which is usually
different too.