Wan 3.0 Video

by AlibabaVideo

Wan 3.0 Video is an Alibaba video generation API on TTAPI. Generate native clips up to 30 seconds from text, frames, reference media, files, and optional audio.

InputMultimodaltext · media · files
Duration2–30 secor smart duration
FlowAsynctask result URL
Price3.5 q/s480P promotional rate
Example clipVideo preview
Heritage headwear showcase

参考上传的5张关键帧图片,生成一支10秒、16:9横版的国风人物头饰展示短视频。整体风格延续参考视频的节奏与展示逻辑,画面精致、干净、克制,带有时尚国风与中式视觉设计感。5位人物按照上传顺序依次出场,人物始终固定在画面正中央,构图稳定,不要人物跑位,不要大幅运镜,不要夸张肢体动作。每位人物停留约2秒,主要动态为轻微左右摇头、缓慢偏头、微微抬颌或回正,让头饰得到充分展示。头饰上的流苏、花朵、喜鹊、纸鸢、蝴蝶等装饰可以有轻微摆动或颤动,背景中的中式纹样如祥云、亭台楼阁、山水纹样保持相对静止,仅有极轻微的呼吸感或视差感。 人物之间的切换方式要统一为“水平翻牌式切换”:当前人物画面像一张竖立在中心的卡片,围绕竖直中轴做平滑的水平翻转,翻到侧面时短暂变窄,随后下一位人物从翻转的另一面自然出现,形成连续的卡牌翻面效果。切换过程干净利落、平滑顺畅,节奏清晰,不要飞散,不要缩放乱跳,不要人物从边缘飞入。切换时人物位置不能偏移,始终保持中心对齐,像在同一个展示台上不断翻牌更换不同人物。 每一位人物的单独动态要求如下:人物保持半身肖像式展示,眼神自然,表情平静高级,身体整体稳定,只有头部轻微向左偏转再回正,或向右偏转再回正,动作幅度小而优雅,用来展示头饰层次。耳饰、流苏、花枝、鸟羽、纸鸢细节可以随动作产生轻微延迟摆动,增强精致感。不要说话,不要夸张眨眼,不要大笑,不要大幅度扭身。 整体镜头语言以固定镜头为主,可有极轻微缓慢推进,营造高级海报动态感,但不能破坏人物始终居中的展示效果。视频整体节奏是“停留展示—水平翻牌—停留展示—水平翻牌”,重复5位人物,最终在最后一位人物上稳定停留结束。 分镜时间轴版 0s - 2s 第一位人物居中展示。 人物轻微将头从正面偏向一侧,再缓慢回正,展示头顶花朵、喜鹊与流苏头饰。 流苏和小装饰轻微摆动,背景纹样保持稳定。 2s - 2.4s 进入第一次翻牌切换。 当前人物画面围绕竖直中轴做水平翻转,像卡片翻面一样,翻到最窄处后自然切换到下一位人物。 翻转时人物仍固定在中心。 2.4s - 4.4s 第二位人物居中展示。 侧脸人物可做轻微抬颌、微小回头或轻轻转回一点,再停住,突出纸鸢、花朵和喜鹊头饰。 花瓣或小装饰可以极轻微漂浮。 4.4s - 4.8s 第二次翻牌切换。 同样采用水平翻牌式过渡,干净利落,不要多余特效。 4.8s - 6.8s 第三位人物居中展示。 人物轻微左右摇头,小幅展示夸张头饰的不同层次。 头饰中的鸟、纸鸢、流苏可有细微摆动,背景云纹轻微呼吸感。 6.8s - 7.2s 第三次翻牌切换。 继续保持统一的中心翻牌节奏。 7.2s - 8.6s 第四位人物居中展示。 人物轻微偏头,露出头饰正面与侧向层次,动作幅度小,精致克制。 8.6s - 9.0s 第四次翻牌切换。 水平翻转,人物不偏移,节奏平稳。 9.0s - 10s 第五位人物居中展示并作为结尾。 人物轻微抬头或缓慢转头回正,头饰小元素轻轻晃动。 最后稳定停住,定格结束。 不要人物走动,不要人物缩放漂移,不要头饰变形,不要人物位置变化,不要背景乱飞,不要复杂粒子,不要夸张镜头旋转,不要卡顿,不要多人同屏。

Reference media · 1080P · 10 seconds · With audio
Examples

Wan 3.0 creator examples

Capabilities

What you can build with Wan 3.0 Video

Native 30-second generation

Generate integer durations from 2 to 30 seconds, or let Wan select a smart duration for the prompt.

Multimodal references

Guide a result with images, video, audio, a webpage, or a document while preserving reference details and timing.

First and last frame control

Constrain the opening frame, or provide both opening and closing frames for a controlled transition.

Synchronized audio

Create video with an audio track by default, or disable audio for silent production workflows.

Pricing

Per-action, in quota

Usage is metered in quota by model, action, and output settings.

Wan 3.0 official channel · limited-time 30% offOfficial pricing
ModelResolutionTTAPI priceOfficial price
wan3.0-video480P3.5 quota / second5 quota / second
wan3.0-video720P7 quota / second10 quota / second
wan3.0-video1080P14 quota / second20 quota / second
API request

Wan 3.0 Video operations and request fields

Use the endpoint below with your TTAPI key and the request fields shown.

Text to videoPOST
/wan/api/v1/services/aigc/video-generation/video-synthesis

Generate a Wan 3.0 video from a text prompt without media input.

Official reference
First framePOST
/wan/api/v1/services/aigc/video-generation/video-synthesis

Generate a video while strictly preserving a supplied first frame.

Official reference
First and last framesPOST
/wan/api/v1/services/aigc/video-generation/video-synthesis

Generate a transition constrained by supplied first and last frames.

Official reference
Reference mediaPOST
/wan/api/v1/services/aigc/video-generation/video-synthesis

Generate from reference images, video, audio, or a web link.

Official reference
Reference filePOST
/wan/api/v1/services/aigc/video-generation/video-synthesis

Generate from a referenced document file that Wan interprets as creative context.

Official reference
FetchGET
/wan/api/v1/tasks/{task_id}

Retrieve task status, generated video URL, and usage by task ID.

Official reference
POSThttps://api.ttapi.io/wan/api/v1/services/aigc/video-generation/video-synthesis

Headers

TT-API-KEYstringrequired

Your TTAPI API key.

Content-Typestringrequired

Use application/json.

Body

modelenumrequired

Use wan3.0-video for the standard Wan 3.0 model.

input.promptstringconditional

Text prompt describing the video. Supports Chinese or English up to 20,000 characters; either a prompt or media input is required.

input.mediaobject[]conditional

Optional media inputs for first/last frames, reference images, reference video, reference audio, a web link, or a document file.

input.media[].typeenumconditional

first_frame, last_frame, reference_image, reference_video, reference_audio, link, or file. Reference media and file/link modes cannot be mixed in one request.

input.media[].urlstringconditional

Public URL or supported Base64 data for the corresponding media item.

parameters.resolutionenumoptional

Output resolution: 480P, 720P, or 1080P. Defaults to 1080P.

parameters.ratioenumoptional

adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16. Defaults to adaptive.

parameters.durationintegeroptional

Output duration from 2 to 30 seconds, or -1 for smart duration. Defaults to 5 seconds.

parameters.audiobooleanoptional

Generate synchronized audio with the video. Defaults to true.

parameters.seedintegeroptional

Seed from 0 to 2,147,483,647 for reproducible generations.

parameters.prompt_extendbooleanoptional

Enable automatic prompt enhancement. Defaults to true.

parameters.watermarkbooleanoptional

Add a watermark when true. Defaults to false.

Open the official TTAPI documentation
Integrate

From key to first result

Copy the resolved endpoint, authentication headers, and request body for this model.

curl --request POST \
  --url https://api.ttapi.io/wan/api/v1/services/aigc/video-generation/video-synthesis \
  --header 'TT-API-KEY: $TTAPI_KEY' \
  --header 'Content-Type: application/json' \
  --data '{"model":"wan3.0-video","input":{"prompt":"这是一个正面的近景镜头,画面主要呈现一位年轻黑人女孩的头部和上半身,她穿着一件蓝色的毕业袍。镜头稍微偏一点点,可以看到左边还有一个人,但背景被虚化处理了,看起来很模糊。女孩正对着镜头说话,嘴巴在动,眉头稍微皱着,表情看起来有点严肃和忧虑。光线从上方照下来,把她脸上的轮廓衬托得很自然,甚至能看到皮肤真实的质感。镜头虽然固定,但带着一点点人手持拍摄时的轻微呼吸感和晃动,视线从看着前面慢慢变成了低头看手里。"},"parameters":{"resolution":"1080P","ratio":"16:9","duration":10,"audio":true,"seed":12345,"prompt_extend":true,"watermark":false}}'