WanWan 3.0 Prime

Wan 3.01080p. Create, edit, extend.

Generate, edit, and extend complete AI video with native audio, precise frame control, and up to twenty multimodal references.

2–30soutput duration
1080pmaximum resolution
20combined references
Nativeaudio generation

Interactive Wan 3.0 workspace

Direct the model, not just the prompt.

Explore the actual Wan 3.0 controls, assemble a shot, then open the complete generation workspace in Unify.

Generate

Create a new sequence from text, frames, or a multimodal reference pack.

Source setup

Prompt-only generation

No source media required. Describe subject, motion, camera, light, and sound below.

206 / 20,000
Resolution
Aspect ratio
Duration

Model intelligence

Wan 3.0 Prime1080p16:9 · 10sAudio on
Continue in the Wan workspace

This page previews the workflow. Source files are selected securely again inside the generator.

20 combined references First + last frame Native sound 30 fps output

Creative direction board

One model. Every kind of motion.

Explore the visual formats available across Unify, then use Wan 3.0’s text, frame, video, image, and audio controls to direct your own version.

Commercial

Fast advertising concepts

Consistent motion

Controlled continuity

Complete capability map

Create. Control. Transform.

Wan 3.0 uses one model for new generation, multimodal direction, editing, and extension—without splitting the workflow across separate tools.

Generate

Text to video

Create a complete scene from a prompt, including subject, action, camera, pacing, and sound.

Reference

Image-guided video

Use reference images to preserve a person, product, environment, or visual language.

Frames

First and last frame

Define the opening frame alone or pair it with an ending frame for controlled transitions.

Multimodal

Up to 20 references

Combine as many as 10 images, 5 short videos, and 5 audio clips in one reference-led brief.

Sound

Native audio

Generate dialogue, background music, ambience, and sound effects with the moving image.

Edit

Prompt-led video editing

Upload a source video and describe the transformation while retaining the material you want to keep.

Extend

Video extension

Continue an existing video with adaptive framing and model-selected duration based on the source.

Context

File input

Supply one supported file as source context when a prompt alone is not enough.

Context

Web link input

Use one supported link as source context instead of uploading a file.

Output

480p to 1080p

Move from efficient 480p drafts to 720p or full-HD 1080p delivery.

Format

Six aspect modes

Choose adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16 for film, feeds, and vertical video.

Duration

2 to 30 seconds

Select any whole-second duration across the supported range, rendered at 30 frames per second.

Inputs and output

A larger creative brief.

Mix the sources a production actually uses while keeping every hard limit visible before upload.

Reference images

Up to 10

JPEG, PNG, WebP, or BMP · 20 MB each

Reference videos

Up to 5

MP4 or MOV · 15 seconds and 100 MB each

Reference audio

Up to 5

MP3 or WAV · 15 seconds combined · 15 MB each

Combined references

Up to 20

Images, video, audio, and one file or link

Prompt

Up to 20,000 characters

Prompt extension is enabled by default

Output

2–30 seconds

480p, 720p, or 1080p · 30 fps

Edit and extend

The first render is only the start.

Upload a selected take, isolate the change in words, or continue the next beat from the same source video. Wan 3.0 keeps generation, editing, and extension inside one model.

Four-step workflow

From brief to continuation.

01

Generate a new scene

Start from text, a first frame, a first-and-last-frame pair, or a set of references. Name each source in the prompt and describe exactly what it controls.

02

Direct motion and sound

Describe the action, camera path, timing, dialogue, ambience, music, and sound effects. Wan 3.0 can resolve picture and native audio together.

03

Edit the selected take

Switch to Edit, provide the source video, and isolate the requested change. State what must remain untouched to protect continuity.

04

Extend the sequence

Switch to Extend and describe the next beat. Wan follows the source framing and selects a compatible continuation length automatically.

Wan 3.0 prompts

Write for motion, continuity, and sound.

01

First-and-last-frame product reveal

Use Image 1 as the first frame and Image 2 as the final frame. Begin on an extreme macro of condensation moving across the bottle, then pull back in one smooth arc as the studio light changes from cool blue to warm amber. End precisely on Image 2. Premium commercial lighting, realistic reflections, soft room tone and a restrained bass hit at the reveal.

02

Multimodal character scene

Keep the lead character identical to Reference Image 1 and use Reference Video 1 only for walking rhythm. She crosses a windswept train platform at night, coat moving naturally, then stops as the train arrives behind her. Slow handheld push-in, realistic weight, distant rail ambience, no dialogue, preserve face and wardrobe in every frame.

03

Edit an existing advertisement

Edit the supplied video. Keep the camera movement, product position, hands, timing, and soundtrack unchanged. Replace only the environment with a minimalist gallery made from pale concrete and brushed steel. Match the original shadows and reflections so the product remains grounded in the shot.

FAQ

Everything before render.

What is Wan 3.0?

+

Wan 3.0 is Alibaba's all-in-one multimodal AI video model for text-to-video, frame-guided generation, reference-led generation, video editing, extension, and native audio. Unify provides the Wan 3.0 Prime quality tier in its video workspace.

How long can Wan 3.0 videos be?

+

A generated output can be any whole-second duration from 2 to 30 seconds. When reference video is supplied, its duration and the requested output must fit within the model's 30-second combined budget.

Which resolutions and aspect ratios does Wan 3.0 support?

+

Wan 3.0 supports 480p, 720p, and 1080p. Available aspect modes are adaptive, 16:9, 4:3, 1:1, 3:4, and 9:16.

Can Wan 3.0 generate audio?

+

Yes. Audio is enabled by default and can include dialogue, background music, ambience, and sound effects generated with the video.

How many references can I use?

+

A request can contain up to 20 references overall: as many as 10 images, 5 short videos, and 5 audio clips, plus the supported file-or-link input rules. First/last frame mode cannot be mixed with the multimodal reference mode in the same request.

Can Wan 3.0 edit or extend a video?

+

Yes. Edit and Extend are explicit modes in Unify. Both require a source video and use adaptive framing with smart duration so the model can follow the supplied footage.

Can I try Wan 3.0 on Unify?

+

Yes. Open the Unify video generator from this page and Wan 3.0 Prime will already be selected. The current credit quote is shown before submission.

Wan

Direct the whole sequence.

Choose Wan 3.0 Prime in Unify and move from prompt to reference-led generation, editing, and extension.