Arab AI
Futuristic 3D graphic displaying the Wan 3.0 model converting digital charts and office documents into high-quality AI-generated videos.

Alibaba Wan 3.0: Generating 30-Second AI Videos from Documents and Slides

August 7, 2026
6 minutes

On August 6, 2026, Alibaba officially announced the public beta of its next-generation AI video generation model, Wan 3.0. Following iterative updates from version 1.0 through 2.7-which were limited to 15-second generations-this major release introduces a shift in AI video production. By combining native long-duration rendering with multimodal document-to-video capabilities, Wan 3.0 goes beyond traditional text and image prompts.

Screenshot of the official announcement for Wan 3.0 public beta highlighting key features.
Official launch post of the Wan 3.0 public beta highlighting native 30-second video generation.

Native 30-Second Generations: From Short Clips to Narrative Arcs

Wan 3.0 addresses a common limitation in AI video generation: narrative consistency. The model supports native 30-second video generation in a single pass, doubling the 15-second limit of Wan 2.7. This allows creators to execute complex, continuous camera movements (such as slow zooms, panning, orbit shots, or tracking) without stitching multiple short clips together, ensuring visual and temporal consistency.

Advertisement

Intelligent Video Features: The model includes an “Intelligent Duration” feature that suggests the optimal video length based on the prompt’s context. It also features “Video Extension,” which appends consistent new scenes to an existing clip while retaining character identities, environment details, and artistic style to build longer narratives.

Omni-Reference Input: Transforming Office Documents into Video

The core update in Wan 3.0 is its “Omni-Reference” input system. This framework extends input capabilities beyond standard text, images, audio, and video reference files to support direct reading and parsing of office and web documents.

Supported file formats include:

Advertisement
  • Office & Document Files: doc, xls, ppt, pdf, key, pages, numbers, txt
  • Web Formats: md (Markdown) and live web pages via URLs
  • Limits: 1 file per request, maximum file size of 100 MB, and up to 50 pages per document.

Practical Application: Slide-to-Video

A marketing professional can upload a product presentation (such as a PPT or Keynote file) and enter a prompt like: “Create a 20-second promotional video highlighting the key features of this product with realistic cinematic shots.” The model analyzes the slides, extracts key data and structures, and generates a cohesive video featuring animated product models, dynamic graphs, and matching visual narratives.

Visual Realism and Scene Consistency

Alibaba has focused on improving visual outputs by optimizing two key areas:

  1. Character Realism:The model reduces the artificial, overly smooth “plastic” look common in older AI generation models. It delivers higher fidelity in facial features, skin textures, and micro-expressions. Even in multi-character scenes, individuals display distinct, natural emotions that align with the narrative.
  2. Temporal and Asset Consistency:Maintaining asset consistency across different frames remains a challenge in AI video. Wan 3.0 maintains consistency across:
    • Characters: Preserving facial structures, hair types, clothing, and accessories.
    • Objects: Maintaining structural details, branding logos, and material textures.
    • Environments: Keeping the spatial relationship between characters and the camera perspective intact.
    • Cinematography: Retaining the specified cinematic lighting and grading across multiple angles.

Performance Comparison: Wan 3.0 vs. Global Competitors

To understand Wan 3.0’s position in the current landscape, here is a comparison with other leading AI video models:

1. Alibaba Wan 3.0

  • Max Duration (Single Pass): 30 seconds
  • Native Resolution: Up to 1080P
  • Supported Inputs: Text, images, audio, video, documents (doc, xls, ppt, pdf, key, pages, numbers, txt, md), and web links
  • Core Advantage: Direct document-to-video conversion
  • API Pricing per Second: 0.3 Yuan (480P), 0.6 Yuan (720P), 1.2 Yuan (1080P)
  • Current Status: Public Beta

2. Seedance 2.5

  • Max Duration (Single Pass): 30 seconds
  • Native Resolution: Up to 4K
  • Supported Inputs: Text, images, audio, and video (supports up to 50 references)
  • Core Advantage: High-throughput rendering and generation speed
  • API Pricing per Second: ~1.51 Yuan (720P)
  • Current Status: Available in Production

3. Kling 3.0

  • Max Duration (Single Pass): ~10 seconds (extendable)
  • Native Resolution: Up to 2048 x 1080
  • Supported Inputs: Text, images, audio, and video
  • Core Advantage: High physical accuracy in movement simulation
  • API Pricing per Second: Varies by plan and usage tiers
  • Current Status: Available in Production

4. Google Veo 3.1

  • Max Duration (Single Pass): 8 seconds
  • Native Resolution: Up to 1080P
  • Supported Inputs: Text, images, audio, and video
  • Core Advantage: Consistent temporal synchronization and audio alignment
  • API Pricing per Second: Varies by Google Cloud usage
  • Current Status: Available in Production

5. Google Gemini Omni

  • Max Duration (Single Pass): 10 seconds (chainable for longer flows)
  • Native Resolution: Up to 1080P (720P supported on Flash)
  • Supported Inputs: Multimodal Any-to-Any (text, image, audio, video) with up to 5 reference images
  • Core Advantage: Interactive conversational editing, native integrated audio generation, and digital avatars
  • API Pricing per Second: $0.10 per second (~0.72 Yuan/sec for Gemini Omni Flash via API)
  • Current Status: Preview for developers and Google AI subscribers

Comparison Summary: Wan 3.0 is distinguished by its native 30-second duration and document-to-video processing at a competitive price point. Meanwhile, competitors excel in other niches: Seedance 2.5 offers 4K outputs, Kling 3.0 focuses on physical motion accuracy, and Gemini Omni offers native audio generation and conversational video editing.

Advanced Editing and Integrated Workflows

Wan 3.0 includes post-generation editing tools that offer granular control over the final output:

  • Multi-Dimensional Editing: Allows creators to modify scenes, plot directions, and dialogues while maintaining continuity.
  • Reference Editing: Users can change camera angles or lighting in an existing shot while keeping character poses and backgrounds static, reducing the need for complete regenerations.
  • Timestamp Control: Allows creators to trigger specific visual effects (e.g., a Dolly Zoom) at a designated second on the timeline.

Practical Use Cases for Creators and Enterprises

The capabilities of Wan 3.0 support workflows across multiple domains:

Creative & Entertainment Industry

  • Generating initial visual pitches and animatics for indie films and web series.
  • Creating music videos with motion synchronized to audio cues.
  • Developing storyboards and visual assets for gaming prototypes.

Marketing & Advertising

  • Converting product specification sheets and sales slide decks into visual ads.
  • Localizing promotional campaigns for global markets quickly.
  • Rendering product ads in diverse virtual settings without physical set production.

Education & Corporate Training

  • Transforming text-based curricula and manuals into interactive video lessons.
  • Recreating historical events or scientific concepts for training environments.
  • Producing training materials for corporate HR onboarding.

Product Design & Prototyping

  • Visualizing 3D product designs in realistic environments.
  • Generating motion previews for mobile and web interface mockups.
  • Prototyping dynamic user experience flows.

Market Adoption and Pricing

Wan 3.0 enters the market with a competitive pricing structure. While final pricing depends on cloud resource usage, the public beta rates are among the lowest in the industry: 0.6 Yuan per second for 720P outputs, which translates to roughly 18 Yuan for a 30-second video. This represents a significant cost reduction compared to alternative models like Seedance 2.5.

Limitations and Future Outlook

Despite these updates, some limitations remain in the current beta release:

  1. Audio and Text Overlay: Alibaba has noted that ambient audio generation, spatial sound effects, and text rendering accuracy inside generated videos are still being optimized.
  2. Beta Environment Stability: Being in public beta, users may experience API quota limits, occasional pricing adjustments, and limited documentation compared to older, more established model APIs.
  3. Open-Source Status: Unlike Wan 2.1, which was released as an open-source model, Wan 3.0 is currently restricted to Alibaba’s cloud platform. This has led to discussions in the developer community regarding when or if open weights will be released.

Access and Integration

Alibaba has provided several endpoints for users to test Wan 3.0 during the public beta phase:

Web Platforms

  • Alibaba Cloud Model Studio: The primary interface for developer API integration and scaling.
  • Wanxiang Official Website: A web interface designed for direct prompt-based generation.
  • Qwen Creation PC: A web-based hub optimized for content generation workflows.

Mobile Applications

  • Qwen App: Currently rolling out to mobile users in select regions.

API Access

  • Full API access is rolling out via Alibaba Cloud, allowing developers to integrate document-to-video capabilities directly into third-party software.

Conclusion

Alibaba’s Wan 3.0 represents a step forward in generative AI, moving from short clip creation to document-driven video generation. Its ability to parse office file formats into consistent 30-second video segments, combined with competitive pricing, makes it a viable tool for creators, marketers, and educators. While certain technical challenges remain to be addressed in post-beta updates, Wan 3.0 establishes a competitive benchmark in the AI video landscape.

Related Articles

Comments

No Comments Yet

Be the first to comment on this content.