Integrate Alibaba's comprehensive model family through a single unified API endpoint. Access both the Qwen suite for state-of-the-art language and multimodal vision capabilities, and the Wan model series for high-definition video generation up to 1080p.
Alibaba
All Models
Wan 3.0
12 variants available

Wan 3.0 Prime Reference-to-Video Spicy
Wan 3.0 Prime Reference-to-Video Spicy is Alibaba's flagship uncensored high-motion video model engineered for native 4K visual synthesis. Delivering native 4K output at 60 fps, it processes up to 8 reference images or 2 reference videos simultaneously to preserve character identity and material fidelity across complex sequences. With unrestricted generation capabilities, it synthesizes explosive physical movements, dramatic camera trajectories, and physical interactions—powering high-energy cinematic pre-visualization, high-impact commercial ads, AAA gaming assets, and action pipelines.

Wan 3.0 Reference-to-Video Spicy
Wan 3.0 Reference-to-Video Spicy is Alibaba's high-motion video synthesis model engineered for high-energy visual dynamics and stylized action generation. Built on an expanded spatio-temporal attention mechanism alongside robust reference feature anchoring, it synthesizes large-scale physical movements, aggressive camera trajectories, and dramatic temporal transitions while preserving strict subject identity and garment texture fidelity. Wan 3.0 Reference-to-Video Spicy powers dynamic action video production, cinematic FX pre-visualization, interactive gaming visual assets, and high-impact commercial advertising pipelines.

Wan 3.0 Prime Reference-to-Video
Wan 3.0 Prime Reference-to-Video is Alibaba's flagship controllable video generation model built for precise reference-conditioned visual synthesis. Powered by an upgraded spatio-temporal decoupling architecture and multi-reference feature fusion, it preserves character identity, garment textures, ambient lighting, and complex camera trajectories across extended sequences while generating native 4K high-frame-rate video. Operating with strict temporal consistency and fluid motion dynamics, Wan 3.0 Prime drives e-commerce video production, cinematic pre-visualization, digital human animation, and commercial advertising pipelines.
Qwen
8 variants available

Qwen Image 3.0 Spicy
Qwen Image 3.0 Spicy is a flagship uncensored image generation model built upon Alibaba's advanced Qwen vision architecture. Operating as a Spicy uncensored model, it completely bypasses standard content filters to unlock maximum creative freedom across photorealistic portraiture, complex visual storytelling, and high-detail stylized concept art without content restrictions. Leveraging enhanced rendering capacities, it delivers sharp textures, realistic lighting physics, and exceptional prompt adherence. Qwen Image 3.0 Spicy is the premier choice for professional designers, digital creators, marketing agencies, and visual artists requiring unconstrained, high-quality image synthesis.

Qwen Image 3.0 Pro Spicy
Qwen Image 3.0 Pro Spicy is a flagship uncensored image generation model built upon Alibaba's advanced Qwen vision architecture. Operating as a Spicy uncensored model, it completely bypasses standard content filters to unlock unrestricted artistic freedom across photorealistic portraiture, complex visual storytelling, and high-detail stylized concept art. Leveraging Pro-tier enhancements, it delivers ultra-sharp textures, master-level lighting physics, fine typography rendering inside images, and exceptional prompt adherence. Qwen Image 3.0 Pro Spicy is the premier choice for professional designers, digital creators, marketing agencies, and visual artists requiring unconstrained, studio-quality image synthesis.

Qwen Image 3.0
Qwen Image 3.0 is Alibaba's next-generation open-weights AI image generation and editing model built on a Diffusion Transformer (DiT) architecture. Supporting both text-to-image (T2I) and image-to-image instruction editing (I2I), it excels at rich content rendering, hyper-realistic detail, and deep contextual understanding. It achieves industry-leading text-in-image typography, rendering clear small text down to 10px across 12 native languages and 20+ fonts for complex multi-column layouts, UI designs, and graphic documents. With open weights available for self-hosting, Qwen Image 3.0 provides highly cost-effective, high-precision visual generation for localized advertising, product UI prototyping, e-commerce graphics, and creative publishing.
Happy Horse
6 variants available

HappyHorse 1.1 Reference-to-Video Spicy
HappyHorse 1.1 Reference-to-Video Spicy is Alibaba's high-expressiveness, high-dynamism variant in the HappyHorse model series. Building upon multi-image subject locking and native audio-visual synchronization, the Spicy mode is a uncensored and fine-tuned for high-intensity action, aggressive camera tracking, strong visual impact, and dramatic visual effects. It leverages reference images to drive bold, highly expressive motion sequences and complex camera maneuvers while maintaining multilingual lip-sync and ambient audio—ideal for action-packed short dramas, viral social video, game CG FX, and high-impact commercial ads.

HappyHorse 1.1 Reference-to-Video
HappyHorse 1.1 Reference-to-Video is Alibaba's next-generation AI reference-guided video generation model. Supporting 1 to 9 reference images (covering character identity, outfits, product silhouette, and scene styles), it achieves high-precision multi-image fusion and subject locking to prevent visual drift across shots. It generates native 720P/1080P HD videos up to 15 seconds long per run.

HappyHorse 1.1 Text-to-Video
HappyHorse 1.1 Text-to-Video is Alibaba's next-generation AI text-to-video model. Built on an integrated audio-video joint generation architecture, it generates native 720P/1080P HD videos directly from text, supporting up to 15 seconds of rendering. It natively supports multilingual lip-sync, ambient sound effects, and audio-visual synchronization without extra dubbing, while delivering exceptional motion smoothness, subject consistency, and camera control—ideal for short dramas, commercial ads, and social media video production.
Wan 2.7
4 variants available

Wan 2.7 Image-to-Video
Tongyi Wanxiang Wan 2.7 Image-to-Video is Alibaba Cloud's next-generation image-to-video model. It supports first-frame generation, first-to-last frame transitions, and short video extension, generating 720P/1080P HD videos up to 15 seconds per run. Featuring powerful camera control and realistic physics simulation, it natively supports audio-driven lip-sync and action alignment while adaptively supporting mainstream aspect ratios like 16:9 and 9:16—ideal for e-commerce, VFX, and film post-production.

Wan 2.7 Text-to-Video
Tongyi Wanxiang Wan 2.7 Text-to-Video is Alibaba Cloud's next-generation text-to-video model. Supporting long Chinese and English prompts with smart expansion, it generates native 720P/1080P HD videos directly from text, up to 15 seconds per run. Featuring native audio-visual coordination and sound effect sync, it adaptively supports multiple aspect ratios like 16:9 and 9:16 with strong motion continuity, realistic physics simulation, and lighting rendering—ideal for commercial ads, short videos, and anime creation.

Wan 2.7 Text-to-Video Spicy
Wan 2.7 Text-to-Video Spicy turns simple text prompts into short cinematic clips, blending impressive temporal stability with expressive, nuanced character movement.
Wan 2.6
3 variants available

Wan 2.6 Reference-to-Video Spicy
Wan 2.6 Reference-to-Video Spicy is Alibaba's advanced uncensored model designed for high-fidelity, reference-guided video synthesis. Operating as a Spicy uncensored model, it completely bypasses standard safety guardrails to perform dynamic character replication, motion transfer, and style transformation without content restrictions. By accurately tracking facial features, body postures, and spatial trajectories from reference images or video clips, it generates fluid, high-amplitude action sequences while preserving key visual identity elements. Wan 2.6 Reference-to-Video Spicy efficiently powers unrestricted choreography, motion transfer pre-visualization, digital human performances, and high-impact cinematic visual effects.

Wan 2.6 Text-to-Video Spicy
Wan 2.6 Text-to-Video Spicy is an uncensored model built for high-energy video creation. As a Spicy uncensored model, it breaks standard limits to render fluid, complex physical interactions and cinematic scenes directly from text. Perfect for VFX, action concept art, and creative media.

Wan 2.6 Image-to-Video Spicy
Wan 2.6 Image-to-Video Spicy turns a first-frame image into short cinematic motion with stable temporal detail and expressive character movement.
Wan 2.2
2 variants available

Wan 2.2 Text-to-Video Spicy
Wan 2.2 Text-to-Video Spicy is an uncensored model built for high-energy video synthesis. As a Spicy uncensored model, it bypasses standard restrictions to render fluid physical movements, dramatic action, and cinematic scenes from text prompts. Ideal for action scenes and creative media.

Wan 2.2 Image-to-Video Spicy
Wan 2.2 Image-to-Video Spicy is an uncensored model built for high-energy video creation from images. As a Spicy uncensored model, it breaks standard limits to render fluid physical movement and dramatic action sequences while preserving image identity. Perfect for VFX and creative media.
Alibaba Models API Pricing Details
| Model | Pricing (USD) | Our Pricing (USD) | Discount | |
|---|---|---|---|---|
| Qwen Image 3.0 Spicy | $0.03/PIC | Start from$0.03/PIC | — | |
| Qwen Image 3.0 Pro Spicy | $0.04/PIC | Start from$0.04/PIC | — | |
| Wan 2.6 Reference-to-Video Spicy | $0.1/SEC | Start from$0.1/SEC | — | |
| Wan 2.6 Text-to-Video Spicy | $0.1/SEC | Start from$0.1/SEC | — | |
| Wan 2.2 Text-to-Video Spicy | $0.02/SEC | Start from$0.02/SEC | — | |
| Wan 2.2 Image-to-Video Spicy | $0.02/SEC | Start from$0.02/SEC | — | |
| Wan 2.6 Image-to-Video Spicy | $0.1/SEC | Start from$0.1/SEC | — | |
| Wan 3.0 Prime Reference-to-Video Spicy | $0.068/SEC | Start from$0.054/SEC | -20% | |
| Wan 3.0 Reference-to-Video Spicy | $0.05/SEC | Start from$0.035/SEC | -30% | |
| Wan 3.0 Prime Reference-to-Video | $0.068/SEC | Start from$0.054/SEC | -20% | |
| Wan 3.0 Reference-to-Video | $0.05/SEC | Start from$0.035/SEC | -30% | |
| Wan 3.0 Prime Image-to-Video | $0.068/SEC | Start from$0.054/SEC | -20% | |
| Wan 3.0 Prime Text-to-Video | $0.068/SEC | Start from$0.054/SEC | -20% | |
| Qwen Image 3.0 | $0.03/PIC | Start from$0.03/PIC | — | |
| Qwen Image 3.0 Pro | $0.04/PIC | Start from$0.04/PIC | — | |
| Wan 3.0 Image-to-Video Spicy | $0.05/SEC | Start from$0.035/SEC | -30% | |
| Wan 3.0 Prime Image-to-Video Spicy | $0.068/SEC | Start from$0.054/SEC | -20% | |
| Wan 3.0 Prime Text-to-Video Spicy | $0.068/SEC | Start from$0.054/SEC | -20% | |
| Wan 3.0 Text-to-Video Spicy | $0.05/SEC | Start from$0.035/SEC | -30% | |
| Wan 3.0 Text-to-Video | $0.05/SEC | Start from$0.035/SEC | -30% | |
| Wan 3.0 Image-to-Video | $0.05/SEC | Start from$0.035/SEC | -30% | |
| Qwen 3.8 Max | $2/M Tokens | Start from$2/M Tokens | — | |
| HappyHorse 1.1 Reference-to-Video Spicy | $0.14/SEC | Start from$0.084/SEC | -40% | |
| HappyHorse 1.1 Reference-to-Video | $0.14/SEC | Start from$0.084/SEC | -40% | |
| Qwen 3.5 Omni Plus | $7.25/M Tokens | Start from$7.25/M Tokens | — | |
| Wan 2.7 Image-to-Video | $0.1/SEC | Start from$0.1/SEC | — | |
| Wan 2.7 Text-to-Video | $0.1/SEC | Start from$0.1/SEC | — | |
| HappyHorse 1.1 Text-to-Video | $0.14/SEC | Start from$0.084/SEC | -40% | |
| HappyHorse 1.1 Image-to-Video | $0.14/SEC | Start from$0.084/SEC | -40% | |
| Qwen 3.7 Plus | $0.4/M Tokens | Start from$0.4/M Tokens | — | |
| Qwen 3.7 Max | $2.5/M Tokens | Start from$2.5/M Tokens | — | |
| HappyHorse 1.1 Image-to-Video Spicy | $0.14/SEC | Start from$0.084/SEC | -40% | |
| HappyHorse 1.1 Text-to-Video Spicy | $0.14/SEC | Start from$0.084/SEC | -40% | |
| Wan 2.7 Image-to-Video Spicy | $0.1/SEC | Start from$0.1/SEC | — | |
| Wan 2.7 Text-to-Video Spicy | $0.1/SEC | Start from$0.1/SEC | — |