模型广场

选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。

重置筛选
全部 27文本 0图像 0音频 0视频 27

happyhorse-1.0-i2v

图像视频

Alibaba

Extends a single starting image into a full shot, preserving its composition and style, and outputs 3–15 second clips at up to 1080p. Suited for turning still assets into motion.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent Generation

happyhorse-1.0-r2v

图像视频

Alibaba

Takes a set of reference images to lock character, scene, and style, holding the subject consistent across shots. Outputs 3–15 second clips at up to 1080p.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent Generation

happyhorse-1.1-i2v

图像视频

Alibaba

Smoother motion than 1.0 across camera moves and body movement. Extends a single starting image into a full shot, outputting 3–15 second clips at up to 1080p.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent Generation

happyhorse-1.1-r2v

图像视频

Alibaba

Holds characters more consistent across frames than 1.0. Takes a set of reference images to lock character, scene, and style, outputting 3–15 second clips at up to 1080p.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent Generation

wan2.6-i2v

文本图像音频视频

Alibaba

Alibaba Tongyi Wanxiang Wan 2.6 image-to-video model animates a reference image into 720p/1080p clips with synchronized audio, preserving subject consistency with cinematic motion.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationMultimodal Understanding

wan2.6-i2v-flash

文本图像音频视频

Alibaba

Speed-optimized Wan 2.6 image-to-video variant with optional audio output, generating 720p/1080p clips quickly and cost-effectively for high-throughput image animation.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveReal-time ResponseCreative Generation

wan2.6-r2v

文本视频图像视频

Alibaba

Alibaba Tongyi Wanxiang Wan 2.6 reference-to-video model replicates the actions, effects, and camera movement of a reference video to produce up to 10 second 720p/1080p clips with synchronized audio.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationMultimodal Understanding

wan2.6-r2v-flash

文本视频图像视频

Alibaba

Speed-optimized Wan 2.6 reference-to-video variant with optional audio output, replicating reference-video motion into up to 10 second 720p/1080p clips for fast, cost-effective production.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveReal-time ResponseCreative Generation

seedance-2.0

文本图像视频音频视频

bytedance

Standard tier of the Seedance family. Supports text-to-video, image-to-video, and reference-to-video, with strong character, style, and camera consistency. Built for ads and short-form creative.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
VisionFrontierCreative GenerationContent Generation

seedance-2.0-fast

文本图像视频音频视频

bytedance

Lightweight Seedance 2.0 variant tuned for speed and cost rather than peak fidelity, with text-to-video, first/last frame control, and multi-reference inputs.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveCreative GenerationContent GenerationBatch Generation

seedance-2.5

文本图像视频音频视频

bytedance

A video generation model built for long-form storytelling, with first-frame and first-and-last-frame control, up to 50 multimodal reference assets, and optional audio generation. It also supports video editing and extension.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
FrontierCreative GenerationHigh-Quality GenerationContent Generation

kling-v1-5

图像视频

kling

The oldest generation still available, limited to image-to-video with first- and last-frame control and no multi-image reference. Suited for basic, cost-sensitive clip generation.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveCreative GenerationContent Generation

kling-v1-6

文本图像视频

kling

Kling 1.6 supports text, first or last frame, and multi-reference-image video generation.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveCreative GenerationContent Generation

kling-v2-master

文本图像视频

kling

Flagship of the 2.0 line, tuned for instruction following and cinematic aesthetics, with stylized effect presets. Outputs 1080p only.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationContent GenerationHigh-Quality Generation

kling-v2-1

图像文本视频

kling

Kling 2.1 image-to-video generation supports first and last frame control.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationCost EffectiveContent Generation

kling-v2-1-master

文本图像视频

kling

Quality-oriented tier of the 2.1 line, tuned for motion performance and semantic responsiveness. Supports text-to-video and image-to-video, suited for single-shot clips.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
High-Quality GenerationContent GenerationCreative Generation

kling-v2-5-turbo

文本图像视频

kling

Cost-efficient tier in the 2.5 generation, with improved prompt adherence and physics under large-amplitude motion. Outputs silent 5- or 10-second clips.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveCreative GenerationContent Generation

kling-2.6

文本图像视频

kling

First version to add native audio, producing visuals, voiceover, sound effects, and ambience in one pass. Suited for finished clips up to 10 seconds.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent Generation

kling-3.0

文本图像视频

kling

Core model of the 3.0 line, with first- and last-frame control, multi-shot storytelling, and native audio produced in the same pass. Suited for finished short-form work.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent Generation

kling-v3-omni

文本图像视频视频

kling

Fully multimodal variant of the 3.0 line, accepting text, image, and video references in a single generation, with voice-driven characters and native audio.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent GenerationMultimodal Understanding

kling-3.0-turbo

文本图像视频

kling

Speed-and-cost tier of the 3.0 line, covering text-to-video and first-frame image-to-video with multi-shot output. Suited for high-volume production.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationContent GenerationCost Effective

kling-o1

文本图像视频视频

kling

Cinematic-oriented model with first- and last-frame control and video reference input, including edits driven by an existing clip. Suited for narrative shot sequences.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent GenerationMultimodal Understanding

hailuo-02

文本图像视频

minimax

It supports Vincentian and Tusheng videos, can output 6-10 seconds, 512p/768p/1080p multi-level image quality, supports camera movement command control, and flexibly balances image quality and cost.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveCreative GenerationContent GenerationMultimodal Understanding

hailuo-2.3

文本图像视频

minimax

It supports Vincentian and Tusheng videos, can output 6-10 seconds, 768p/1080p images, supports fine camera movement command control, has smooth movement and stable subject, and is suitable for short video creative content.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent GenerationMultimodal Understanding

hailuo-2.3-fast

文本图像视频

minimax

Speed and cost optimized variant of MiniMax Hailuo 2.3 for rapid iteration and bulk production, generating 6 to 10 second 768p/1080p clips from text or image input with camera-movement control.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationCost Effective

minimax-h3

文本图像视频音频视频

minimax

A lightweight open-weight video generation model focused on instruction-guided editing, text and brand rendering, and video-to-video motion transfer, with native audiovisual output. Built for advertising, e-commerce, and interface design workflows.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
LightweightCreative GenerationContent GenerationMarketing

sora-2

文本图像视频

OpenAI

It can generate detailed, dynamic 720p videos from text or reference images with synchronized audio tracks. It has excellent understanding of three-dimensional space, motion and lens continuity, and is suitable for the creation of 16-20 second cinematic short films.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
FrontierCreative GenerationHigh-Quality GenerationMultimodal Understanding
需要完整模型清单或企业接入方案?进入控制台