模型广场

选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。

重置筛选
全部 8文本 0图像 0音频 0视频 8

wan2.6-r2v

文本视频图像视频

Alibaba

Alibaba Tongyi Wanxiang Wan 2.6 reference-to-video model replicates the actions, effects, and camera movement of a reference video to produce up to 10 second 720p/1080p clips with synchronized audio.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationMultimodal Understanding

wan2.6-r2v-flash

文本视频图像视频

Alibaba

Speed-optimized Wan 2.6 reference-to-video variant with optional audio output, replicating reference-video motion into up to 10 second 720p/1080p clips for fast, cost-effective production.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveReal-time ResponseCreative Generation

seedance-2.0

文本图像视频音频视频

bytedance

Standard tier of the Seedance family. Supports text-to-video, image-to-video, and reference-to-video, with strong character, style, and camera consistency. Built for ads and short-form creative.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
VisionFrontierCreative GenerationContent Generation

seedance-2.0-fast

文本图像视频音频视频

bytedance

Lightweight Seedance 2.0 variant tuned for speed and cost rather than peak fidelity, with text-to-video, first/last frame control, and multi-reference inputs.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveCreative GenerationContent GenerationBatch Generation

seedance-2.5

文本图像视频音频视频

bytedance

A video generation model built for long-form storytelling, with first-frame and first-and-last-frame control, up to 50 multimodal reference assets, and optional audio generation. It also supports video editing and extension.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
FrontierCreative GenerationHigh-Quality GenerationContent Generation

kling-v3-omni

文本图像视频视频

kling

Fully multimodal variant of the 3.0 line, accepting text, image, and video references in a single generation, with voice-driven characters and native audio.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent GenerationMultimodal Understanding

kling-o1

文本图像视频视频

kling

Cinematic-oriented model with first- and last-frame control and video reference input, including edits driven by an existing clip. Suited for narrative shot sequences.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent GenerationMultimodal Understanding

minimax-h3

文本图像视频音频视频

minimax

A lightweight open-weight video generation model focused on instruction-guided editing, text and brand rendering, and video-to-video motion transfer, with native audiovisual output. Built for advertising, e-commerce, and interface design workflows.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
LightweightCreative GenerationContent GenerationMarketing
需要完整模型清单或企业接入方案?进入控制台