模型广场

选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。

重置筛选
全部 26文本 18图像 0音频 0视频 8

wan2.5-t2v-preview

文本音频视频

Alibaba

Generate 5-10 second, 480p/720p/1080p video from text prompts with native synced audio track.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent GenerationMultimodal Understanding

wan2.6-i2v

文本图像音频视频

Alibaba

Alibaba Tongyi Wanxiang Wan 2.6 image-to-video model animates a reference image into 720p/1080p clips with synchronized audio, preserving subject consistency with cinematic motion.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationMultimodal Understanding

wan2.6-i2v-flash

文本图像音频视频

Alibaba

Speed-optimized Wan 2.6 image-to-video variant with optional audio output, generating 720p/1080p clips quickly and cost-effectively for high-throughput image animation.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveReal-time ResponseCreative Generation

wan2.6-t2v

文本音频视频

Alibaba

Generate 720p/1080p video from text prompts and come with synchronized audio tracks. The picture has movie-level aesthetics and complex motion performance, suitable for creative video creation.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Creative GenerationHigh-Quality GenerationContent GenerationMultimodal Understanding

ernie-5.0

文本图像视频音频文本

baidu

A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.

上下文
128K
输入 / 1M tokens
$0.84
输出 / 1M tokens
$3.28
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningVisionFrontierMultimodal Understanding

seed-character

文本图像音频文本

bytedance

Character roleplay model for virtual companion scenarios, with stable persona adherence and plot progression in multi-turn dialogue, supporting multimodal inputs.

上下文
128K
输入 / 1M tokens
$0.12
输出 / 1M tokens
$0.29
缓存读 / 1M
$0.02
缓存写 / 1M
$0
ChatVisionChineseCreative Generation

doubao-seed-2.0-lite-260428

文本图像视频音频文本

bytedance

Lightweight tier of the Seed 2.0 family tuned for low-latency agent, coding, and GUI workloads.

上下文
256K
输入 / 1M tokens
$0.25
输出 / 1M tokens
$2
缓存读 / 1M
$0.05
缓存写 / 1M
$0.008333
Cost EffectiveCodingReal-time ResponseAgent

doubao-seed-2.0-mini-260428

文本图像视频音频文本

bytedance

Lightweight tier in the Seed 2.0 family, tuned for high-throughput, low-latency tasks with adjustable reasoning levels. Fits batch processing, moderation, and classification.

上下文
256K
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.02
缓存写 / 1M
$0.008333
Real-time ResponseLightweightClassificationBatch Generation

seedance-2.0

文本图像视频音频视频

bytedance

Standard tier of the Seedance family. Supports text-to-video, image-to-video, and reference-to-video, with strong character, style, and camera consistency. Built for ads and short-form creative.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
VisionFrontierCreative GenerationContent Generation

seedance-2.0-fast

文本图像视频音频视频

bytedance

Lightweight Seedance 2.0 variant tuned for speed and cost rather than peak fidelity, with text-to-video, first/last frame control, and multi-reference inputs.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveCreative GenerationContent GenerationBatch Generation

seedance-2.5

文本图像视频音频视频

bytedance

A video generation model built for long-form storytelling, with first-frame and first-and-last-frame control, up to 50 multimodal reference assets, and optional audio generation. It also supports video editing and extension.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
FrontierCreative GenerationHigh-Quality GenerationContent Generation

gemini-2.5-flash

文本图像视频音频文本

Google

A high-performance general-purpose model from Google, designed for advanced reasoning, coding, mathematics, and scientific tasks. Its built-in thinking capabilities improve response accuracy and enable deeper contextual understanding.

上下文
1.1M
输入 / 1M tokens
$0.3
输出 / 1M tokens
$2.5
缓存读 / 1M
$0.03
缓存写 / 1M
$0.08333
TechCodingTranslation

gemini-2.5-flash-lite

文本图像视频音频文本

Google

A lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency, with faster token generation and higher throughput. Thinking is disabled by default but can be enabled via API.

上下文
1.1M
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.01
缓存写 / 1M
$0.08333
FinanceCodingTranslation

gemini-2.5-pro

文本图像视频音频文本

Google

Massive context window. Great for video understanding.

上下文
1M
输入 / 1M tokens
$2.5
输出 / 1M tokens
$15
缓存读 / 1M
$0.25
缓存写 / 1M
$0
Vision

gemini-3-flash-preview

文本图像视频音频文本

Google

Speed-tier variant in the Gemini 3 line, pairing Pro-grade reasoning with Flash-level latency and cost, built for agent workflows and high-throughput interactive apps.

上下文
1M
输入 / 1M tokens
$0.5
输出 / 1M tokens
$3
缓存读 / 1M
$0.05
缓存写 / 1M
$0
CodingTranslationFinance

gemini-3.1-flash-lite

文本图像视频音频文本

Google

A high-efficiency multimodal lite model for low-latency, high-volume workloads like translation, classification, and data extraction, priced at about half of Gemini 3 Flash.

上下文
1M
输入 / 1M tokens
$0.25
输出 / 1M tokens
$1.5
缓存读 / 1M
$0.025
缓存写 / 1M
$0.08333
Instruction FollowingTask AutomationStructured OutputClassification

gemini-3.1-pro-preview

文本图像视频音频文本

Google

Google's latest. Strong reasoning with native image and video.

上下文
1M
输入 / 1M tokens
$2
输出 / 1M tokens
$12
缓存读 / 1M
$0.2
缓存写 / 1M
$0
Vision

gemini-3.5-flash

文本图像视频音频文本

Google

Designed for efficient multimodal AI tasks, offering strong coding, reasoning, real-time chat, and agent execution at Flash-tier cost and speed.

上下文
1M
输入 / 1M tokens
$1.5
输出 / 1M tokens
$9
缓存读 / 1M
$0.15
缓存写 / 1M
$0.08333
Production CodeRefactoringAgent CodingDaily Dev

gemma-3n-e4b-it:free

文本图像视频音频文本

Google

Fast and cost-efficient.

上下文
33K
输入 / 1M tokens
$0.06
输出 / 1M tokens
$0.12
缓存读 / 1M
$0
缓存写 / 1M
$0
ChatReal-time ResponseClassification

minimax-h3

文本图像视频音频视频

minimax

A lightweight open-weight video generation model focused on instruction-guided editing, text and brand rendering, and video-to-video motion transfer, with native audiovisual output. Built for advertising, e-commerce, and interface design workflows.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
LightweightCreative GenerationContent GenerationMarketing

gpt-4o-mini-transcribe

文本音频文本

OpenAI

GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o Mini audio capabilities.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveSpeech-to-Text

gpt-4o-transcribe

文本音频文本

OpenAI

GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o audio capabilities.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Speech-to-Text

gpt-4o-transcribe-diarize

音频文本文本

OpenAI

GPT-4o Transcribe Diarize is an automatic speech recognition (ASR) model with built-in speaker diarization, meaning it associates audio segments with different speakers in a conversation.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Speech-to-Text

gpt-5.5-beta

文本图像音频文本

OpenAI

GPT-5.5 natively supports video comprehension, speech emotion recognition, and autonomous tool calling, delivering omni-modal foundational capabilities for agent-based execution of complex tasks.

上下文
1.1M
输入 / 1M tokens
$10
输出 / 1M tokens
$45
缓存读 / 1M
$1
缓存写 / 1M
$0
CodingProduction CodeInstruction Following

stt-whisper-1

音频文本

OpenAI

Whisper is OpenAI's open-source automatic speech recognition model. It supports transcription and translation across 50+ languages from audio files up to 25 MB. Accepts formats including mp3, mp4, wav, and webm.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
MultilingualSpeech-to-Text

mimo-v2.5

文本图像音频视频文本

xiaomi

Mid-tier V2.5 model with native image, audio, and video understanding, built for agent workflows and coding at lower cost than Pro.

上下文
1M
输入 / 1M tokens
$0.14
输出 / 1M tokens
$0.28
缓存读 / 1M
$0.0428
缓存写 / 1M
$0
CodingMultimodal UnderstandingAgent CodingAgent
需要完整模型清单或企业接入方案?进入控制台