模型广场

选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。

重置筛选
全部 18文本 18图像 0音频 0视频 0

ernie-5.0

文本图像视频音频文本

baidu

A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.

上下文
128K
输入 / 1M tokens
$0.84
输出 / 1M tokens
$3.28
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningVisionFrontierMultimodal Understanding

seed-character

文本图像音频文本

bytedance

Character roleplay model for virtual companion scenarios, with stable persona adherence and plot progression in multi-turn dialogue, supporting multimodal inputs.

上下文
128K
输入 / 1M tokens
$0.12
输出 / 1M tokens
$0.29
缓存读 / 1M
$0.02
缓存写 / 1M
$0
ChatVisionChineseCreative Generation

doubao-seed-2.0-lite-260428

文本图像视频音频文本

bytedance

Lightweight tier of the Seed 2.0 family tuned for low-latency agent, coding, and GUI workloads.

上下文
256K
输入 / 1M tokens
$0.25
输出 / 1M tokens
$2
缓存读 / 1M
$0.05
缓存写 / 1M
$0.008333
Cost EffectiveCodingReal-time ResponseAgent

doubao-seed-2.0-mini-260428

文本图像视频音频文本

bytedance

Lightweight tier in the Seed 2.0 family, tuned for high-throughput, low-latency tasks with adjustable reasoning levels. Fits batch processing, moderation, and classification.

上下文
256K
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.02
缓存写 / 1M
$0.008333
Real-time ResponseLightweightClassificationBatch Generation

gemini-2.5-flash

文本图像视频音频文本

Google

A high-performance general-purpose model from Google, designed for advanced reasoning, coding, mathematics, and scientific tasks. Its built-in thinking capabilities improve response accuracy and enable deeper contextual understanding.

上下文
1.1M
输入 / 1M tokens
$0.3
输出 / 1M tokens
$2.5
缓存读 / 1M
$0.03
缓存写 / 1M
$0.08333
TechCodingTranslation

gemini-2.5-flash-lite

文本图像视频音频文本

Google

A lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency, with faster token generation and higher throughput. Thinking is disabled by default but can be enabled via API.

上下文
1.1M
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.01
缓存写 / 1M
$0.08333
FinanceCodingTranslation

gemini-2.5-pro

文本图像视频音频文本

Google

Massive context window. Great for video understanding.

上下文
1M
输入 / 1M tokens
$2.5
输出 / 1M tokens
$15
缓存读 / 1M
$0.25
缓存写 / 1M
$0
Vision

gemini-3-flash-preview

文本图像视频音频文本

Google

Speed-tier variant in the Gemini 3 line, pairing Pro-grade reasoning with Flash-level latency and cost, built for agent workflows and high-throughput interactive apps.

上下文
1M
输入 / 1M tokens
$0.5
输出 / 1M tokens
$3
缓存读 / 1M
$0.05
缓存写 / 1M
$0
CodingTranslationFinance

gemini-3.1-flash-lite

文本图像视频音频文本

Google

A high-efficiency multimodal lite model for low-latency, high-volume workloads like translation, classification, and data extraction, priced at about half of Gemini 3 Flash.

上下文
1M
输入 / 1M tokens
$0.25
输出 / 1M tokens
$1.5
缓存读 / 1M
$0.025
缓存写 / 1M
$0.08333
Instruction FollowingTask AutomationStructured OutputClassification

gemini-3.1-pro-preview

文本图像视频音频文本

Google

Google's latest. Strong reasoning with native image and video.

上下文
1M
输入 / 1M tokens
$2
输出 / 1M tokens
$12
缓存读 / 1M
$0.2
缓存写 / 1M
$0
Vision

gemini-3.5-flash

文本图像视频音频文本

Google

Designed for efficient multimodal AI tasks, offering strong coding, reasoning, real-time chat, and agent execution at Flash-tier cost and speed.

上下文
1M
输入 / 1M tokens
$1.5
输出 / 1M tokens
$9
缓存读 / 1M
$0.15
缓存写 / 1M
$0.08333
Production CodeRefactoringAgent CodingDaily Dev

gemma-3n-e4b-it:free

文本图像视频音频文本

Google

Fast and cost-efficient.

上下文
33K
输入 / 1M tokens
$0.06
输出 / 1M tokens
$0.12
缓存读 / 1M
$0
缓存写 / 1M
$0
ChatReal-time ResponseClassification

gpt-4o-mini-transcribe

文本音频文本

OpenAI

GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o Mini audio capabilities.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Cost EffectiveSpeech-to-Text

gpt-4o-transcribe

文本音频文本

OpenAI

GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o audio capabilities.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Speech-to-Text

gpt-4o-transcribe-diarize

音频文本文本

OpenAI

GPT-4o Transcribe Diarize is an automatic speech recognition (ASR) model with built-in speaker diarization, meaning it associates audio segments with different speakers in a conversation.

上下文
131K
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
Speech-to-Text

gpt-5.5-beta

文本图像音频文本

OpenAI

GPT-5.5 natively supports video comprehension, speech emotion recognition, and autonomous tool calling, delivering omni-modal foundational capabilities for agent-based execution of complex tasks.

上下文
1.1M
输入 / 1M tokens
$10
输出 / 1M tokens
$45
缓存读 / 1M
$1
缓存写 / 1M
$0
CodingProduction CodeInstruction Following

stt-whisper-1

音频文本

OpenAI

Whisper is OpenAI's open-source automatic speech recognition model. It supports transcription and translation across 50+ languages from audio files up to 25 MB. Accepts formats including mp3, mp4, wav, and webm.

上下文
待核
输入 / 1M tokens
-
输出 / 1M tokens
-
缓存读 / 1M
-
缓存写 / 1M
-
MultilingualSpeech-to-Text

mimo-v2.5

文本图像音频视频文本

xiaomi

Mid-tier V2.5 model with native image, audio, and video understanding, built for agent workflows and coding at lower cost than Pro.

上下文
1M
输入 / 1M tokens
$0.14
输出 / 1M tokens
$0.28
缓存读 / 1M
$0.0428
缓存写 / 1M
$0
CodingMultimodal UnderstandingAgent CodingAgent
需要完整模型清单或企业接入方案?进入控制台