模型广场

选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。

重置筛选
全部 35文本 35图像 0音频 0视频 0

ernie-5.0

文本图像视频音频文本

baidu

A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.

上下文
128K
输入 / 1M tokens
$0.84
输出 / 1M tokens
$3.28
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningVisionFrontierMultimodal Understanding

seed-2-1-turbo

文本图像视频文本

bytedance

Turbo tier in Seed 2.1: multimodal model for coding and long-horizon agents, suited to multi-step task execution and visual content understanding.

上下文
262K
输入 / 1M tokens
$0.5
输出 / 1M tokens
$2.5
缓存读 / 1M
$0.1
缓存写 / 1M
$0.008
VisionCost EffectiveCodingMultimodal Understanding

seed-1.8

文本图像视频文本

bytedance

Agent-oriented foundation model tuned for tool calling and complex instruction following, with use cases spanning GUI agents, search agents, and multi-step task orchestration.

上下文
256K
输入 / 1M tokens
$0.5
输出 / 1M tokens
$4
缓存读 / 1M
$0.05
缓存写 / 1M
$0.008333
Instruction FollowingTask AutomationComplex PlanningAgent

doubao-seed-2.0-code-preview-260215

文本图像视频文本

bytedance

Code-focused preview model for agent workflows, with competitive SWE-Bench and LiveCodeBench scores, aimed at autonomous coding agents and multi-language software tasks.

上下文
256K
输入 / 1M tokens
$0.5
输出 / 1M tokens
$3
缓存读 / 1M
$0.1
缓存写 / 1M
$0.008333
CodingProduction CodeRefactoringAgent Coding

doubao-seed-2.0-lite-260428

文本图像视频音频文本

bytedance

Lightweight tier of the Seed 2.0 family tuned for low-latency agent, coding, and GUI workloads.

上下文
256K
输入 / 1M tokens
$0.25
输出 / 1M tokens
$2
缓存读 / 1M
$0.05
缓存写 / 1M
$0.008333
Cost EffectiveCodingReal-time ResponseAgent

doubao-seed-2.0-mini-260428

文本图像视频音频文本

bytedance

Lightweight tier in the Seed 2.0 family, tuned for high-throughput, low-latency tasks with adjustable reasoning levels. Fits batch processing, moderation, and classification.

上下文
256K
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.02
缓存写 / 1M
$0.008333
Real-time ResponseLightweightClassificationBatch Generation

doubao-seed-2.0-pro-260215

文本图像视频文本

bytedance

Flagship of the Seed 2.0 series for complex reasoning and long-horizon agent workflows with reliable tool use.

上下文
256K
输入 / 1M tokens
$0.5
输出 / 1M tokens
$3
缓存读 / 1M
$0.1
缓存写 / 1M
$0.008333
ReasoningInstruction FollowingComplex PlanningAgent

gemini-2.5-flash

文本图像视频音频文本

Google

A high-performance general-purpose model from Google, designed for advanced reasoning, coding, mathematics, and scientific tasks. Its built-in thinking capabilities improve response accuracy and enable deeper contextual understanding.

上下文
1.1M
输入 / 1M tokens
$0.3
输出 / 1M tokens
$2.5
缓存读 / 1M
$0.03
缓存写 / 1M
$0.08333
TechCodingTranslation

gemini-2.5-flash-lite

文本图像视频音频文本

Google

A lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency, with faster token generation and higher throughput. Thinking is disabled by default but can be enabled via API.

上下文
1.1M
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.01
缓存写 / 1M
$0.08333
FinanceCodingTranslation

gemini-2.5-pro

文本图像视频音频文本

Google

Massive context window. Great for video understanding.

上下文
1M
输入 / 1M tokens
$2.5
输出 / 1M tokens
$15
缓存读 / 1M
$0.25
缓存写 / 1M
$0
Vision

gemini-3-flash-preview

文本图像视频音频文本

Google

Speed-tier variant in the Gemini 3 line, pairing Pro-grade reasoning with Flash-level latency and cost, built for agent workflows and high-throughput interactive apps.

上下文
1M
输入 / 1M tokens
$0.5
输出 / 1M tokens
$3
缓存读 / 1M
$0.05
缓存写 / 1M
$0
CodingTranslationFinance

gemini-3.1-flash-lite

文本图像视频音频文本

Google

A high-efficiency multimodal lite model for low-latency, high-volume workloads like translation, classification, and data extraction, priced at about half of Gemini 3 Flash.

上下文
1M
输入 / 1M tokens
$0.25
输出 / 1M tokens
$1.5
缓存读 / 1M
$0.025
缓存写 / 1M
$0.08333
Instruction FollowingTask AutomationStructured OutputClassification

gemini-3.1-pro-preview

文本图像视频音频文本

Google

Google's latest. Strong reasoning with native image and video.

上下文
1M
输入 / 1M tokens
$2
输出 / 1M tokens
$12
缓存读 / 1M
$0.2
缓存写 / 1M
$0
Vision

gemini-3.5-flash

文本图像视频音频文本

Google

Designed for efficient multimodal AI tasks, offering strong coding, reasoning, real-time chat, and agent execution at Flash-tier cost and speed.

上下文
1M
输入 / 1M tokens
$1.5
输出 / 1M tokens
$9
缓存读 / 1M
$0.15
缓存写 / 1M
$0.08333
Production CodeRefactoringAgent CodingDaily Dev

gemma-3n-e4b-it:free

文本图像视频音频文本

Google

Fast and cost-efficient.

上下文
33K
输入 / 1M tokens
$0.06
输出 / 1M tokens
$0.12
缓存读 / 1M
$0
缓存写 / 1M
$0
ChatReal-time ResponseClassification

minimax-m3

文本图像视频文本

minimax

MiniMax\\'s first multimodal release, accepting image and video input. Long-context inference runs cheaper and faster, and it handles multi-step agent workflows and computer-use scenarios.

上下文
1M
输入 / 1M tokens
$1.2
输出 / 1M tokens
$4.8
缓存读 / 1M
$0.12
缓存写 / 1M
$0
VisionLong Text ProcessingMultimodal UnderstandingAgent Coding

kimi-k2.5

文本图像视频文本

Moonshot

Top Chinese-English bilingual. Strong long context.

上下文
262K
输入 / 1M tokens
$0.6
输出 / 1M tokens
$3.011
缓存读 / 1M
$0.15
缓存写 / 1M
$0.718
ChineseBilingual

kimi-k2.6

文本图像视频文本

Moonshot

Latest multimodal model in the K2 series, built for long-horizon coding, code-driven UI/UX generation, and multi-agent orchestration.

上下文
256K
输入 / 1M tokens
$0.8939
输出 / 1M tokens
$3.7131
缓存读 / 1M
$0.34
缓存写 / 1M
$1.1174
Vision

kimi-k2.7-code

文本视频图像文本

Moonshot

Coding-focused variant in the Kimi K2 family. Accepts text and image inputs, with reasoning on by default. Suited for long-horizon coding and agentic workflows.

上下文
262K
输入 / 1M tokens
$0.95
输出 / 1M tokens
$4
缓存读 / 1M
$0.19
缓存写 / 1M
$1.1174
CodingAgent CodingChain of Thought

qwen3-vl-8b-instruct

文本图像视频文本

qwen

Lightweight multimodal model with balanced performance.

上下文
128K
输入 / 1M tokens
$0.25
输出 / 1M tokens
$0.75
缓存读 / 1M
$0.12
缓存写 / 1M
$0
Visual QAUI UnderstandingImage Parsing

qwen3-vl-plus

文本图像视频文本

qwen

Plus-tier vision-language model in the Qwen3-VL line, handling text, image and video inputs with strengths in document parsing, video understanding, spatial grounding and agent tool use.

上下文
262K
输入 / 1M tokens
$0.144
输出 / 1M tokens
$1.434
缓存读 / 1M
$0.029
缓存写 / 1M
$0.539
VisionUI UnderstandingDocument OCRMultimodal Understanding

qwen3.5-122b-a10b

文本图像视频文本

qwen

Open-weight mid-tier MoE text model with long chain-of-thought reasoning, strong on function-calling and multi-step planning benchmarks for self-hosted agentic workflows.

上下文
262K
输入 / 1M tokens
$0.115
输出 / 1M tokens
$0.917
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningChain of ThoughtLocal DeploymentAgent

qwen3.5-27b

文本图像视频文本

qwen

Mid-tier Qwen model tuned for multi-step reasoning and agent workflows, compact enough for private and on-prem deployment.

上下文
262K
输入 / 1M tokens
$0.086
输出 / 1M tokens
$0.688
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningCodingChain of ThoughtLocal Deployment

qwen3.5-flash

文本图像视频文本

qwen

Lightweight mid-tier variant tuned for low-latency agent workflows with coding and tool-calling support.

上下文
1M
输入 / 1M tokens
$0.029
输出 / 1M tokens
$0.287
缓存读 / 1M
$0.003
缓存写 / 1M
$0.036
CodingLong Text ProcessingReal-time ResponseLightweight

qwen3.5-plus

文本图像视频文本

qwen

Multimodal model from the Qwen3.5 line with toggleable thinking mode, tuned for image and video understanding, document parsing, and multimodal agents.

上下文
1M
输入 / 1M tokens
$0.115
输出 / 1M tokens
$0.688
缓存读 / 1M
$0.057
缓存写 / 1M
$0.717
FrontierDocument OCRMultimodal UnderstandingAgent

qwen3.6-27b

文本图像视频文本

qwen

Thinking-mode variant of the dense open-weight model that outputs step-by-step reasoning; at 27B it surpasses the prior open-weight 397B-A17B flagship across coding, math and multi-step reasoning benchmarks.

上下文
262K
输入 / 1M tokens
$0.412564
输出 / 1M tokens
$2.475384
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningMathChain of ThoughtScience

qwen3.6-flash

文本图像视频文本

qwen

Qwen 3.6 lightweight tier with text, image, and video input. Low latency and low cost, well-suited for high-volume classification, extraction, summarization, and simple agent workflows.

上下文
1M
输入 / 1M tokens
$0.165
输出 / 1M tokens
$0.99
缓存读 / 1M
$0
缓存写 / 1M
$0
Cost EffectiveReal-time ResponseLightweightMultimodal Understanding

qwen3.6-plus

文本图像视频文本

qwen

Enhanced Qwen model with strong bilingual (Chinese/English) comprehension, excelling at long-document analysis and structured output.

上下文
1M
输入 / 1M tokens
$0.276
输出 / 1M tokens
$1.651
缓存读 / 1M
$0
缓存写 / 1M
$0
CodingOfficeTranslation

qwen3.7-flash

文本图像视频文本

qwen

Qwen3.7 Flash is the lightweight tier in the series, a vision-language reasoning model with strengths in object recognition, spatial understanding, and real-world visual perception. Suited for multimodal agents, visual coding, search, and computer interaction.

上下文
1M
输入 / 1M tokens
$0.03
输出 / 1M tokens
$0.13
缓存读 / 1M
$0.006
缓存写 / 1M
$0.038
ReasoningVisionCost EffectiveLong Text Processing

qwen3.7-plus

文本图像视频文本

qwen

Qwen 3.7 Plus is a mid-tier multimodal model that reads screens, operates GUIs, and navigates mobile apps end-to-end. Suited for agent workflows and tool calling.

上下文
1M
输入 / 1M tokens
$0.276
输出 / 1M tokens
$1.101
缓存读 / 1M
$0.056
缓存写 / 1M
$0
VisionBalanced PerformanceTask AutomationUI Understanding

qwen3.8-flash

文本图像视频文本

qwen

Flash tier in Qwen3.8: multimodal reasoning for coding and agent workflows, with visual understanding across documents, codebases, and long video.

上下文
1M
输入 / 1M tokens
$0.113
输出 / 1M tokens
$0.382
缓存读 / 1M
$0.014
缓存写 / 1M
$0.177
ReasoningVisionCodingMultimodal Understanding

qwen3.8-max

文本图像视频文本

qwen

Qwen3.8 series flagship, the general-availability successor to Max Preview. A multimodal reasoning model built for complex reasoning, visual understanding, coding, and agentic workflows.

上下文
1M
输入 / 1M tokens
$2
输出 / 1M tokens
$6
缓存读 / 1M
$0.25
缓存写 / 1M
$2.5
CodingReasoningVisionFrontier

mimo-v2.5

文本图像音频视频文本

xiaomi

Mid-tier V2.5 model with native image, audio, and video understanding, built for agent workflows and coding at lower cost than Pro.

上下文
1M
输入 / 1M tokens
$0.14
输出 / 1M tokens
$0.28
缓存读 / 1M
$0.0428
缓存写 / 1M
$0
CodingMultimodal UnderstandingAgent CodingAgent

glm-4.6v

文本图像视频文本

Z.ai

Multimodal model with strong image + text understanding.

上下文
128K
输入 / 1M tokens
$0.3
输出 / 1M tokens
$0.9
缓存读 / 1M
$0.055
缓存写 / 1M
$0
VisionVisual QAMultimodal Understanding

glm-5.3-flash

文本图像视频文本

Z.ai

Flash tier of GLM-5.3: native multimodal inputs, suited for efficient coding and long-horizon agents with stable long-context behavior in production.

上下文
1M
输入 / 1M tokens
$0.15
输出 / 1M tokens
$0.5
缓存读 / 1M
$0.03
缓存写 / 1M
$0
VisionCost EffectiveCodingMultimodal Understanding
需要完整模型清单或企业接入方案?进入控制台