模型广场

选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。

重置筛选
全部 50文本 50图像 0音频 0视频 0

ernie-5.0

文本图像视频音频文本

baidu

A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.

上下文
128K
输入 / 1M tokens
$0.84
输出 / 1M tokens
$3.28
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningVisionFrontierMultimodal Understanding

ernie-5.1

文本图像文本

baidu

An efficient iteration of the flagship line, delivering major gains in reasoning, agentic handling, and search with far fewer parameters, balancing high performance and low-cost deployment.

上下文
128K
输入 / 1M tokens
$0.56
输出 / 1M tokens
$2.53
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningCost EffectiveChain of ThoughtAgent

seed-2-1-turbo

文本图像视频文本

bytedance

Turbo tier in Seed 2.1: multimodal model for coding and long-horizon agents, suited to multi-step task execution and visual content understanding.

上下文
262K
输入 / 1M tokens
$0.5
输出 / 1M tokens
$2.5
缓存读 / 1M
$0.1
缓存写 / 1M
$0.008
VisionCost EffectiveCodingMultimodal Understanding

seed-1.8

文本图像视频文本

bytedance

Agent-oriented foundation model tuned for tool calling and complex instruction following, with use cases spanning GUI agents, search agents, and multi-step task orchestration.

上下文
256K
输入 / 1M tokens
$0.5
输出 / 1M tokens
$4
缓存读 / 1M
$0.05
缓存写 / 1M
$0.008333
Instruction FollowingTask AutomationComplex PlanningAgent

doubao-seed-2.0-code-preview-260215

文本图像视频文本

bytedance

Code-focused preview model for agent workflows, with competitive SWE-Bench and LiveCodeBench scores, aimed at autonomous coding agents and multi-language software tasks.

上下文
256K
输入 / 1M tokens
$0.5
输出 / 1M tokens
$3
缓存读 / 1M
$0.1
缓存写 / 1M
$0.008333
CodingProduction CodeRefactoringAgent Coding

doubao-seed-2.0-lite-260428

文本图像视频音频文本

bytedance

Lightweight tier of the Seed 2.0 family tuned for low-latency agent, coding, and GUI workloads.

上下文
256K
输入 / 1M tokens
$0.25
输出 / 1M tokens
$2
缓存读 / 1M
$0.05
缓存写 / 1M
$0.008333
Cost EffectiveCodingReal-time ResponseAgent

doubao-seed-2.0-pro-260215

文本图像视频文本

bytedance

Flagship of the Seed 2.0 series for complex reasoning and long-horizon agent workflows with reliable tool use.

上下文
256K
输入 / 1M tokens
$0.5
输出 / 1M tokens
$3
缓存读 / 1M
$0.1
缓存写 / 1M
$0.008333
ReasoningInstruction FollowingComplex PlanningAgent

deepseek-r1

文本文本

DeepSeek

Open-source reasoning model with visible chain-of-thought, tuned for math, coding, and multi-step problem solving.

上下文
164K
输入 / 1M tokens
$0.7
输出 / 1M tokens
$2.5
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningMathChain of ThoughtDeep Analysis

deepseek-v3

文本文本

DeepSeek

Open-source general-purpose LLM with representative instruction-following and coding skills within the open-weights ecosystem, suited for chat, code assistance, and enterprise text workflows.

上下文
131K
输入 / 1M tokens
$0.287
输出 / 1M tokens
$1.147
缓存读 / 1M
$0.057
缓存写 / 1M
-
FrontierCodingInstruction FollowingHigh-Quality Generation

deepseek-v3.1

文本文本

DeepSeek

Hybrid model that switches between thinking and non-thinking modes on a single endpoint, post-trained for tool use and code agent workflows.

上下文
164K
输入 / 1M tokens
$0.574
输出 / 1M tokens
$1.721
缓存读 / 1M
$0.13
缓存写 / 1M
$0
ReasoningCodingAgent CodingAgent

deepseek-v4-pro

文本文本

DeepSeek

DeepSeek V4 flagship with adjustable reasoning depth, built for full-codebase analysis, multi-step automation, and complex reasoning tasks.

上下文
1M
输入 / 1M tokens
$0.66
输出 / 1M tokens
$1.98
缓存读 / 1M
$0.022
缓存写 / 1M
$0
Reasoning

gemini-2.5-flash

文本图像视频音频文本

Google

A high-performance general-purpose model from Google, designed for advanced reasoning, coding, mathematics, and scientific tasks. Its built-in thinking capabilities improve response accuracy and enable deeper contextual understanding.

上下文
1.1M
输入 / 1M tokens
$0.3
输出 / 1M tokens
$2.5
缓存读 / 1M
$0.03
缓存写 / 1M
$0.08333
TechCodingTranslation

gemini-3-flash-preview

文本图像视频音频文本

Google

Speed-tier variant in the Gemini 3 line, pairing Pro-grade reasoning with Flash-level latency and cost, built for agent workflows and high-throughput interactive apps.

上下文
1M
输入 / 1M tokens
$0.5
输出 / 1M tokens
$3
缓存读 / 1M
$0.05
缓存写 / 1M
$0
CodingTranslationFinance

gemini-3.1-flash-lite

文本图像视频音频文本

Google

A high-efficiency multimodal lite model for low-latency, high-volume workloads like translation, classification, and data extraction, priced at about half of Gemini 3 Flash.

上下文
1M
输入 / 1M tokens
$0.25
输出 / 1M tokens
$1.5
缓存读 / 1M
$0.025
缓存写 / 1M
$0.08333
Instruction FollowingTask AutomationStructured OutputClassification

minimax-m2.1

文本文本

minimax

Cost-effective general model for balanced speed and quality.

上下文
128K
输入 / 1M tokens
$0.31
输出 / 1M tokens
$1.23
缓存读 / 1M
$0.031
缓存写 / 1M
$0.39
ChatMarketing CopyBatch Generation

minimax-m2.5

文本文本

minimax

Built for conversation and bilingual chat.

上下文
197K
输入 / 1M tokens
$0.304
输出 / 1M tokens
$1.213
缓存读 / 1M
$0.061
缓存写 / 1M
$0
ChatBilingual

minimax-m2.5-highspeed

文本文本

minimax

A high-throughput variant of M2.5 with matching quality at roughly triple the speed, tuned for coding and agentic tool use under latency-sensitive workloads.

上下文
200K
输入 / 1M tokens
$0.6
输出 / 1M tokens
$2.4
缓存读 / 1M
$0.03
缓存写 / 1M
$0.375
CodingReal-time ResponseBatch GenerationAgent

minimax-m2.7

文本文本

minimax

A next-generation autonomous language model that uses multi-agent collaboration to plan, execute, and continuously refine complex tasks. It supports production-grade workflows including live debugging, root cause analysis, financial modeling, and document generation across Word, Excel, and PowerPoint.

上下文
205K
输入 / 1M tokens
$0.3
输出 / 1M tokens
$1.2
缓存读 / 1M
$0.06
缓存写 / 1M
$0
FinanceCodingTranslation

minimax-m2.7-highspeed

文本文本

minimax

Speed-optimized variant of M2.7 with identical outputs, tuned for low-latency coding and agent tool-calling workloads.

上下文
200K
输入 / 1M tokens
$0.6
输出 / 1M tokens
$2.4
缓存读 / 1M
$0.06
缓存写 / 1M
$0.375
CodingReal-time ResponseOfficeAgent

minimax-m3

文本图像视频文本

minimax

MiniMax\\'s first multimodal release, accepting image and video input. Long-context inference runs cheaper and faster, and it handles multi-step agent workflows and computer-use scenarios.

上下文
1M
输入 / 1M tokens
$1.2
输出 / 1M tokens
$4.8
缓存读 / 1M
$0.12
缓存写 / 1M
$0
VisionLong Text ProcessingMultimodal UnderstandingAgent Coding

kimi-k2-instruct

文本文本

Moonshot

Instruction-tuned for agent workflows with reliable tool calling and multilingual coding.

上下文
131K
输入 / 1M tokens
$0.574
输出 / 1M tokens
$2.3
缓存读 / 1M
$0.115
缓存写 / 1M
$0
ReasoningCodingTask AutomationAgent

kimi-k2-thinking

文本文本

Moonshot

Open-source thinking model built for long-horizon reasoning and hundreds of sequential tool calls across autonomous research, coding, and agent workflows.

上下文
262K
输入 / 1M tokens
$0.6
输出 / 1M tokens
$2.5
缓存读 / 1M
$0.15
缓存写 / 1M
$0
ReasoningChain of ThoughtComplex PlanningAgent

kimi-k2.5

文本图像视频文本

Moonshot

Top Chinese-English bilingual. Strong long context.

上下文
262K
输入 / 1M tokens
$0.6
输出 / 1M tokens
$3.011
缓存读 / 1M
$0.15
缓存写 / 1M
$0.718
ChineseBilingual

kimi-k2.6

文本图像视频文本

Moonshot

Latest multimodal model in the K2 series, built for long-horizon coding, code-driven UI/UX generation, and multi-agent orchestration.

上下文
256K
输入 / 1M tokens
$0.8939
输出 / 1M tokens
$3.7131
缓存读 / 1M
$0.34
缓存写 / 1M
$1.1174
Vision

kimi-k2.7-code

文本视频图像文本

Moonshot

Coding-focused variant in the Kimi K2 family. Accepts text and image inputs, with reasoning on by default. Suited for long-horizon coding and agentic workflows.

上下文
262K
输入 / 1M tokens
$0.95
输出 / 1M tokens
$4
缓存读 / 1M
$0.19
缓存写 / 1M
$1.1174
CodingAgent CodingChain of Thought

gpt-4.1-mini

文本图像文本

OpenAI

Near GPT-4o performance at lower latency and cost, suited for high-frequency interactions, coding, and vision tasks.

上下文
1M
输入 / 1M tokens
$0.4
输出 / 1M tokens
$1.6
缓存读 / 1M
$0.1
缓存写 / 1M
-
ChatReasoningTranslation

gpt-5-mini

文本图像文本

OpenAI

Fast and low-cost; for high-concurrency text and reasoning tasks.

上下文
400K
输入 / 1M tokens
$0.25
输出 / 1M tokens
$2
缓存读 / 1M
$0.025
缓存写 / 1M
$0
Cost EffectiveUniversal

gpt-5.4-mini

文本图像文本

OpenAI

Lightweight GPT-5.4 sibling tuned for coding, tool use, and agent workflows at roughly twice the speed of the 5.4 Pro tier.

上下文
400K
输入 / 1M tokens
$0.75
输出 / 1M tokens
$4.5
缓存读 / 1M
$0.075
缓存写 / 1M
$0
CodingBalanced PerformanceAgent CodingAgent

gpt-5.4-nano

文本图像文本

OpenAI

Lightweight variant in the gpt-5.4 family, tuned for low-latency, high-volume tasks like classification, extraction, and sub-agent execution.

上下文
400K
输入 / 1M tokens
$0.2
输出 / 1M tokens
$1.25
缓存读 / 1M
$0.02
缓存写 / 1M
$0
Cost EffectiveLightweightClassificationBatch Generation

gpt-5.6-luna

文本图像文本

OpenAI

Entry point of the GPT-5.6 lineup. A small, low-latency workhorse for chat, tagging, and lightweight agent loops, keeping reasoning solid enough for routine work.

上下文
1.1M
输入 / 1M tokens
$0.2
输出 / 1M tokens
$1.2
缓存读 / 1M
$0.02
缓存写 / 1M
$0.25
ChatCost EffectiveReal-time ResponseLightweight

o3-mini

文本文本

OpenAI

Cost-efficient reasoning model with adjustable thinking effort, tuned for STEM and coding tasks where deliberation matters.

上下文
200K
输入 / 1M tokens
$1.1
输出 / 1M tokens
$4.4
缓存读 / 1M
$0.55
缓存写 / 1M
-
ReasoningMathChain of ThoughtDeep Analysis

o4-mini

文本图像文本

OpenAI

Lightweight o-series reasoning model with built-in chain-of-thought, delivering steady math, coding, and tool-use quality at low latency and cost.

上下文
2M
输入 / 1M tokens
$1.1
输出 / 1M tokens
$4.4
缓存读 / 1M
$0.275
缓存写 / 1M
$0
ReasoningCost EffectiveChain of ThoughtAgent

qwen3-235b-a22b-2507

文本文本

qwen

Instruction-tuned Qwen3 variant without thinking mode, strong on multilingual reasoning, math, code, and tool use for agent workflows.

上下文
262K
输入 / 1M tokens
$0.287
输出 / 1M tokens
$2.868
缓存读 / 1M
-
缓存写 / 1M
-
ReasoningCodingMathLong Text Processing

qwen3-32b-s1-v2604

文本文本

qwen

Internal fine-tune of Qwen3-32B for text-only Chinese workloads, tuned to in-house instruction style for QA, summarization and rewriting.

上下文
131K
输入 / 1M tokens
$0.287
输出 / 1M tokens
$1.147
缓存读 / 1M
-
缓存写 / 1M
-
Instruction FollowingContent GenerationTranslation

qwen3-max

文本文本

qwen

Qwen3-generation thinking-mode reasoning model with native tool use, tuned for math, coding, and multi-step agentic workflows.

上下文
262K
输入 / 1M tokens
$0.359
输出 / 1M tokens
$1.434
缓存读 / 1M
$0.072
缓存写 / 1M
$2.438
ReasoningMathChain of ThoughtAgent

qwen3-max-preview

文本文本

qwen

Deep-thinking preview that reasons step-by-step before answering, holding up on multi-step math, code, and Chinese-language tasks.

上下文
256K
输入 / 1M tokens
$0.861
输出 / 1M tokens
$3.441
缓存读 / 1M
$0.173
缓存写 / 1M
-
ReasoningMathChain of ThoughtAgent

qwen3-next-80b-a3b-instruct

文本文本

qwen

Instruction-tuned non-thinking chat model that answers directly without exposing reasoning, suited to agent workflows needing deterministic output like coding assistance and tool calling.

上下文
262K
输入 / 1M tokens
$0.15
输出 / 1M tokens
$1.2
缓存读 / 1M
-
缓存写 / 1M
-
Instruction FollowingTask AutomationStructured OutputAgent

qwen3-vl-plus

文本图像视频文本

qwen

Plus-tier vision-language model in the Qwen3-VL line, handling text, image and video inputs with strengths in document parsing, video understanding, spatial grounding and agent tool use.

上下文
262K
输入 / 1M tokens
$0.144
输出 / 1M tokens
$1.434
缓存读 / 1M
$0.029
缓存写 / 1M
$0.539
VisionUI UnderstandingDocument OCRMultimodal Understanding

qwen3.6-27b

文本图像视频文本

qwen

Thinking-mode variant of the dense open-weight model that outputs step-by-step reasoning; at 27B it surpasses the prior open-weight 397B-A17B flagship across coding, math and multi-step reasoning benchmarks.

上下文
262K
输入 / 1M tokens
$0.412564
输出 / 1M tokens
$2.475384
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningMathChain of ThoughtScience

qwen3.6-35b-a3b

文本文本

qwen

Open-weight coding model tuned for agentic terminal tasks and repo-scale reasoning with low inference cost.

上下文
262K
输入 / 1M tokens
$0.248
输出 / 1M tokens
$1.485
缓存读 / 1M
-
缓存写 / 1M
-
CodingTask AutomationAgent CodingAgent

qwen3.6-plus

文本图像视频文本

qwen

Enhanced Qwen model with strong bilingual (Chinese/English) comprehension, excelling at long-document analysis and structured output.

上下文
1M
输入 / 1M tokens
$0.276
输出 / 1M tokens
$1.651
缓存读 / 1M
$0
缓存写 / 1M
$0
CodingOfficeTranslation

qwen3.7-max

文本文本

qwen

Flagship agent-centric reasoning model with a 1M context window, excelling at coding, productivity, and long-horizon autonomous tasks.

上下文
1M
输入 / 1M tokens
$1.65
输出 / 1M tokens
$4.951
缓存读 / 1M
$0.33
缓存写 / 1M
$2.063
CodingProduction CodeRefactoringAgent Coding

qwen3.7-plus

文本图像视频文本

qwen

Qwen 3.7 Plus is a mid-tier multimodal model that reads screens, operates GUIs, and navigates mobile apps end-to-end. Suited for agent workflows and tool calling.

上下文
1M
输入 / 1M tokens
$0.276
输出 / 1M tokens
$1.101
缓存读 / 1M
$0.056
缓存写 / 1M
$0
VisionBalanced PerformanceTask AutomationUI Understanding

glm-4.6

文本文本

Z.ai

Text-only model in the GLM-4 family with improved coding benchmarks and tool calls during reasoning, suited for software development and agent workflows.

上下文
203K
输入 / 1M tokens
$0.574
输出 / 1M tokens
$2.294
缓存读 / 1M
$0.115
缓存写 / 1M
$0
FrontierCodingAgent CodingAgent

glm-4.7

文本文本

Z.ai

Balanced model with strong Chinese understanding.

上下文
128K
输入 / 1M tokens
$0.59
输出 / 1M tokens
$2.35
缓存读 / 1M
$0.55
缓存写 / 1M
$0
ChineseContent GenerationReasoning

glm-5

文本文本

Z.ai

Competitive general-purpose Chinese model.

上下文
203K
输入 / 1M tokens
$0.88
输出 / 1M tokens
$3.23
缓存读 / 1M
$0.172
缓存写 / 1M
$0
ChineseUniversal

glm-5-turbo

文本文本

Z.ai

Designed for agent-based workflows such as OpenClaw.

上下文
128K
输入 / 1M tokens
$1.2
输出 / 1M tokens
$4
缓存读 / 1M
$0.24
缓存写 / 1M
$0
CodingReasoning

glm-5.1

文本文本

Z.ai

Offers significantly improved coding and long-horizon task capabilities, autonomously planning, executing, and refining a single task for over eight hours to deliver complete, engineering-grade results.

上下文
128K
输入 / 1M tokens
$0.826
输出 / 1M tokens
$3.303
缓存读 / 1M
$1.035
缓存写 / 1M
$1.376
CodingTask Execution

glm-5.2

文本文本

Z.ai

GLM 5.2 is a large-scale reasoning model with a 1M-token context window, suited for long-horizon agents, repo-level coding, and multi-step automation.

上下文
1M
输入 / 1M tokens
$1.4
输出 / 1M tokens
$4.4
缓存读 / 1M
$0.275
缓存写 / 1M
$0
CodingReasoningLong Text ProcessingChain of Thought

glm-5.3

文本文本

Z.ai

Flagship-tier model for complex software engineering and long-horizon agent tasks, with stable execution across multi-turn tool calls and large-scale codebase refactoring.

上下文
1M
输入 / 1M tokens
$1.4
输出 / 1M tokens
$4.4
缓存读 / 1M
$0.26
缓存写 / 1M
$0
ReasoningFrontierCodingLong Text Processing
需要完整模型清单或企业接入方案?进入控制台