模型广场

选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。

重置筛选
全部 18文本 18图像 0音频 0视频 0

seed-character

文本图像音频文本

bytedance

Character roleplay model for virtual companion scenarios, with stable persona adherence and plot progression in multi-turn dialogue, supporting multimodal inputs.

上下文
128K
输入 / 1M tokens
$0.12
输出 / 1M tokens
$0.29
缓存读 / 1M
$0.02
缓存写 / 1M
$0
ChatVisionChineseCreative Generation

doubao-seed-2.0-mini-260428

文本图像视频音频文本

bytedance

Lightweight tier in the Seed 2.0 family, tuned for high-throughput, low-latency tasks with adjustable reasoning levels. Fits batch processing, moderation, and classification.

上下文
256K
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.02
缓存写 / 1M
$0.008333
Real-time ResponseLightweightClassificationBatch Generation

gemini-2.5-flash-lite

文本图像视频音频文本

Google

A lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency, with faster token generation and higher throughput. Thinking is disabled by default but can be enabled via API.

上下文
1.1M
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.01
缓存写 / 1M
$0.08333
FinanceCodingTranslation

gemma-4-31b

文本图像文本

Google

Open-weights dense model with stronger math and coding, plus native function calling for self-hosted agent workflows.

上下文
262K
输入 / 1M tokens
$0.14
输出 / 1M tokens
$0.4
缓存读 / 1M
$0
缓存写 / 1M
$0
CodingMathLocal DeploymentAgent

gpt-4.1-nano

文本图像文本

OpenAI

Fastest GPT. Sub-second responses.

上下文
1M
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.4
缓存读 / 1M
$0.025
缓存写 / 1M
-
Cost Effective

gpt-4o-mini

文本图像文本

OpenAI

A lightweight model tuned for low-cost, low-latency text tasks like chatbots, summarization, and bulk classification.

上下文
128K
输入 / 1M tokens
$0.15
输出 / 1M tokens
$0.6
缓存读 / 1M
$0.075
缓存写 / 1M
$0
Cost EffectiveLightweightClassificationBatch Generation

qwen3-next-80b-a3b-instruct

文本文本

qwen

Instruction-tuned non-thinking chat model that answers directly without exposing reasoning, suited to agent workflows needing deterministic output like coding assistance and tool calling.

上下文
262K
输入 / 1M tokens
$0.15
输出 / 1M tokens
$1.2
缓存读 / 1M
-
缓存写 / 1M
-
Instruction FollowingTask AutomationStructured OutputAgent

qwen3-vl-plus

文本图像视频文本

qwen

Plus-tier vision-language model in the Qwen3-VL line, handling text, image and video inputs with strengths in document parsing, video understanding, spatial grounding and agent tool use.

上下文
262K
输入 / 1M tokens
$0.144
输出 / 1M tokens
$1.434
缓存读 / 1M
$0.029
缓存写 / 1M
$0.539
VisionUI UnderstandingDocument OCRMultimodal Understanding

qwen3.5-122b-a10b

文本图像视频文本

qwen

Open-weight mid-tier MoE text model with long chain-of-thought reasoning, strong on function-calling and multi-step planning benchmarks for self-hosted agentic workflows.

上下文
262K
输入 / 1M tokens
$0.115
输出 / 1M tokens
$0.917
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningChain of ThoughtLocal DeploymentAgent

qwen3.5-plus

文本图像视频文本

qwen

Multimodal model from the Qwen3.5 line with toggleable thinking mode, tuned for image and video understanding, document parsing, and multimodal agents.

上下文
1M
输入 / 1M tokens
$0.115
输出 / 1M tokens
$0.688
缓存读 / 1M
$0.057
缓存写 / 1M
$0.717
FrontierDocument OCRMultimodal UnderstandingAgent

qwen3.6-flash

文本图像视频文本

qwen

Qwen 3.6 lightweight tier with text, image, and video input. Low latency and low cost, well-suited for high-volume classification, extraction, summarization, and simple agent workflows.

上下文
1M
输入 / 1M tokens
$0.165
输出 / 1M tokens
$0.99
缓存读 / 1M
$0
缓存写 / 1M
$0
Cost EffectiveReal-time ResponseLightweightMultimodal Understanding

qwen3.8-flash

文本图像视频文本

qwen

Flash tier in Qwen3.8: multimodal reasoning for coding and agent workflows, with visual understanding across documents, codebases, and long video.

上下文
1M
输入 / 1M tokens
$0.113
输出 / 1M tokens
$0.382
缓存读 / 1M
$0.014
缓存写 / 1M
$0.177
ReasoningVisionCodingMultimodal Understanding

step-3.5-flash

文本文本

stepfun

StepFun’s most capable open-source foundation model, built on a sparse MoE architecture that activates only 11B of its 196B parameters per token. It delivers efficient reasoning and fast performance, even with long contexts.

上下文
256K
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.3
缓存读 / 1M
$0
缓存写 / 1M
$0
CodingTranslationFinance

hy3

文本文本

tencent

A 295B MoE reasoning model from Tencent with 21B active parameters, configurable reasoning effort, and a 262K context window. It is designed for agentic workflows, long-horizon tasks, coding, document processing, and financial analysis, with an emphasis on reliable tool use and reduced hallucinations.

上下文
262K
输入 / 1M tokens
$0.132
输出 / 1M tokens
$0.528
缓存读 / 1M
$0.033
缓存写 / 1M
$0
ReasoningChain of ThoughtComplex PlanningAgent

hy3-preview

文本文本

tencent

Hy3 preview in high-reasoning mode, built for complex problem solving and multi-step agent workflows that need dependable execution.

上下文
262K
输入 / 1M tokens
$0.18
输出 / 1M tokens
$0.6
缓存读 / 1M
$0.06
缓存写 / 1M
$0
ReasoningChain of ThoughtComplex PlanningAgent

mimo-v2.5

文本图像音频视频文本

xiaomi

Mid-tier V2.5 model with native image, audio, and video understanding, built for agent workflows and coding at lower cost than Pro.

上下文
1M
输入 / 1M tokens
$0.14
输出 / 1M tokens
$0.28
缓存读 / 1M
$0.0428
缓存写 / 1M
$0
CodingMultimodal UnderstandingAgent CodingAgent

glm-4.7-flash

文本文本

Z.ai

Ultra-fast for Chinese. Very cheap.

上下文
203K
输入 / 1M tokens
$0.1
输出 / 1M tokens
$0.43
缓存读 / 1M
$0.01
缓存写 / 1M
$0
ChineseCost Effective

glm-5.3-flash

文本图像视频文本

Z.ai

Flash tier of GLM-5.3: native multimodal inputs, suited for efficient coding and long-horizon agents with stable long-context behavior in production.

上下文
1M
输入 / 1M tokens
$0.15
输出 / 1M tokens
$0.5
缓存读 / 1M
$0.03
缓存写 / 1M
$0
VisionCost EffectiveCodingMultimodal Understanding
需要完整模型清单或企业接入方案?进入控制台