模型广场

选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。

重置筛选
全部 12文本 12图像 0音频 0视频 0

deepseek-v4-flash

文本文本

DeepSeek

Faster, lower-cost variant of DeepSeek V4, optimized for coding assistants, high-throughput chat, and latency-sensitive Agent workflows.

上下文
1M
输入 / 1M tokens
$0.22
输出 / 1M tokens
$0.66
缓存读 / 1M
$0.007
缓存写 / 1M
$0
Coding

gpt-4o-mini

文本图像文本

OpenAI

A lightweight model tuned for low-cost, low-latency text tasks like chatbots, summarization, and bulk classification.

上下文
128K
输入 / 1M tokens
$0.15
输出 / 1M tokens
$0.6
缓存读 / 1M
$0.075
缓存写 / 1M
$0
Cost EffectiveLightweightClassificationBatch Generation

qwen3-vl-8b-instruct

文本图像视频文本

qwen

Lightweight multimodal model with balanced performance.

上下文
128K
输入 / 1M tokens
$0.25
输出 / 1M tokens
$0.75
缓存读 / 1M
$0.12
缓存写 / 1M
$0
Visual QAUI UnderstandingImage Parsing

qwen3.5-122b-a10b

文本图像视频文本

qwen

Open-weight mid-tier MoE text model with long chain-of-thought reasoning, strong on function-calling and multi-step planning benchmarks for self-hosted agentic workflows.

上下文
262K
输入 / 1M tokens
$0.115
输出 / 1M tokens
$0.917
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningChain of ThoughtLocal DeploymentAgent

qwen3.5-27b

文本图像视频文本

qwen

Mid-tier Qwen model tuned for multi-step reasoning and agent workflows, compact enough for private and on-prem deployment.

上下文
262K
输入 / 1M tokens
$0.086
输出 / 1M tokens
$0.688
缓存读 / 1M
$0
缓存写 / 1M
$0
ReasoningCodingChain of ThoughtLocal Deployment

qwen3.5-plus

文本图像视频文本

qwen

Multimodal model from the Qwen3.5 line with toggleable thinking mode, tuned for image and video understanding, document parsing, and multimodal agents.

上下文
1M
输入 / 1M tokens
$0.115
输出 / 1M tokens
$0.688
缓存读 / 1M
$0.057
缓存写 / 1M
$0.717
FrontierDocument OCRMultimodal UnderstandingAgent

qwen3.6-flash

文本图像视频文本

qwen

Qwen 3.6 lightweight tier with text, image, and video input. Low latency and low cost, well-suited for high-volume classification, extraction, summarization, and simple agent workflows.

上下文
1M
输入 / 1M tokens
$0.165
输出 / 1M tokens
$0.99
缓存读 / 1M
$0
缓存写 / 1M
$0
Cost EffectiveReal-time ResponseLightweightMultimodal Understanding

hy3

文本文本

tencent

A 295B MoE reasoning model from Tencent with 21B active parameters, configurable reasoning effort, and a 262K context window. It is designed for agentic workflows, long-horizon tasks, coding, document processing, and financial analysis, with an emphasis on reliable tool use and reduced hallucinations.

上下文
262K
输入 / 1M tokens
$0.132
输出 / 1M tokens
$0.528
缓存读 / 1M
$0.033
缓存写 / 1M
$0
ReasoningChain of ThoughtComplex PlanningAgent

hy3-preview

文本文本

tencent

Hy3 preview in high-reasoning mode, built for complex problem solving and multi-step agent workflows that need dependable execution.

上下文
262K
输入 / 1M tokens
$0.18
输出 / 1M tokens
$0.6
缓存读 / 1M
$0.06
缓存写 / 1M
$0
ReasoningChain of ThoughtComplex PlanningAgent

mimo-v2.5-pro

文本文本

xiaomi

Flagship coding and agent model that sustains thousand-tool-call autonomous workflows for complex software engineering.

上下文
1M
输入 / 1M tokens
$0.435
输出 / 1M tokens
$0.87
缓存读 / 1M
$0.0036
缓存写 / 1M
$0
CodingProduction CodeAgent CodingAgent

glm-4.6v

文本图像视频文本

Z.ai

Multimodal model with strong image + text understanding.

上下文
128K
输入 / 1M tokens
$0.3
输出 / 1M tokens
$0.9
缓存读 / 1M
$0.055
缓存写 / 1M
$0
VisionVisual QAMultimodal Understanding

glm-5.3-flash

文本图像视频文本

Z.ai

Flash tier of GLM-5.3: native multimodal inputs, suited for efficient coding and long-horizon agents with stable long-context behavior in production.

上下文
1M
输入 / 1M tokens
$0.15
输出 / 1M tokens
$0.5
缓存读 / 1M
$0.03
缓存写 / 1M
$0
VisionCost EffectiveCodingMultimodal Understanding
需要完整模型清单或企业接入方案?进入控制台