ernie-5.0
baidu
A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.
- 上下文
- 128K
- 输入 / 1M tokens
- $0.84
- 输出 / 1M tokens
- $3.28
- 缓存读 / 1M
- $0
- 缓存写 / 1M
- $0
选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。
baidu
A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.
bytedance
Turbo tier in Seed 2.1: multimodal model for coding and long-horizon agents, suited to multi-step task execution and visual content understanding.
bytedance
Agent-oriented foundation model tuned for tool calling and complex instruction following, with use cases spanning GUI agents, search agents, and multi-step task orchestration.
bytedance
Code-focused preview model for agent workflows, with competitive SWE-Bench and LiveCodeBench scores, aimed at autonomous coding agents and multi-language software tasks.
bytedance
Lightweight tier of the Seed 2.0 family tuned for low-latency agent, coding, and GUI workloads.
bytedance
Lightweight tier in the Seed 2.0 family, tuned for high-throughput, low-latency tasks with adjustable reasoning levels. Fits batch processing, moderation, and classification.
bytedance
Flagship of the Seed 2.0 series for complex reasoning and long-horizon agent workflows with reliable tool use.
A high-performance general-purpose model from Google, designed for advanced reasoning, coding, mathematics, and scientific tasks. Its built-in thinking capabilities improve response accuracy and enable deeper contextual understanding.
A lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency, with faster token generation and higher throughput. Thinking is disabled by default but can be enabled via API.
Massive context window. Great for video understanding.
Speed-tier variant in the Gemini 3 line, pairing Pro-grade reasoning with Flash-level latency and cost, built for agent workflows and high-throughput interactive apps.
A high-efficiency multimodal lite model for low-latency, high-volume workloads like translation, classification, and data extraction, priced at about half of Gemini 3 Flash.
Google's latest. Strong reasoning with native image and video.
Designed for efficient multimodal AI tasks, offering strong coding, reasoning, real-time chat, and agent execution at Flash-tier cost and speed.
Fast and cost-efficient.
minimax
MiniMax\\'s first multimodal release, accepting image and video input. Long-context inference runs cheaper and faster, and it handles multi-step agent workflows and computer-use scenarios.
Moonshot
Top Chinese-English bilingual. Strong long context.
Moonshot
Latest multimodal model in the K2 series, built for long-horizon coding, code-driven UI/UX generation, and multi-agent orchestration.
Moonshot
Coding-focused variant in the Kimi K2 family. Accepts text and image inputs, with reasoning on by default. Suited for long-horizon coding and agentic workflows.
qwen
Lightweight multimodal model with balanced performance.
qwen
Plus-tier vision-language model in the Qwen3-VL line, handling text, image and video inputs with strengths in document parsing, video understanding, spatial grounding and agent tool use.
qwen
Open-weight mid-tier MoE text model with long chain-of-thought reasoning, strong on function-calling and multi-step planning benchmarks for self-hosted agentic workflows.
qwen
Mid-tier Qwen model tuned for multi-step reasoning and agent workflows, compact enough for private and on-prem deployment.
qwen
Lightweight mid-tier variant tuned for low-latency agent workflows with coding and tool-calling support.
qwen
Multimodal model from the Qwen3.5 line with toggleable thinking mode, tuned for image and video understanding, document parsing, and multimodal agents.
qwen
Thinking-mode variant of the dense open-weight model that outputs step-by-step reasoning; at 27B it surpasses the prior open-weight 397B-A17B flagship across coding, math and multi-step reasoning benchmarks.
qwen
Qwen 3.6 lightweight tier with text, image, and video input. Low latency and low cost, well-suited for high-volume classification, extraction, summarization, and simple agent workflows.
qwen
Enhanced Qwen model with strong bilingual (Chinese/English) comprehension, excelling at long-document analysis and structured output.
qwen
Qwen3.7 Flash is the lightweight tier in the series, a vision-language reasoning model with strengths in object recognition, spatial understanding, and real-world visual perception. Suited for multimodal agents, visual coding, search, and computer interaction.
qwen
Qwen 3.7 Plus is a mid-tier multimodal model that reads screens, operates GUIs, and navigates mobile apps end-to-end. Suited for agent workflows and tool calling.
qwen
Flash tier in Qwen3.8: multimodal reasoning for coding and agent workflows, with visual understanding across documents, codebases, and long video.
qwen
Qwen3.8 series flagship, the general-availability successor to Max Preview. A multimodal reasoning model built for complex reasoning, visual understanding, coding, and agentic workflows.
xiaomi
Mid-tier V2.5 model with native image, audio, and video understanding, built for agent workflows and coding at lower cost than Pro.
Z.ai
Multimodal model with strong image + text understanding.
Z.ai
Flash tier of GLM-5.3: native multimodal inputs, suited for efficient coding and long-horizon agents with stable long-context behavior in production.