ernie-5.0
baidu
A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.
- 上下文
- 128K
- 输入 / 1M tokens
- $0.84
- 输出 / 1M tokens
- $3.28
- 缓存读 / 1M
- $0
- 缓存写 / 1M
- $0
选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。
baidu
A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.
bytedance
Character roleplay model for virtual companion scenarios, with stable persona adherence and plot progression in multi-turn dialogue, supporting multimodal inputs.
bytedance
Lightweight tier of the Seed 2.0 family tuned for low-latency agent, coding, and GUI workloads.
bytedance
Lightweight tier in the Seed 2.0 family, tuned for high-throughput, low-latency tasks with adjustable reasoning levels. Fits batch processing, moderation, and classification.
A high-performance general-purpose model from Google, designed for advanced reasoning, coding, mathematics, and scientific tasks. Its built-in thinking capabilities improve response accuracy and enable deeper contextual understanding.
A lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency, with faster token generation and higher throughput. Thinking is disabled by default but can be enabled via API.
Massive context window. Great for video understanding.
Speed-tier variant in the Gemini 3 line, pairing Pro-grade reasoning with Flash-level latency and cost, built for agent workflows and high-throughput interactive apps.
A high-efficiency multimodal lite model for low-latency, high-volume workloads like translation, classification, and data extraction, priced at about half of Gemini 3 Flash.
Google's latest. Strong reasoning with native image and video.
Designed for efficient multimodal AI tasks, offering strong coding, reasoning, real-time chat, and agent execution at Flash-tier cost and speed.
Fast and cost-efficient.
OpenAI
GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o Mini audio capabilities.
OpenAI
GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o audio capabilities.
OpenAI
GPT-4o Transcribe Diarize is an automatic speech recognition (ASR) model with built-in speaker diarization, meaning it associates audio segments with different speakers in a conversation.
OpenAI
GPT-5.5 natively supports video comprehension, speech emotion recognition, and autonomous tool calling, delivering omni-modal foundational capabilities for agent-based execution of complex tasks.
OpenAI
Whisper is OpenAI's open-source automatic speech recognition model. It supports transcription and translation across 50+ languages from audio files up to 25 MB. Accepts formats including mp3, mp4, wav, and webm.
xiaomi
Mid-tier V2.5 model with native image, audio, and video understanding, built for agent workflows and coding at lower cost than Pro.