wan2.5-t2v-preview
Alibaba
Generate 5-10 second, 480p/720p/1080p video from text prompts with native synced audio track.
- 上下文
- 待核
- 输入 / 1M tokens
- -
- 输出 / 1M tokens
- -
- 缓存读 / 1M
- -
- 缓存写 / 1M
- -
选择模型后可复制配置给 AI,也可以快速测试当前连接延迟。
Alibaba
Generate 5-10 second, 480p/720p/1080p video from text prompts with native synced audio track.
Alibaba
Alibaba Tongyi Wanxiang Wan 2.6 image-to-video model animates a reference image into 720p/1080p clips with synchronized audio, preserving subject consistency with cinematic motion.
Alibaba
Speed-optimized Wan 2.6 image-to-video variant with optional audio output, generating 720p/1080p clips quickly and cost-effectively for high-throughput image animation.
Alibaba
Generate 720p/1080p video from text prompts and come with synchronized audio tracks. The picture has movie-level aesthetics and complex motion performance, suitable for creative video creation.
baidu
A natively omni-modal model with unified understanding of text, image, audio, and video, providing a flagship foundation for full-modality capabilities.
bytedance
Character roleplay model for virtual companion scenarios, with stable persona adherence and plot progression in multi-turn dialogue, supporting multimodal inputs.
bytedance
Lightweight tier of the Seed 2.0 family tuned for low-latency agent, coding, and GUI workloads.
bytedance
Lightweight tier in the Seed 2.0 family, tuned for high-throughput, low-latency tasks with adjustable reasoning levels. Fits batch processing, moderation, and classification.
bytedance
Standard tier of the Seedance family. Supports text-to-video, image-to-video, and reference-to-video, with strong character, style, and camera consistency. Built for ads and short-form creative.
bytedance
Lightweight Seedance 2.0 variant tuned for speed and cost rather than peak fidelity, with text-to-video, first/last frame control, and multi-reference inputs.
bytedance
A video generation model built for long-form storytelling, with first-frame and first-and-last-frame control, up to 50 multimodal reference assets, and optional audio generation. It also supports video editing and extension.
A high-performance general-purpose model from Google, designed for advanced reasoning, coding, mathematics, and scientific tasks. Its built-in thinking capabilities improve response accuracy and enable deeper contextual understanding.
A lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency, with faster token generation and higher throughput. Thinking is disabled by default but can be enabled via API.
Massive context window. Great for video understanding.
Speed-tier variant in the Gemini 3 line, pairing Pro-grade reasoning with Flash-level latency and cost, built for agent workflows and high-throughput interactive apps.
A high-efficiency multimodal lite model for low-latency, high-volume workloads like translation, classification, and data extraction, priced at about half of Gemini 3 Flash.
Google's latest. Strong reasoning with native image and video.
Designed for efficient multimodal AI tasks, offering strong coding, reasoning, real-time chat, and agent execution at Flash-tier cost and speed.
Fast and cost-efficient.
minimax
A lightweight open-weight video generation model focused on instruction-guided editing, text and brand rendering, and video-to-video motion transfer, with native audiovisual output. Built for advertising, e-commerce, and interface design workflows.
OpenAI
GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o Mini audio capabilities.
OpenAI
GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o audio capabilities.
OpenAI
GPT-4o Transcribe Diarize is an automatic speech recognition (ASR) model with built-in speaker diarization, meaning it associates audio segments with different speakers in a conversation.
OpenAI
GPT-5.5 natively supports video comprehension, speech emotion recognition, and autonomous tool calling, delivering omni-modal foundational capabilities for agent-based execution of complex tasks.
OpenAI
Whisper is OpenAI's open-source automatic speech recognition model. It supports transcription and translation across 50+ languages from audio files up to 25 MB. Accepts formats including mp3, mp4, wav, and webm.
xiaomi
Mid-tier V2.5 model with native image, audio, and video understanding, built for agent workflows and coding at lower cost than Pro.