View all available AI models. High-speed global routing, unified billing.
Pay-as-you-go: Top up your account in USD, and credits are consumed per API call at the rates shown below. No subscriptions, no hidden fees.
ApiHub model
DeepSeek-V4-Flash is a preview Mixture-of-Experts (MoE) language model in the DeepSeek-V4 series, with 284 billion total parameters, 13 billion active parameters, and support for an ultra-long context window of up to 1 million tokens.
DeepSeek-V4-Pro is the flagship Mixture-of-Experts (MoE) language model in the DeepSeek-V4 series, featuring 1.6 trillion total parameters, 49 billion active parameters, and native support for an ultra-long context window of up to 1 million tokens.
Doubao-Seed-2.0-Code is optimized for enterprise-grade programming needs. Building on the strong agent and vision-language model capabilities of Seed 2.0, it delivers enhanced coding performance, with particular strengths in front-end development and additional optimization for multilingual programming scenarios commonly encountered by enterprises. It is well suited for integration with a wide range of AI coding tools.
Doubao-Seed-2.0-Lite natively supports unified understanding across video, images, audio, and text. It is the first omni-modal understanding model in the Doubao model family, with enhanced capabilities in agents, coding, and GUI interaction. At the same compute cost, it offers a more cost-effective solution for enterprises deploying omni-modal reasoning tasks at scale and in batch workloads.
Doubao-Seed-2.0-Pro is a flagship, all-purpose general model built for complex reasoning and long-horizon task execution in the agentic era, with enhanced capabilities in multimodal understanding, long-context reasoning, structured generation, and tool-augmented execution.
Doubao-Seed-2.1 is a next-generation foundation model built for the era of coding and AI agents. It is available in two versions—Pro and Turbo—designed respectively for exploring highly complex tasks and supporting large-scale production workloads. Compared with the previous generation, Seed 2.1 delivers comprehensive upgrades in three key areas: production-grade coding and engineering delivery, long-horizon agent task execution, and multimodal understanding. With stronger autonomous planning and dynamic error-recovery capabilities, it is well equipped to handle real-world software development and high-value production tasks.
GLM-5.1 is a next-generation flagship model designed for agentic engineering. It adopts a Mixture-of-Experts (MoE) architecture with 754 billion parameters. The model delivers significantly enhanced coding capabilities and achieves leading performance on SWE-Bench Pro.
GLM-5.2 is Z.ai’s latest flagship model, designed for long-horizon task scenarios. Compared with GLM-5.1, it delivers significant improvements in long-horizon task performance. The 753B-parameter model supports a stable 1-million-token context window, offers stronger coding capabilities, and provides multiple thinking-effort levels, enabling a flexible balance between performance and latency. GLM-5.2 introduces the IndexShare architecture optimization, reducing per-token FLOPs by 2.9× at a 1-million-token context length. It also improves the MTP layer to support speculative decoding, increasing the maximum acceptance length by up to 20%.
GLM-5V-Turbo is Zhipu AI’s first multimodal coding foundation model, purpose-built for visual programming tasks. It can natively process multimodal inputs such as images, video, and text, while excelling at long-horizon planning, complex coding, and action execution. Deeply optimized for agent workflows, it can work closely with agents such as Claude Code and OpenClaw to complete the full loop of understanding the environment, planning actions, and executing tasks.
The production release of Hy3 has been optimized for real-world business scenarios. It features a Mixture-of-Experts architecture with 295 billion total parameters and 21 billion active parameters, natively supports a 256K context window, and offers multiple reasoning modes: no_think for ultra-fast responses, think_low for fast reasoning, and think_high for deep reasoning. This enables Hy3 to strike a balance between response speed, complex reasoning capabilities, and inference costs.
MiniMax-M3 is the latest language model in the M series, designed for agentic reasoning, tool use, coding, and long-context tasks. It is a frontier-level coding model with native multimodal capabilities and support for a 1-million-token context window.
Qwen3.7-Max is the largest and most capable model in the Qwen3.7 series. Currently, its text-only capabilities are available for preview. Qwen3.7 is a next-generation flagship model built for the agentic era, with core strengths in both the breadth and depth of its agentic capabilities. It excels across coding, office and productivity workflows, and long-horizon autonomous task execution.
Qwen3.7-Plus is a cost-effective model in the Qwen3.7 series. Building on its strong text capabilities, it delivers comprehensive upgrades in vision-language understanding while retaining full agentic capabilities across coding, tool use, and productivity workflows.