{"alibaba/qwen-3-14b":{"maxTokens":16384,"contextWindow":40960,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.12,"outputPrice":0.24,"description":"Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support"},"alibaba/qwen-3-235b":{"maxTokens":16384,"contextWindow":262144,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.22,"outputPrice":0.88,"description":""},"alibaba/qwen-3-30b":{"maxTokens":16384,"contextWindow":40960,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.08,"outputPrice":0.29,"description":"Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support"},"alibaba/qwen-3-32b":{"maxTokens":8192,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.16,"outputPrice":0.64,"description":"Qwen3-32B is a world-class model with comparable quality to DeepSeek R1 while outperforming GPT-4.1 and Claude Sonnet 3.7. It excels in code-gen, tool-calling, and advanced reasoning, making it an exceptional model for a wide range of production use cases."},"alibaba/qwen-3.6-max-preview":{"maxTokens":64000,"contextWindow":240000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":1.3,"outputPrice":7.8,"cacheWritesPrice":1.625,"cacheReadsPrice":0.26,"description":"Compared with the previously released Qwen3-Max and Qwen3.6-Plus, this model features enhanced vibe coding abilities, more efficient coding agent execution, and significantly improved front-end development skills. Additionally, its long-tail knowledge retention has been further upgraded."},"alibaba/qwen3-235b-a22b-thinking":{"maxTokens":32768,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.39999999999999997,"outputPrice":4,"description":"Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade."},"alibaba/qwen3-coder":{"maxTokens":65536,"contextWindow":262144,"supportsImages":true,"supportsPromptCache":false,"inputPrice":1.5,"outputPrice":7.5,"cacheReadsPrice":0.3,"description":"Qwen3-Coder-480B-A35B-Instruct is a cutting-edge open coding model from Qwen, matching Claude Sonnet’s performance in agentic programming, browser automation, and core development tasks."},"alibaba/qwen3-coder-30b-a3b":{"maxTokens":8192,"contextWindow":262144,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":0.6,"description":"Efficient coding specialist balancing performance with cost-effectiveness for daily development tasks while maintaining strong tool integration capabilities."},"alibaba/qwen3-coder-next":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.5,"outputPrice":1.2,"description":"Qwen3-Coder-Next is an open-weight language model built specifically for coding, with strong performance on large-scale software engineering and agentic coding benchmarks. It uses a hybrid Mixture-of-Experts architecture to offer high capability at relatively modest active parameter counts, improving efficiency for real-world deployments. The model is trained on diverse code and natural language data so it can handle tasks like code generation, refactoring, debugging, repository-level reasoning, and technical explanation across multiple programming languages. It is also optimized for tool use and function calling, making it suitable as the core of coding agents that interact with shells, editors, issue trackers, and other developer tools."},"alibaba/qwen3-coder-plus":{"maxTokens":65536,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1,"outputPrice":5,"cacheReadsPrice":0.19999999999999998,"description":"Powered by Qwen3 this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities."},"alibaba/qwen3-max":{"maxTokens":32768,"contextWindow":262144,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.2,"outputPrice":6,"cacheReadsPrice":0.24,"description":"The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios."},"alibaba/qwen3-max-preview":{"maxTokens":32768,"contextWindow":262144,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.2,"outputPrice":6,"cacheReadsPrice":0.24,"description":"Qwen3-Max-Preview shows substantial gains over the 2.5 series in overall capability, with significant enhancements in Chinese-English text understanding, complex instruction following, handling of subjective open-ended tasks, multilingual ability, and tool invocation; model knowledge hallucinations are reduced."},"alibaba/qwen3-max-thinking":{"maxTokens":65536,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.2,"outputPrice":6,"cacheReadsPrice":0.24,"description":"Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model’s overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026."},"alibaba/qwen3-next-80b-a3b-instruct":{"maxTokens":32768,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":1.2,"description":"A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507)."},"alibaba/qwen3-next-80b-a3b-thinking":{"maxTokens":32768,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":1.2,"description":"A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507)."},"alibaba/qwen3-vl-235b-a22b-instruct":{"maxTokens":129024,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.39999999999999997,"outputPrice":1.5999999999999999,"description":"The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement."},"alibaba/qwen3-vl-instruct":{"maxTokens":129024,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.39999999999999997,"outputPrice":1.5999999999999999,"description":"The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement."},"alibaba/qwen3-vl-thinking":{"maxTokens":32768,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.39999999999999997,"outputPrice":4,"description":"Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade."},"alibaba/qwen3.5-flash":{"maxTokens":64000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.09999999999999999,"outputPrice":0.39999999999999997,"cacheWritesPrice":0.125,"cacheReadsPrice":0.001,"description":"The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance."},"alibaba/qwen3.5-plus":{"maxTokens":64000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.39999999999999997,"outputPrice":2.4,"cacheWritesPrice":0.5,"cacheReadsPrice":0.04,"description":"The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities."},"alibaba/qwen3.6-27b":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.6,"outputPrice":3.5999999999999996,"description":"The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance."},"alibaba/qwen3.6-plus":{"maxTokens":64000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.5,"outputPrice":3,"cacheWritesPrice":0.625,"cacheReadsPrice":0.09999999999999999,"description":"The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization."},"alibaba/qwen3.7-max":{"maxTokens":64000,"contextWindow":991000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":1.25,"outputPrice":3.75,"cacheWritesPrice":1.5625,"cacheReadsPrice":0.25,"description":"Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution."},"amazon/nova-2-lite":{"maxTokens":1000000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.3,"outputPrice":2.5,"cacheReadsPrice":0.075,"description":"Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text."},"amazon/nova-lite":{"maxTokens":8192,"contextWindow":300000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.06,"outputPrice":0.24,"description":"A very low cost multimodal model that is lightning fast for processing image, video, and text inputs."},"amazon/nova-micro":{"maxTokens":8192,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.035,"outputPrice":0.14,"description":"A text-only model that delivers the lowest latency responses at very low cost."},"amazon/nova-pro":{"maxTokens":8192,"contextWindow":300000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.7999999999999999,"outputPrice":3.1999999999999997,"description":"A highly capable multimodal model with the best combination of accuracy, speed, and cost for a wide range of tasks."},"anthropic/claude-3-haiku":{"maxTokens":4096,"contextWindow":200000,"supportsImages":true,"supportsPromptCache":true,"inputPrice":0.25,"outputPrice":1.25,"cacheWritesPrice":0.3,"cacheReadsPrice":0.03,"description":"Claude 3 Haiku is Anthropic's fastest, most compact model for near-instant responsiveness. It answers simple queries and requests with speed. Customers will be able to build seamless AI experiences that mimic human interactions. Claude 3 Haiku can process images and return text outputs, and features a 200K context window."},"anthropic/claude-3.5-haiku":{"maxTokens":8192,"contextWindow":200000,"supportsImages":true,"supportsPromptCache":true,"inputPrice":0.7999999999999999,"outputPrice":4,"cacheWritesPrice":1,"cacheReadsPrice":0.08,"description":"Claude 3 Haiku is Anthropic's fastest, most compact model for near-instant responsiveness. It answers simple queries and requests with speed. Customers will be able to build seamless AI experiences that mimic human interactions. Claude 3 Haiku can process images and return text outputs, and features a 200K context window."},"anthropic/claude-haiku-4.5":{"maxTokens":64000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":1,"outputPrice":5,"cacheWritesPrice":1.25,"cacheReadsPrice":0.09999999999999999,"description":"Claude Haiku 4.5 matches Sonnet 4's performance on coding, computer use, and agent tasks at substantially lower cost and faster speeds. It delivers near-frontier performance and Claude’s unique character at a price point that works for scaled sub-agent deployments, free tier products, and intelligence-sensitive applications with budget constraints."},"anthropic/claude-opus-4":{"maxTokens":32000,"contextWindow":200000,"supportsImages":true,"supportsPromptCache":true,"inputPrice":15,"outputPrice":75,"cacheWritesPrice":18.75,"cacheReadsPrice":1.5,"description":"Claude Opus 4 is Anthropic's most powerful model yet and the best coding model in the world, leading on SWE-bench (72.5%) and Terminal-bench (43.2%). It delivers sustained performance on long-running tasks that require focused effort and thousands of steps, with the ability to work continuously for several hours—dramatically outperforming all Sonnet models and significantly expanding what AI agents can accomplish."},"anthropic/claude-opus-4.1":{"maxTokens":32000,"contextWindow":200000,"supportsImages":true,"supportsPromptCache":true,"inputPrice":15,"outputPrice":75,"cacheWritesPrice":18.75,"cacheReadsPrice":1.5,"description":"Claude Opus 4.1 is a drop-in replacement for Opus 4 that delivers superior performance and precision for real-world coding and agentic tasks. Opus 4.1 advances state-of-the-art coding performance to 74.5% on SWE-bench Verified, and handles complex, multi-step problems with more rigor and attention to detail."},"anthropic/claude-opus-4.5":{"maxTokens":64000,"contextWindow":200000,"supportsImages":true,"supportsPromptCache":true,"inputPrice":5,"outputPrice":25,"cacheWritesPrice":6.25,"cacheReadsPrice":0.5,"description":"Claude Opus 4.5 is Anthropic’s latest model in the Opus series, meant for demanding reasoning tasks and complex problem solving. This model has improvements in general intelligence and vision compared to previous iterations. In addition, it is suited for difficult coding tasks and agentic workflows, especially those with computer use and tool use, and can effectively handle context usage and external memory files."},"anthropic/claude-opus-4.6":{"maxTokens":128000,"contextWindow":1000000,"supportsImages":true,"supportsPromptCache":true,"inputPrice":5,"outputPrice":25,"cacheWritesPrice":6.25,"cacheReadsPrice":0.5,"description":"Opus 4.6 is the world’s best model for coding and professional work, built to power agents that take on whole categories of real-world work. It excels across the entire SDLC, breaking through on hard problems, identifying complex bugs, and demonstrating deeper codebase understanding. It also delivers a step-change in knowledge work, with near-production-ready documents, presentations, and spreadsheets on the first pass."},"anthropic/claude-opus-4.7":{"maxTokens":128000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":5,"outputPrice":25,"cacheWritesPrice":6.25,"cacheReadsPrice":0.5,"description":"Opus 4.7 builds on the coding and agentic strengths of Opus 4.6 with stronger performance on complex, multi-step tasks and more reliable agentic execution. It also brings improved performance on knowledge work, from drafting documents to building presentations and analyzing data."},"anthropic/claude-opus-4.8":{"maxTokens":128000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":5,"outputPrice":25,"cacheWritesPrice":6.25,"cacheReadsPrice":0.5,"description":"Opus 4.8 is a focused upgrade to Opus 4.7 and is Anthropic's best generally available model for coding, agentic tasks, and enterprise workflows. It builds on the strengths of previous Opus models with stronger performance on complex, multi-step coding tasks. Anthropic recommends using it on long-horizon coding and agentic tasks. It is also stronger on professional work, including document drafting, data analysis, and presentations."},"anthropic/claude-sonnet-4":{"maxTokens":64000,"contextWindow":1000000,"supportsImages":true,"supportsPromptCache":true,"inputPrice":3,"outputPrice":15,"cacheWritesPrice":3.75,"cacheReadsPrice":0.3,"description":"Claude Sonnet 4 significantly improves on Sonnet 3.7's industry-leading capabilities, excelling in coding with a state-of-the-art 72.7% on SWE-bench. The model balances performance and efficiency for internal and external use cases, with enhanced steerability for greater control over implementations. While not matching Opus 4 in most domains, it delivers an optimal mix of capability and practicality."},"anthropic/claude-sonnet-4.5":{"maxTokens":64000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":3,"outputPrice":15,"cacheWritesPrice":3.75,"cacheReadsPrice":0.3,"description":"Claude Sonnet 4.5 is the newest model in the Sonnet series, offering improvements and updates over Sonnet 4."},"anthropic/claude-sonnet-4.6":{"maxTokens":128000,"contextWindow":1000000,"supportsImages":true,"supportsPromptCache":true,"inputPrice":3,"outputPrice":15,"cacheWritesPrice":3.75,"cacheReadsPrice":0.3,"description":"Claude Sonnet 4.6 is the most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation."},"arcee-ai/trinity-large-preview":{"maxTokens":131000,"contextWindow":131000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":1,"description":"Trinity Large (Preview) is a 400B-parameter (13B active) sparse mixture-of-experts language model, engineered to scale model capacity while maintaining inference efficiency over long contexts, with strong performance in reasoning-heavy workloads including math, coding-related tasks, and multi-step agent workflows."},"arcee-ai/trinity-large-thinking":{"maxTokens":80000,"contextWindow":262100,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":0.8999999999999999,"description":"Trinity-Large-Thinking is a reasoning-optimized variant of Arcee AI's Trinity-Large family — a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token. Built on Trinity-Large-Base and post-trained with extended chain-of-thought reasoning and agentic RL, Trinity-Large-Thinking delivers state-of-the-art performance on agentic benchmarks while maintaining strong general capabilities."},"arcee-ai/trinity-mini":{"maxTokens":131072,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.045,"outputPrice":0.15,"description":"Trinity Mini is a 26B-parameter (3B active) sparse mixture-of-experts language model, engineered for efficient inference over long contexts with robust function calling and multi-step agent workflows."},"bytedance/seed-1.6":{"maxTokens":32000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":2,"cacheReadsPrice":0.049999999999999996,"description":"ByteDance's new multimodal deep-thinking model, supporting both text and visual inputs with enhanced reasoning capabilities."},"bytedance/seed-1.8":{"maxTokens":64000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":2,"cacheReadsPrice":0.049999999999999996,"description":"Bytedance Seed 1.8 features stronger multimodal understanding and agent capabilities. The model delivers superior performance across a wide range of complex real-world tasks, helping enterprises create greater value."},"cohere/command-a":{"maxTokens":8000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":2.5,"outputPrice":10,"description":"Command A is Cohere's most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024."},"deepseek/deepseek-r1":{"maxTokens":8192,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.35,"outputPrice":5.4,"description":"DeepSeek-R1 provides customers a state-of-the-art reasoning model, optimized for general reasoning tasks, math, science, and code generation."},"deepseek/deepseek-v3":{"maxTokens":16384,"contextWindow":163840,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.77,"outputPrice":0.77,"description":"Fast general-purpose LLM with enhanced reasoning capabilities"},"deepseek/deepseek-v3.1":{"maxTokens":8192,"contextWindow":163840,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.56,"outputPrice":1.68,"cacheReadsPrice":0.28,"description":"DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. DeepSeek has expanded their dataset by collecting additional long documents and substantially extending both training phases."},"deepseek/deepseek-v3.1-terminus":{"maxTokens":65536,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.27,"outputPrice":1,"cacheReadsPrice":0.135,"description":"DeepSeek-V3.1-Terminus delivers more stable & reliable outputs across benchmarks compared to the previous version and addresses user feedback (i.e. language consistency and agent upgrades)."},"deepseek/deepseek-v3.2":{"maxTokens":8000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.28,"outputPrice":0.42,"cacheReadsPrice":0.028,"description":"DeepSeek-V3.2: Official successor to V3.2-Exp."},"deepseek/deepseek-v3.2-thinking":{"maxTokens":8000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.62,"outputPrice":1.85,"description":"DeepSeek‑V3.2 from DeepSeek harmonizes high computational efficiency with superior reasoning and agent performance. It builds on three main techniques: DeepSeek Sparse Attention for long‑context efficiency, a scalable reinforcement learning framework, and a large‑scale agentic task synthesis pipeline. This model excels at long-context reasoning and agentic tasks, efficiently handling extended inputs while maintaining strong accuracy. Its sparse attention design enables it to process complex, multi-step workflows without excessive compute costs. Overall, DeepSeek‑V3.2 targets long‑context reasoning, tool‑using agents, and efficient deployment in production environments."},"deepseek/deepseek-v4-flash":{"maxTokens":384000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.14,"outputPrice":0.28,"cacheReadsPrice":0.0028,"description":"DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA)\nand Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) ManifoldConstrained Hyper-Connections (mHC) that enhance conventional residual connections; (3)\nand the Muon optimizer for faster convergence and greater training stability"},"deepseek/deepseek-v4-pro":{"maxTokens":384000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.435,"outputPrice":0.87,"cacheReadsPrice":0.0036,"description":"DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA)\nand Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) ManifoldConstrained Hyper-Connections (mHC) that enhance conventional residual connections; (3)\nand the Muon optimizer for faster convergence and greater training stability"},"google/gemini-2.0-flash":{"maxTokens":8192,"contextWindow":1048576,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":0.6,"cacheReadsPrice":0.024999999999999998,"description":"Gemini 2.0 Flash delivers next-gen features and improved capabilities, including superior speed, built-in tool use, multimodal generation, and a 1M token context window."},"google/gemini-2.0-flash-lite":{"maxTokens":8192,"contextWindow":1048576,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.075,"outputPrice":0.3,"cacheReadsPrice":0.02,"description":"Gemini 2.0 Flash delivers next-gen features and improved capabilities, including superior speed, built-in tool use, multimodal generation, and a 1M token context window."},"google/gemini-2.5-flash":{"maxTokens":65536,"contextWindow":1000000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.3,"outputPrice":2.5,"cacheReadsPrice":0.03,"description":"Gemini 2.5 Flash is a thinking model that offers great, well-rounded capabilities. It is designed to offer a balance between price and performance with multimodal support and a 1M token context window."},"google/gemini-2.5-flash-image":{"maxTokens":65536,"contextWindow":32768,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.3,"outputPrice":2.5,"cacheReadsPrice":0.03,"description":"Nano Banana (Gemini 2.5 Flash Image) is Google's first fully hybrid reasoning model, letting developers turn thinking on or off and set thinking budgets to balance quality, cost, and latency. Upgraded for rapid creative workflows, it can generate interleaved text and images and supports conversational, multi‑turn image editing in natural language. It’s also locale‑aware, enabling culturally and linguistically appropriate image generation for audiences worldwide."},"google/gemini-2.5-flash-lite":{"maxTokens":65536,"contextWindow":1048576,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.09999999999999999,"outputPrice":0.39999999999999997,"cacheReadsPrice":0.01,"description":"Gemini 2.5 Flash-Lite is a balanced, low-latency model with configurable thinking budgets and tool connectivity (e.g., Google Search grounding and code execution). It supports multimodal input and offers a 1M-token context window."},"google/gemini-2.5-pro":{"maxTokens":65536,"contextWindow":1048576,"supportsImages":true,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":10,"cacheReadsPrice":0.125,"description":"Gemini 2.5 Pro is our most advanced reasoning Gemini model, capable of solving complex problems. Gemini 2.5 Pro can comprehend vast datasets and challenging problems from different information sources, including text, audio, images, video, and even entire code repositories."},"google/gemini-3-flash":{"maxTokens":65000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.5,"outputPrice":3,"cacheReadsPrice":0.049999999999999996,"description":"Google's most intelligent model built for speed, combining frontier intelligence with superior search and grounding."},"google/gemini-3-pro-image":{"maxTokens":32768,"contextWindow":65536,"supportsImages":false,"supportsPromptCache":false,"inputPrice":2,"outputPrice":12,"cacheReadsPrice":0.19999999999999998,"description":"Nano Banana Pro (Gemini 3 Pro Image) builds on Nano Banana's generation capabilities into a new era of studio-quality, functional design to help you create and edit high-fidelity, production-ready visuals with unparalleled precision and control. Improvements include enhanced world knowledge and reasoning, dynamic text and translation, and studio level controls."},"google/gemini-3-pro-preview":{"maxTokens":64000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":2,"outputPrice":12,"cacheReadsPrice":0.19999999999999998,"description":"This model improves upon Gemini 2.5 Pro and is catered towards challenging tasks, especially those involving complex reasoning or agentic workflows. Improvements highlighted include use cases for coding, multi-step function calling, planning, reasoning, deep knowledge tasks, and instruction following."},"google/gemini-3.1-flash-image-preview":{"maxTokens":32768,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.5,"outputPrice":3,"cacheReadsPrice":0.049999999999999996,"description":"Gemini 3.1 Flash Image is optimized for image understanding and generation and offers a balance of price and performance."},"google/gemini-3.1-flash-lite":{"maxTokens":65000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":1.5,"cacheReadsPrice":0.03,"description":"Gemini 3.1 Flash Lite outperforms 2.5 Flash Lite on overall quality and lands close to 2.5 Flash performance across key capability areas. It is a workhorse model for high-volume use cases, with improvements across audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion."},"google/gemini-3.1-flash-lite-preview":{"maxTokens":65000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":1.5,"cacheReadsPrice":0.03,"description":"Gemini 3.1 Flash Lite Preview outperforms 2.5 Flash Lite on overall quality and lands close to 2.5 Flash performance across key capability areas. It is a workhorse model for high-volume use cases, with improvements across audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion."},"google/gemini-3.1-pro-preview":{"maxTokens":64000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":2,"outputPrice":12,"cacheReadsPrice":0.19999999999999998,"description":"This model improves upon Gemini 2.5 Pro and is catered towards challenging tasks, especially those involving complex reasoning or agentic workflows. Improvements highlighted include use cases for coding, multi-step function calling, planning, reasoning, deep knowledge tasks, and instruction following."},"google/gemini-3.5-flash":{"maxTokens":64000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.5,"outputPrice":9,"cacheReadsPrice":0.15,"description":"Google's latest model, highly optimized for coding proficiency and parallel agentic execution loops. Defaults to medium thinking effort for faster and more cost-efficient responses."},"google/gemma-4-26b-a4b-it":{"maxTokens":131072,"contextWindow":262144,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.13,"outputPrice":0.39999999999999997,"description":"Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages."},"google/gemma-4-31b-it":{"maxTokens":131072,"contextWindow":262144,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.14,"outputPrice":0.39999999999999997,"description":"Gemma 4 31B is engineered to tackle the most demanding enterprise workloads and complex reasoning tasks. With an expansive 256K-token context window, the 31B model can effortlessly ingest entire codebases, and massive sets of images in a single prompt."},"inception/mercury-2":{"maxTokens":128000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":0.75,"cacheReadsPrice":0.024999999999999998,"description":"A diffusion-based reasoning LLM that generates text via parallel refinement (not token-by-token), delivering real-time latency with ~1k tokens/sec plus 128K context and built-in tool/JSON support."},"inception/mercury-coder-small":{"maxTokens":16384,"contextWindow":32000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":1,"description":"Mercury Coder Small is ideal for code generation, debugging, and refactoring tasks with minimal latency."},"interfaze/interfaze-beta":{"maxTokens":32000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.5,"outputPrice":3.5,"description":"Interfaze is an AI model built on a new architecture that merges specialized DNN/CNN models with LLMs for developer tasks that require deterministic output and high consistency like OCR, scraping, classification, STT and more."},"kwaipilot/kat-coder-pro-v1":{"maxTokens":32000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.03,"outputPrice":1.2,"cacheReadsPrice":0.06,"description":"KAT-Coder-Pro V1 is KwaiKAT's most advanced agentic coding model in the KwaiKAT series. Designed specifically for agentic coding tasks, it excels in real-world software engineering scenarios, achieving a remarkable 73.4% solve rate on the SWE-Bench Verified benchmark. KAT-Coder-Pro V1 delivers top-tier coding performance and has been rigorously tested by thousands of in-house engineers. The model has been optimized for tool-use capability, multi-turn interaction, instruction following, generalization and comprehensive capabilities through a multi-stage training process, including mid-training, supervised fine-tuning (SFT), reinforcement fine-tuning (RFT), and scalable agentic RL."},"kwaipilot/kat-coder-pro-v2":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.3,"outputPrice":1.2,"cacheReadsPrice":0.06,"description":"A high-performance edition designed for complex enterprise projects and SaaS integration."},"meituan/longcat-flash-chat":{"maxTokens":100000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"description":"LongCat-Flash-Chat is a high-throughput MoE chat model (128k context) optimized for agentic tasks."},"meituan/longcat-flash-thinking-2601":{"maxTokens":32768,"contextWindow":32768,"supportsImages":false,"supportsPromptCache":false,"description":"A version built for deep and general agentic thinking."},"meta/llama-3.1-70b":{"maxTokens":8192,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.72,"outputPrice":0.72,"description":"An update to Meta Llama 3 70B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities."},"meta/llama-3.1-8b":{"maxTokens":8192,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.22,"outputPrice":0.22,"description":"An update to Meta Llama 3 8B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities."},"meta/llama-3.2-11b":{"maxTokens":8192,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.16,"outputPrice":0.16,"description":"Instruction-tuned image reasoning generative model (text + images in / text out) optimized for visual recognition, image reasoning, captioning and answering general questions about the image."},"meta/llama-3.2-1b":{"maxTokens":8192,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.09999999999999999,"outputPrice":0.09999999999999999,"description":"Text-only model, supporting on-device use cases such as multilingual local knowledge retrieval, summarization, and rewriting."},"meta/llama-3.2-3b":{"maxTokens":8192,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":0.15,"description":"Text-only model, fine-tuned for supporting on-device use cases such as multilingual local knowledge retrieval, summarization, and rewriting."},"meta/llama-3.2-90b":{"maxTokens":8192,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.72,"outputPrice":0.72,"description":"Instruction-tuned image reasoning generative model (text + images in / text out) optimized for visual recognition, image reasoning, captioning and answering general questions about the image."},"meta/llama-3.3-70b":{"maxTokens":8192,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.72,"outputPrice":0.72,"description":"Where performance meets efficiency. This model supports high-performance conversational AI designed for content creation, enterprise applications, and research, offering advanced language understanding capabilities, including text summarization, classification, sentiment analysis, and code generation."},"meta/llama-4-maverick":{"maxTokens":8192,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.24,"outputPrice":0.9700000000000001,"description":"As a general purpose LLM, Llama 4 Maverick contains 17 billion active parameters, 128 experts, and 400 billion total parameters, offering high quality at a lower price compared to Llama 3.3 70B."},"meta/llama-4-scout":{"maxTokens":8192,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.16999999999999998,"outputPrice":0.66,"description":"Llama 4 Scout is the best multimodal model in the world in its class and is more powerful than our Llama 3 models, while fitting in a single H100 GPU. Additionally, Llama 4 Scout supports an industry-leading context window of up to 10M tokens."},"minimax/minimax-m2":{"maxTokens":205000,"contextWindow":205000,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.3,"outputPrice":1.2,"cacheWritesPrice":0.375,"cacheReadsPrice":0.03,"description":"MiniMax-M2 redefines efficiency for agents. It is a compact, fast, and cost-effective MoE model (230 billion total parameters with 10 billion active parameters) built for elite performance in coding and agentic tasks, all while maintaining powerful general intelligence."},"minimax/minimax-m2.1":{"maxTokens":131072,"contextWindow":204800,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.3,"outputPrice":1.2,"cacheWritesPrice":0.375,"cacheReadsPrice":0.03,"description":"MiniMax 2.1 is MiniMax's latest model, optimized specifically for robustness in coding, tool use, instruction following, and long-horizon planning."},"minimax/minimax-m2.1-lightning":{"maxTokens":131072,"contextWindow":204800,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.3,"outputPrice":2.4,"cacheWritesPrice":0.375,"cacheReadsPrice":0.03,"description":"MiniMax-M2.1-lightning is a faster version of MiniMax-M2.1, offering the same performance but with significantly higher throughput (output speed ~100 TPS, MiniMax-M2 output speed ~60 TPS)."},"minimax/minimax-m2.5":{"maxTokens":131000,"contextWindow":204800,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.3,"outputPrice":1.2,"cacheWritesPrice":0.375,"cacheReadsPrice":0.03,"description":"MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. It is capable of handling the entire development process of various complex systems. It covers full-stack projects across multiple platforms including Web, Android, iOS, Windows, and Mac, encompassing server-side APIs, functional logic, and databases."},"minimax/minimax-m2.5-highspeed":{"maxTokens":131000,"contextWindow":204800,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.6,"outputPrice":2.4,"cacheWritesPrice":0.375,"cacheReadsPrice":0.03,"description":"M2.5 highspeed: Same performance, faster and more agile (output speed approximately 100 tps)"},"minimax/minimax-m2.7":{"maxTokens":131000,"contextWindow":204800,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.3,"outputPrice":1.2,"cacheWritesPrice":0.375,"cacheReadsPrice":0.06,"description":"M2.7 delivers outstanding performance in real-world software engineering, including end-to-end full project delivery, log analysis and bug troubleshooting, code security, machine learning, and more."},"minimax/minimax-m2.7-highspeed":{"maxTokens":131100,"contextWindow":204800,"supportsImages":false,"supportsPromptCache":true,"inputPrice":0.6,"outputPrice":2.4,"cacheWritesPrice":0.375,"cacheReadsPrice":0.06,"description":"M2.7 Highspeed: Same performance, faster and more agile (output speed approximately 100 tps)"},"mistral/codestral":{"maxTokens":4000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.3,"outputPrice":0.8999999999999999,"description":"Mistral's cutting-edge language model for coding released end of July 2025, Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation."},"mistral/devstral-2":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.39999999999999997,"outputPrice":2,"description":"An enterprise-grade text model that excels at using tools to explore codebases, editing multiple files, and powering software engineering agents."},"mistral/devstral-small":{"maxTokens":64000,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.09999999999999999,"outputPrice":0.3,"description":"Devstral is an agentic LLM for software engineering tasks built under a collaboration between Mistral AI and All Hands AI 🙌. Devstral excels at using tools to explore codebases, editing multiple files and power software engineering agents."},"mistral/devstral-small-2":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.09999999999999999,"outputPrice":0.3,"description":"Our open source model that excels at using tools to explore codebases, editing multiple files, and powering software engineering agents."},"mistral/magistral-medium":{"maxTokens":64000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":2,"outputPrice":5,"description":"Complex thinking, backed by deep understanding, with transparent reasoning you can follow and verify. The model excels in maintaining high-fidelity reasoning across numerous languages, even when switching between languages mid-task."},"mistral/magistral-small":{"maxTokens":64000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.5,"outputPrice":1.5,"description":"Complex thinking, backed by deep understanding, with transparent reasoning you can follow and verify. The model excels in maintaining high-fidelity reasoning across numerous languages, even when switching between languages mid-task."},"mistral/ministral-14b":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.19999999999999998,"outputPrice":0.19999999999999998,"description":"Ministral 3 14B is the largest model in the Ministral 3 family, offering state-of-the-art capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. Optimized for local deployment, it delivers high performance across diverse hardware, including local setups."},"mistral/ministral-3b":{"maxTokens":4000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.09999999999999999,"outputPrice":0.09999999999999999,"description":"A compact, efficient model for on-device tasks like smart assistants and local analytics, offering low-latency performance."},"mistral/ministral-8b":{"maxTokens":4000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":0.15,"description":"A more powerful model with faster, memory-efficient inference, ideal for complex workflows and demanding edge applications."},"mistral/mistral-large-3":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.5,"outputPrice":1.5,"description":"Mistral Large 3 2512 is Mistral’s most capable model to date. It has a sparse mixture-of-experts architecture with 41B active parameters (675B total)."},"mistral/mistral-medium":{"maxTokens":64000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.39999999999999997,"outputPrice":2,"description":"Mistral Medium 3 delivers frontier performance while being an order of magnitude less expensive. For instance, the model performs at or above 90% of Claude Sonnet 3.7 on benchmarks across the board at a significantly lower cost."},"mistral/mistral-medium-3.5":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.5,"outputPrice":7.5,"description":"Mistral's frontier-class multimodal model optimized for agentic and coding use cases."},"mistral/mistral-nemo":{"maxTokens":131072,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.02,"outputPrice":0.04,"description":"12B model trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size."},"mistral/mistral-small":{"maxTokens":4000,"contextWindow":32000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.09999999999999999,"outputPrice":0.3,"description":"Mistral Small is the ideal choice for simple tasks that one can do in bulk - like Classification, Customer Support, or Text Generation. It offers excellent performance at an affordable price point."},"mistral/pixtral-12b":{"maxTokens":4000,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":0.15,"description":"A 12B model with image understanding capabilities in addition to text."},"mistral/pixtral-large":{"maxTokens":4000,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":2,"outputPrice":6,"description":"Pixtral Large is the second model in our multimodal family and demonstrates frontier-level image understanding. Particularly, the model is able to understand documents, charts and natural images, while maintaining the leading text-only understanding of Mistral Large 2."},"moonshotai/kimi-k2":{"maxTokens":131072,"contextWindow":131072,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.5700000000000001,"outputPrice":2.3,"description":"Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities."},"moonshotai/kimi-k2-thinking":{"maxTokens":262114,"contextWindow":262114,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.6,"outputPrice":2.5,"cacheReadsPrice":0.15,"description":"Kimi K2 Thinking is an advanced open-source thinking model by Moonshot AI. It can execute up to 200 – 300 sequential tool calls without human interference, reasoning coherently across hundreds of steps to solve complex problems. Built as a thinking agent, it reasons step by step while using tools, achieving state-of-the-art performance on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks, with major gains in reasoning, agentic search, coding, writing, and general capabilities."},"moonshotai/kimi-k2-thinking-turbo":{"maxTokens":262114,"contextWindow":262114,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.15,"outputPrice":8,"cacheReadsPrice":0.15,"description":"High-speed version of kimi-k2-thinking, suitable for scenarios requiring both deep reasoning and extremely fast responses"},"moonshotai/kimi-k2-turbo":{"maxTokens":16384,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.15,"outputPrice":8,"cacheReadsPrice":0.15,"description":"Kimi K2 Turbo is the high-speed version of kimi-k2, with the same model parameters as kimi-k2, but the output speed is increased to 60 tokens per second, with a maximum of 100 tokens per second, the context length is 256k"},"moonshotai/kimi-k2.5":{"maxTokens":262114,"contextWindow":262114,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.6,"outputPrice":3,"cacheReadsPrice":0.09999999999999999,"description":"kimi-k2.5 is Kimi's most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and agent tasks."},"moonshotai/kimi-k2.6":{"maxTokens":262000,"contextWindow":262000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.95,"outputPrice":4,"cacheReadsPrice":0.16,"description":"Kimi K2.6 demonstrates particularly strong performance in long-horizon coding tasks and produces professional-grade design with code and vision."},"morph/morph-v3-fast":{"maxTokens":16384,"contextWindow":81920,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.7999999999999999,"outputPrice":1.2,"description":"Morph offers a specialized AI model that applies code changes suggested by frontier models (like Claude or GPT-4o) to your existing code files FAST - 4500+ tokens/second. It acts as the final step in the AI coding workflow. Supports 16k input tokens and 16k output tokens."},"morph/morph-v3-large":{"maxTokens":16384,"contextWindow":81920,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.8999999999999999,"outputPrice":1.9,"description":"Morph offers a specialized AI model that applies code changes suggested by frontier models (like Claude or GPT-4o) to your existing code files FAST - 2500+ tokens/second. It acts as the final step in the AI coding workflow. Supports 16k input tokens and 16k output tokens."},"nvidia/nemotron-3-nano-30b-a3b":{"maxTokens":262144,"contextWindow":262144,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.049999999999999996,"outputPrice":0.24,"description":"NVIDIA Nemotron 3 Nano is an open reasoning model optimized for fast, cost-efficient inference. Built with a hybrid MoE and Mamba architecture and trained on NVIDIA-curated synthetic reasoning data, it delivers strong multi-step reasoning with stable latency and predictable performance for agentic and production workloads."},"nvidia/nemotron-3-super-120b-a12b":{"maxTokens":32000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":0.65,"description":"NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. It delivers up to 7x higher throughput, providing fast, cost-efficient inference for agentic tasks. Additionally, a long context window gives the model long-term memory, preventing AI agents from losing focus on long, multi-step tasks and ensuring high-accuracy results. Fully open with weights, datasets, and recipes, Super allows easy customization and secure deployment anywhere."},"nvidia/nemotron-nano-12b-v2-vl":{"maxTokens":131072,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.19999999999999998,"outputPrice":0.6,"description":"The model is an auto-regressive vision language model that uses an optimized transformer architecture. The model enables multi-image reasoning and video understanding, along with strong document intelligence, visual Q&A and summarization capabilities."},"nvidia/nemotron-nano-9b-v2":{"maxTokens":131072,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.06,"outputPrice":0.22999999999999998,"description":"NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so.\\"},"openai/gpt-3.5-turbo":{"maxTokens":4096,"contextWindow":16385,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.5,"outputPrice":1.5,"description":"OpenAI's most capable and cost effective model in the GPT-3.5 family optimized for chat purposes, but also works well for traditional completions tasks."},"openai/gpt-3.5-turbo-instruct":{"maxTokens":4096,"contextWindow":8192,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.5,"outputPrice":2,"description":"Similar capabilities as GPT-3 era models. Compatible with legacy Completions endpoint and not Chat Completions."},"openai/gpt-4-turbo":{"maxTokens":4096,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":10,"outputPrice":30,"description":"gpt-4-turbo from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It has a knowledge cutoff of April 2023 and a 128,000 token context window."},"openai/gpt-4.1":{"maxTokens":32768,"contextWindow":1047576,"supportsImages":true,"supportsPromptCache":false,"inputPrice":2,"outputPrice":8,"cacheReadsPrice":0.5,"description":"GPT 4.1 is OpenAI's flagship model for complex tasks. It is well suited for problem solving across domains."},"openai/gpt-4.1-mini":{"maxTokens":32768,"contextWindow":1047576,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.39999999999999997,"outputPrice":1.5999999999999999,"cacheReadsPrice":0.09999999999999999,"description":"GPT 4.1 mini provides a balance between intelligence, speed, and cost that makes it an attractive model for many use cases."},"openai/gpt-4.1-nano":{"maxTokens":32768,"contextWindow":1047576,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.09999999999999999,"outputPrice":0.39999999999999997,"cacheReadsPrice":0.024999999999999998,"description":"GPT-4.1 nano is the fastest, most cost-effective GPT 4.1 model."},"openai/gpt-4o":{"maxTokens":16384,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":2.5,"outputPrice":10,"cacheReadsPrice":1.25,"description":"GPT-4o from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It matches GPT-4 Turbo performance with a faster and cheaper API."},"openai/gpt-4o-mini":{"maxTokens":16384,"contextWindow":128000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":0.6,"cacheReadsPrice":0.075,"description":"GPT-4o mini from OpenAI is their most advanced and cost-efficient small model. It is multi-modal (accepting text or image inputs and outputting text) and has higher intelligence than gpt-3.5-turbo but is just as fast."},"openai/gpt-4o-mini-search-preview":{"maxTokens":16384,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.15,"outputPrice":0.6,"description":"GPT-4o mini Search Preview is a specialized model trained to understand and execute web search queries with the Chat Completions API. In addition to token fees, web search queries have a fee per tool call."},"openai/gpt-5":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":10,"cacheReadsPrice":0.125,"description":"GPT-5 is OpenAI's flagship language model that excels at complex reasoning, broad real-world knowledge, code-intensive, and multi-step agentic tasks."},"openai/gpt-5-chat":{"maxTokens":16384,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":10,"cacheReadsPrice":0.125,"description":"GPT-5 Chat points to the GPT-5 snapshot currently used in ChatGPT."},"openai/gpt-5-codex":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":10,"cacheReadsPrice":0.125,"description":"GPT-5-Codex is a version of GPT-5 optimized for agentic coding tasks in Codex or similar environments."},"openai/gpt-5-mini":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":2,"cacheReadsPrice":0.024999999999999998,"description":"GPT-5 mini is a cost optimized model that excels at reasoning/chat tasks. It offers an optimal balance between speed, cost, and capability."},"openai/gpt-5-nano":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.049999999999999996,"outputPrice":0.39999999999999997,"cacheReadsPrice":0.005,"description":"GPT-5 nano is a high throughput model that excels at simple instruction or classification tasks."},"openai/gpt-5-pro":{"maxTokens":272000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":15,"outputPrice":120,"description":"GPT-5 pro uses more compute to think harder and provide consistently better answers. Since GPT-5 pro is designed to tackle tough problems, some requests may take several minutes to finish."},"openai/gpt-5.1-codex":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":10,"cacheReadsPrice":0.125,"description":"GPT-5.1-Codex is a version of GPT-5.1 optimized for agentic coding tasks in Codex or similar environments."},"openai/gpt-5.1-codex-max":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":10,"cacheReadsPrice":0.125,"description":"GPT‑5.1-Codex-Max is purpose-built for agentic coding."},"openai/gpt-5.1-codex-mini":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.25,"outputPrice":2,"cacheReadsPrice":0.024999999999999998,"description":"GPT-5.1 Codex mini is a smaller, faster, and cheaper version of GPT-5.1 Codex."},"openai/gpt-5.1-instant":{"maxTokens":16384,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":10,"cacheReadsPrice":0.125,"description":"GPT-5.1 Instant (or GPT-5.1 chat) is a warmer and more conversational version of GPT-5-chat, with improved instruction following and adaptive reasoning for deciding when to think before responding."},"openai/gpt-5.1-thinking":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":10,"cacheReadsPrice":0.125,"description":"An upgraded version of GPT-5 that adapts thinking time more precisely to the question to spend more time on complex questions and respond more quickly to simpler tasks."},"openai/gpt-5.2":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.75,"outputPrice":14,"cacheReadsPrice":0.175,"description":"GPT-5.2 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks."},"openai/gpt-5.2-chat":{"maxTokens":16384,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.75,"outputPrice":14,"cacheReadsPrice":0.175,"description":"The model powering ChatGPT is gpt-5.2-chat-latest: this is OpenAI's best general-purpose model, part of the GPT-5 flagship model family."},"openai/gpt-5.2-codex":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.75,"outputPrice":14,"cacheReadsPrice":0.175,"description":"GPT‑5.2-Codex is a version of GPT‑5.2⁠ further optimized for agentic coding in Codex, including improvements on long-horizon work through context compaction, stronger performance on large code changes like refactors and migrations, improved performance in Windows environments, and significantly stronger cybersecurity capabilities."},"openai/gpt-5.2-pro":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":21,"outputPrice":168,"description":"Version of GPT-5.2 that produces smarter and more precise responses."},"openai/gpt-5.3-chat":{"maxTokens":16384,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.75,"outputPrice":14,"cacheReadsPrice":0.175,"description":"The model powering ChatGPT is gpt-5.3-chat-latest: this is OpenAI's best general-purpose model, part of the GPT-5 flagship model family."},"openai/gpt-5.3-codex":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.75,"outputPrice":14,"cacheReadsPrice":0.175,"description":"GPT-5.3-Codex advances both the frontier coding performance of GPT‑5.2-Codex and the reasoning and professional knowledge capabilities of GPT‑5.2, together in one model, which is also 25% faster. This enables it to take on long-running tasks that involve research, tool use, and complex execution."},"openai/gpt-5.4":{"maxTokens":128000,"contextWindow":1050000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":2.5,"outputPrice":15,"cacheReadsPrice":0.25,"description":"GPT-5.4 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks."},"openai/gpt-5.4-mini":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.75,"outputPrice":4.5,"cacheReadsPrice":0.075,"description":"GPT-5.4 Mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads."},"openai/gpt-5.4-nano":{"maxTokens":128000,"contextWindow":400000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.19999999999999998,"outputPrice":1.25,"cacheReadsPrice":0.02,"description":"GPT-5.4 Nano is designed for tasks where speed and cost matter most like classification, data extraction, ranking, and sub-agents."},"openai/gpt-5.4-pro":{"maxTokens":128000,"contextWindow":1050000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":30,"outputPrice":180,"description":"GPT-5.4 Pro uses more compute to think harder and provide consistently better answers. It's designed to tackle tough problems."},"openai/gpt-5.5":{"maxTokens":128000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":5,"outputPrice":30,"cacheReadsPrice":0.5,"description":"GPT‑5.5 understands what you’re trying to do faster and can carry more of the work itself. It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Instead of carefully managing every step, you can give GPT‑5.5 a messy, multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going."},"openai/gpt-5.5-pro":{"maxTokens":128000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":30,"outputPrice":180,"description":""},"openai/gpt-oss-120b":{"maxTokens":131000,"contextWindow":131072,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.35,"outputPrice":0.75,"cacheReadsPrice":0.25,"description":"This model excels at efficient reasoning across science, math, and coding applications. It’s ideal for real-time coding assistance, processing large documents for Q&A and summarization, agentic research workflows, and regulated on-premises workloads."},"openai/gpt-oss-20b":{"maxTokens":8192,"contextWindow":131072,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.049999999999999996,"outputPrice":0.19999999999999998,"description":"A compact, open-weight language model optimized for low-latency and resource-constrained environments, including local and edge deployments."},"openai/gpt-oss-safeguard-20b":{"maxTokens":65536,"contextWindow":131072,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.075,"outputPrice":0.3,"cacheReadsPrice":0.037,"description":"OpenAI's first open weight reasoning model specifically trained for safety classification tasks. Fine-tuned from GPT-OSS, this model helps classify text content based on customizable policies, enabling bring-your-own-policy Trust & Safety AI where your own taxonomy, definitions, and thresholds guide classification decisions."},"openai/o1":{"maxTokens":100000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":15,"outputPrice":60,"cacheReadsPrice":7.5,"description":"o1 is OpenAI's flagship reasoning model, designed for complex problems that require deep thinking. It provides strong reasoning capabilities with improved accuracy for complex multi-step tasks."},"openai/o3":{"maxTokens":100000,"contextWindow":200000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":2,"outputPrice":8,"cacheReadsPrice":0.5,"description":"OpenAI's o3 is their most powerful reasoning model, setting new state-of-the-art benchmarks in coding, math, science, and visual perception. It excels at complex queries requiring multi-faceted analysis, with particular strength in analyzing images, charts, and graphics."},"openai/o3-deep-research":{"maxTokens":100000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":10,"outputPrice":40,"cacheReadsPrice":2.5,"description":"o3-deep-research is OpenAI's most advanced model for deep research, designed to tackle complex, multi-step research tasks. It can search and synthesize information from across the internet as well as from your own data—brought in through MCP connectors."},"openai/o3-mini":{"maxTokens":100000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.1,"outputPrice":4.4,"cacheReadsPrice":0.55,"description":"o3-mini is OpenAI's most recent small reasoning model, providing high intelligence at the same cost and latency targets of o1-mini."},"openai/o3-pro":{"maxTokens":100000,"contextWindow":200000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":20,"outputPrice":80,"description":"The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently better answers."},"openai/o4-mini":{"maxTokens":100000,"contextWindow":200000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":1.1,"outputPrice":4.4,"cacheReadsPrice":0.275,"description":"OpenAI's o4-mini delivers fast, cost-efficient reasoning with exceptional performance for its size, particularly excelling in math (best-performing on AIME benchmarks), coding, and visual tasks."},"perplexity/sonar":{"maxTokens":8000,"contextWindow":127000,"supportsImages":false,"supportsPromptCache":false,"description":"Perplexity's lightweight offering with search grounding, quicker and cheaper than Sonar Pro."},"perplexity/sonar-pro":{"maxTokens":8000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":false,"description":"Perplexity's premier offering with search grounding, supporting advanced queries and follow-ups."},"perplexity/sonar-reasoning-pro":{"maxTokens":8000,"contextWindow":127000,"supportsImages":false,"supportsPromptCache":false,"description":"A premium reasoning-focused model that outputs Chain of Thought (CoT) in responses, providing comprehensive explanations with enhanced search capabilities and multiple search queries per request."},"stepfun/step-3.7-flash":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.19999999999999998,"outputPrice":1.15,"cacheReadsPrice":0.04,"description":"StepFun’s flagship multimodal reasoning model. Powered by a 198B-parameter / 11B-activation sparse MoE architecture, with native support for image and video understanding."},"xai/grok-4.1-fast-non-reasoning":{"maxTokens":1000000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.19999999999999998,"outputPrice":0.5,"cacheReadsPrice":0.049999999999999996,"description":""},"xai/grok-4.1-fast-reasoning":{"maxTokens":1000000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.19999999999999998,"outputPrice":0.5,"cacheReadsPrice":0.049999999999999996,"description":""},"xai/grok-4.20-multi-agent":{"maxTokens":2000000,"contextWindow":2000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":2.5,"cacheReadsPrice":0.19999999999999998,"description":"Multiple agents collaborate in parallel to perform deep research tasks."},"xai/grok-4.20-multi-agent-beta":{"maxTokens":2000000,"contextWindow":2000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":2.5,"cacheReadsPrice":0.19999999999999998,"description":"Multiple agents collaborate in parallel to perform deep research tasks."},"xai/grok-4.20-non-reasoning":{"maxTokens":2000000,"contextWindow":2000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":2.5,"cacheReadsPrice":0.19999999999999998,"description":"Grok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses."},"xai/grok-4.20-non-reasoning-beta":{"maxTokens":2000000,"contextWindow":2000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":2.5,"cacheReadsPrice":0.19999999999999998,"description":"Grok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses."},"xai/grok-4.20-reasoning":{"maxTokens":2000000,"contextWindow":2000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":2.5,"cacheReadsPrice":0.19999999999999998,"description":"Grok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses."},"xai/grok-4.20-reasoning-beta":{"maxTokens":2000000,"contextWindow":2000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":2.5,"cacheReadsPrice":0.19999999999999998,"description":"Grok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses."},"xai/grok-4.3":{"maxTokens":1000000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.25,"outputPrice":2.5,"cacheReadsPrice":0.19999999999999998,"description":"Grok 4.3 is a new model matching the scale of Grok 4.20 with an improved architecture and a December 2025 knowledge cutoff."},"xai/grok-build-0.1":{"maxTokens":256000,"contextWindow":256000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1,"outputPrice":2,"cacheReadsPrice":0.19999999999999998,"description":"xAI's fast coding model trained specifically for agentic coding."},"xiaomi/mimo-v2-flash":{"maxTokens":32000,"contextWindow":262144,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.09999999999999999,"outputPrice":0.3,"cacheReadsPrice":0.01,"description":"Xiaomi MiMo-V2-Flash is a proprietary MoE model developed by Xiaomi, designed for extreme inference efficiency with 309B total parameters (15B active). By incorporating an innovative Hybrid attention architecture and multi-layer MTP inference acceleration, it ranks among the top 2 global open-source models across multiple Agent benchmarks."},"xiaomi/mimo-v2-pro":{"maxTokens":128000,"contextWindow":1000000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1,"outputPrice":3,"cacheReadsPrice":0.19999999999999998,"description":"Xiaomi MiMo-V2-Pro is built for demanding real-world Agent workflows. It has over 1T total parameters, with 42B active parameters, uses an innovative hybrid attention architecture, and supports an ultra-long context window of up to 1M tokens."},"xiaomi/mimo-v2.5":{"maxTokens":131100,"contextWindow":1050000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.14,"outputPrice":0.28,"cacheReadsPrice":0.0028,"description":"A native full-modal model supporting text, image, video, and audio understanding, with powerful Agent capabilities."},"xiaomi/mimo-v2.5-pro":{"maxTokens":131000,"contextWindow":1050000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.435,"outputPrice":0.87,"cacheReadsPrice":0.0036,"description":"MiMo V2.5 Pro delivers significant improvements over its predecessor, MiMo-V2-Pro, in general agentic capabilities, complex software engineering, and long-horizon tasks. MiMo-V2.5-Pro is a 1.02T-parameter Mixture-of-Experts model with 42B active parameters, built on a hybrid-attention architecture with a 1M-token context window."},"zai/glm-4.5":{"maxTokens":96000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.6,"outputPrice":2.2,"cacheReadsPrice":0.11,"description":"GLM-4.5 and GLM-4.5-Air are our latest flagship models, purpose-built as foundational models for agent-oriented applications. Both leverage a Mixture-of-Experts (MoE) architecture. GLM-4.5 has a total parameter count of 355B with 32B active parameters per forward pass, while GLM-4.5-Air adopts a more streamlined design with 106B total parameters and 12B active parameters."},"zai/glm-4.5-air":{"maxTokens":96000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.19999999999999998,"outputPrice":1.1,"cacheReadsPrice":0.03,"description":"GLM-4.5 and GLM-4.5-Air are our latest flagship models, purpose-built as foundational models for agent-oriented applications. Both leverage a Mixture-of-Experts (MoE) architecture. GLM-4.5 has a total parameter count of 355B with 32B active parameters per forward pass, while GLM-4.5-Air adopts a more streamlined design with 106B total parameters and 12B active parameters."},"zai/glm-4.5v":{"maxTokens":16000,"contextWindow":66000,"supportsImages":true,"supportsPromptCache":false,"inputPrice":0.6,"outputPrice":1.7999999999999998,"cacheReadsPrice":0.11,"description":"Built on the GLM-4.5-Air base model, GLM-4.5V inherits proven techniques from GLM-4.1V-Thinking while achieving effective scaling through a powerful 106B-parameter MoE architecture."},"zai/glm-4.6":{"maxTokens":96000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.6,"outputPrice":2.2,"cacheReadsPrice":0.11,"description":"As the latest iteration in the GLM series, GLM-4.6 achieves comprehensive enhancements across multiple domains, including real-world coding, long-context processing, reasoning, searching, writing, and agentic applications."},"zai/glm-4.6v":{"maxTokens":24000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.3,"outputPrice":0.8999999999999999,"cacheReadsPrice":0.049999999999999996,"description":"GLM-4.6V series are Z.ai’s iterations in a multimodal large language model. GLM-4.6V scales its context window to 128k tokens in training, and achieves SoTA performance in visual understanding among models of similar parameter scales."},"zai/glm-4.6v-flash":{"maxTokens":24000,"contextWindow":128000,"supportsImages":false,"supportsPromptCache":false,"description":"For local deployment and low-latency applications. GLM-4.6V series are Z.ai’s iterations in a multimodal large language model. GLM-4.6V scales its context window to 128k tokens in training, and achieves SoTA performance in visual understanding among models of similar parameter scales."},"zai/glm-4.7":{"maxTokens":40000,"contextWindow":131000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":2.25,"outputPrice":2.75,"cacheReadsPrice":2.25,"description":"GLM-4.7 is Z.ai’s latest flagship model, with major upgrades focused on two key areas: stronger coding capabilities and more stable multi-step reasoning and execution."},"zai/glm-4.7-flash":{"maxTokens":131000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.07,"outputPrice":0.39999999999999997,"description":"GLM-4.7-Flash balances high performance with efficiency, making it the perfect lightweight deployment option. Beyond coding, it is also recommended for creative writing, translation, long-context tasks, and roleplay."},"zai/glm-4.7-flashx":{"maxTokens":128000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":0.06,"outputPrice":0.39999999999999997,"cacheReadsPrice":0.01,"description":" GLM-4.7-Flash balances high performance with efficiency, making it the perfect lightweight deployment option. "},"zai/glm-5":{"maxTokens":131100,"contextWindow":202800,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1,"outputPrice":3.1999999999999997,"cacheReadsPrice":0.19999999999999998,"description":"GLM-5 is Zai’s new-generation flagship foundation model, designed for Agentic Engineering, capable of providing reliable productivity in complex system engineering and long-range Agent tasks. In terms of Coding and Agent capabilities, GLM-5 has achieved state-of-the-art (SOTA) performance in open source, with its usability in real programming scenarios approaching that of Claude Opus 4.5."},"zai/glm-5-turbo":{"maxTokens":131100,"contextWindow":202800,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.2,"outputPrice":4,"cacheReadsPrice":0.24,"description":"GLM 5 Turbo is a foundation model deeply optimized for the OpenClaw scenario. It has been specifically optimized for the core requirements of OpenClaw tasks since the training phase, enhancing key capabilities such as tool invocation, command following, timed and persistent tasks, and long-chain execution."},"zai/glm-5.1":{"maxTokens":64000,"contextWindow":202800,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.4,"outputPrice":4.4,"cacheReadsPrice":0.26,"description":"GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours—autonomously planning, executing, and improving itself throughout the process—ultimately delivering complete, engineering-grade results."},"zai/glm-5v-turbo":{"maxTokens":128000,"contextWindow":200000,"supportsImages":false,"supportsPromptCache":false,"inputPrice":1.2,"outputPrice":4,"cacheReadsPrice":0.24,"description":"GLM-5V-Turbo is Z.AI’s first multimodal coding foundation model, built for vision-based coding tasks. It can natively process multimodal inputs such as images, video, and text, while also excelling at long-horizon planning, complex coding, and action execution. Deeply optimized for agent workflows, it works seamlessly with agents such as Claude Code and OpenClaw to complete the full loop of “understand the environment → plan actions → execute tasks”."}}