Complete model overview
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| GPT-6 Luna Pro | openai/gpt-6-luna-pro | 1.1M | file, image, text | text | On | $0.1000 | $0.5000 | GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Le... |
| GPT-6 Luna Pro (batch) | openai/gpt-6-luna-pro:batch | 1.1M | file, image, text | text | On | $0.0500 | $0.2500 | GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Le... |
| GPT-6 Luna | openai/gpt-6-luna | 1.1M | file, image, text | text | On | $0.1000 | $0.5000 | GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive ... |
| GPT-6 Luna (batch) | openai/gpt-6-luna:batch | 1.1M | file, image, text | text | On | $0.0500 | $0.2500 | GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive ... |
| GPT-6 Sol Pro | openai/gpt-6-sol-pro | 1.1M | file, image, text | text | On | $2.00 | $10.0 | GPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Lear... |
| GPT-6 Sol Pro (batch) | openai/gpt-6-sol-pro:batch | 1.1M | file, image, text | text | On | $1.00 | $5.00 | GPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Lear... |
| GPT-6 Sol | openai/gpt-6-sol | 1.1M | file, image, text | text | On | $2.00 | $10.0 | GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier.... |
| GPT-6 Sol (batch) | openai/gpt-6-sol:batch | 1.1M | file, image, text | text | On | $1.00 | $5.00 | GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier.... |
| GPT Astra Latest | ~openai/gpt-astra-latest | 1.1M | file, image, text | text | Always on | $10.0 | $50.0 | This model always redirects to the latest model in the GPT Astra family. |
| GPT Sol Latest | ~openai/gpt-sol-latest | 1.1M | file, image, text | text | On | $2.00 | $10.0 | This model always redirects to the latest model in the GPT Sol family. |
| GPT Terra Latest | ~openai/gpt-terra-latest | 1.1M | file, image, text | text | On | $2.00 | $12.0 | This model always redirects to the latest model in the GPT Terra family. |
| GPT Luna Latest | ~openai/gpt-luna-latest | 1.1M | file, image, text | text | On | $0.1000 | $0.5000 | This model always redirects to the latest model in the GPT Luna family. |
| GPT-6 Astra | openai/gpt-6-astra | 1.1M | file, image, text | text | Always on | $10.0 | $50.0 | GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scienti... |
| GPT-6 Astra (batch) | openai/gpt-6-astra:batch | 1.1M | file, image, text | text | Always on | $5.00 | $25.0 | GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scienti... |
| GPT-6 Astra Pro | openai/gpt-6-astra-pro | 1.1M | file, image, text | text | Always on | $10.0 | $50.0 | GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. ... |
| GPT-6 Astra Pro (batch) | openai/gpt-6-astra-pro:batch | 1.1M | file, image, text | text | Always on | $5.00 | $25.0 | GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. ... |
| GPT-5.6 Luna Pro | openai/gpt-5.6-luna-pro | 1.1M | file, image, text | text | On | $0.2000 | $1.20 | GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks... |
| GPT-5.6 Luna Pro (batch) | openai/gpt-5.6-luna-pro:batch | 1.1M | file, image, text | text | On | $0.1000 | $0.6000 | GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks... |
| GPT-5.6 Luna | openai/gpt-5.6-luna | 1.1M | file, image, text | text | On | $0.2000 | $1.20 | GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classific... |
| GPT-5.6 Luna (batch) | openai/gpt-5.6-luna:batch | 1.1M | file, image, text | text | On | $0.1000 | $0.6000 | GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classific... |
| GPT-5.6 Terra Pro | openai/gpt-5.6-terra-pro | 1.1M | file, image, text | text | On | $2.00 | $12.0 | GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tas... |
| GPT-5.6 Terra Pro (batch) | openai/gpt-5.6-terra-pro:batch | 1.1M | file, image, text | text | On | $1.00 | $6.00 | GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tas... |
| GPT-5.6 Terra | openai/gpt-5.6-terra | 1.1M | file, image, text | text | On | $2.00 | $12.0 | GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited ... |
| GPT-5.6 Terra (batch) | openai/gpt-5.6-terra:batch | 1.1M | file, image, text | text | On | $1.00 | $6.00 | GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited ... |
| GPT-5.6 Sol Pro | openai/gpt-5.6-sol-pro | 1.1M | file, image, text | text | On | $2.00 | $10.0 | GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. ... |
| GPT-5.6 Sol Pro (batch) | openai/gpt-5.6-sol-pro:batch | 1.1M | file, image, text | text | On | $1.00 | $5.00 | GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. ... |
| GPT-5.6 Sol | openai/gpt-5.6-sol | 1.1M | file, image, text | text | On | $2.00 | $10.0 | GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly s... |
| GPT-5.6 Sol (batch) | openai/gpt-5.6-sol:batch | 1.1M | file, image, text | text | On | $1.00 | $5.00 | GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly s... |
| GPT Chat Latest | openai/gpt-chat-latest | 400K | text, image, file | text | Off | $5.00 | $30.0 | GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rol... |
| GPT Mini Latest | ~openai/gpt-mini-latest | 400K | file, image, text | text | Off | $0.7500 | $4.50 | This model always redirects to the latest model in the GPT Mini family. |
| GPT-5.5 Pro | openai/gpt-5.5-pro | 1.1M | file, image, text | text | Always on | $30.0 | $180.0 | GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token con... |
| GPT-5.5 Pro (batch) | openai/gpt-5.5-pro:batch | 1.1M | file, image, text | text | Always on | $15.0 | $90.0 | GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token con... |
| GPT-5.5 | openai/gpt-5.5 | 1.1M | file, image, text | text | On | $5.00 | $30.0 | GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and i... |
| GPT-5.5 (batch) | openai/gpt-5.5:batch | 1.1M | file, image, text | text | On | $2.50 | $15.0 | GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and i... |
| GPT-5.4 Nano | openai/gpt-5.4-nano | 400K | file, image, text | text | Off | $0.2000 | $1.25 | GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports... |
| GPT-5.4 Nano (batch) | openai/gpt-5.4-nano:batch | 400K | file, image, text | text | Off | $0.1000 | $0.6250 | GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports... |
| GPT-5.4 Mini | openai/gpt-5.4-mini | 400K | file, image, text | text | Off | $0.7500 | $4.50 | GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and i... |
| GPT-5.4 Mini (batch) | openai/gpt-5.4-mini:batch | 400K | file, image, text | text | Off | $0.3750 | $2.25 | GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and i... |
| GPT-5.4 Pro | openai/gpt-5.4-pro | 1.1M | text, image, file | text | Always on | $30.0 | $180.0 | GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes ... |
| GPT-5.4 Pro (batch) | openai/gpt-5.4-pro:batch | 1.1M | text, image, file | text | Always on | $15.0 | $90.0 | GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes ... |
| GPT-5.4 | openai/gpt-5.4 | 1.1M | text, image, file | text | Off | $2.50 | $15.0 | GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, ... |
| GPT-5.4 (batch) | openai/gpt-5.4:batch | 1.1M | text, image, file | text | Off | $1.25 | $7.50 | GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, ... |
| GPT-5.3-Codex | openai/gpt-5.3-codex | 400K | text, image, file | text | Off | $1.75 | $14.0 | GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broade... |
| GPT Audio | openai/gpt-audio | 128K | text, audio | text, audio | Off | $2.50 | $10.0 | The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices ... |
| GPT Audio Mini | openai/gpt-audio-mini | 128K | text, audio | text, audio | Off | $0.6000 | $2.40 | A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consi... |
| GPT-5.2-Codex | openai/gpt-5.2-codex | 400K | text, image | text | Always on | $1.75 | $14.0 | GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive dev... |
| GPT-5.2 Chat | openai/gpt-5.2-chat | 128K | file, image, text | text | Off | $1.75 | $14.0 | GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligen... |
| GPT-5.2 Pro | openai/gpt-5.2-pro | 400K | image, text, file | text | Always on | $21.0 | $168.0 | GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimize... |
| GPT-5.2 Pro (batch) | openai/gpt-5.2-pro:batch | 400K | image, text, file | text | Always on | $10.5 | $84.0 | GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimize... |
| GPT-5.2 | openai/gpt-5.2 | 400K | file, image, text | text | Off | $1.75 | $14.0 | GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses ada... |
| GPT-5.2 (batch) | openai/gpt-5.2:batch | 400K | file, image, text | text | Off | $0.8750 | $7.00 | GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses ada... |
| GPT-5.1-Codex-Max | openai/gpt-5.1-codex-max | 400K | text, image | text | Always on | $1.25 | $10.0 | GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updat... |
| GPT-5.1 | openai/gpt-5.1 | 400K | image, text, file | text | On | $1.25 | $10.0 | GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a mor... |
| GPT-5.1 (batch) | openai/gpt-5.1:batch | 400K | image, text, file | text | On | $0.6250 | $5.00 | GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a mor... |
| GPT-5.1-Codex | openai/gpt-5.1-codex | 400K | text, image | text | Always on | $1.25 | $10.0 | GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive develop... |
| GPT-5.1-Codex-Mini | openai/gpt-5.1-codex-mini | 400K | image, text | text | Off | $0.2500 | $2.00 | GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex |
| Text Embedding Ada 002 | openai/text-embedding-ada-002 | 8K | text | embeddings | Off | $0.1000 | FREE | text-embedding-ada-002 is OpenAI's legacy text embedding model. |
| Text Embedding Ada 002 (batch) | openai/text-embedding-ada-002:batch | 8K | text | embeddings | Off | $0.0500 | FREE | text-embedding-ada-002 is OpenAI's legacy text embedding model. |
| Text Embedding 3 Large | openai/text-embedding-3-large | 8K | text | embeddings | Off | $0.1300 | FREE | text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english tasks. Embeddings are a numerical representation of t... |
| Text Embedding 3 Large (batch) | openai/text-embedding-3-large:batch | 8K | text | embeddings | Off | $0.0650 | FREE | text-embedding-3-large is OpenAI's most capable embedding model for both english and non-english tasks. Embeddings are a numerical representation of t... |
| Text Embedding 3 Small | openai/text-embedding-3-small | 8K | text | embeddings | Off | $0.0200 | FREE | text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text tha... |
| Text Embedding 3 Small (batch) | openai/text-embedding-3-small:batch | 8K | text | embeddings | Off | $0.0200 | FREE | text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text tha... |
| gpt-oss-safeguard-20b | openai/gpt-oss-safeguard-20b | 131K | text | text | Always on | $0.0750 | $0.3000 | gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE) model o... |
| GPT-5 Pro | openai/gpt-5-pro | 400K | image, text, file | text | Always on | $15.0 | $120.0 | GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex ta... |
| GPT-5 Pro (batch) | openai/gpt-5-pro:batch | 400K | image, text, file | text | Always on | $7.50 | $60.0 | GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex ta... |
| GPT-5 | openai/gpt-5 | 400K | text, image, file | text | Always on | $1.25 | $10.0 | GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks ... |
| GPT-5 (batch) | openai/gpt-5:batch | 400K | text, image, file | text | Always on | $0.6250 | $5.00 | GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks ... |
| GPT-5 Mini | openai/gpt-5-mini | 400K | text, image, file | text | Always on | $0.2500 | $2.00 | GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tun... |
| GPT-5 Mini (batch) | openai/gpt-5-mini:batch | 400K | text, image, file | text | Always on | $0.1250 | $1.00 | GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tun... |
| GPT-5 Nano | openai/gpt-5-nano | 400K | text, image, file | text | Always on | $0.0500 | $0.4000 | GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environme... |
| GPT-5 Nano (batch) | openai/gpt-5-nano:batch | 400K | text, image, file | text | Always on | $0.0250 | $0.2000 | GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environme... |
| gpt-oss-120b | openai/gpt-oss-120b | 131K | text | text | Always on | $0.0300 | $0.1700 | gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-p... |
| gpt-oss-20b | openai/gpt-oss-20b | 131K | text | text | Always on | $0.0180 | $0.0900 | gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture wit... |
| gpt-oss-20b (batch) | openai/gpt-oss-20b:batch | 131K | text | text | Always on | $0.0240 | $0.1120 | gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture wit... |
| o3 Pro | openai/o3-pro | 200K | text, file, image | text | Off | $20.0 | $80.0 | The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more c... |
| o4 Mini High | openai/o4-mini-high | 200K | image, text, file | text | Always on | $1.10 | $4.40 | OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in ... |
| o3 | openai/o3 | 200K | image, text, file | text | Off | $2.00 | $8.00 | o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels a... |
| o3 (batch) | openai/o3:batch | 200K | image, text, file | text | Off | $1.00 | $4.00 | o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels a... |
| o4 Mini | openai/o4-mini | 200K | image, text, file | text | Off | $1.10 | $4.40 | OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agen... |
| o4 Mini (batch) | openai/o4-mini:batch | 200K | image, text, file | text | Off | $0.5500 | $2.20 | OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agen... |
| GPT-4.1 | openai/gpt-4.1 | 1.0M | image, text, file | text | Off | $2.00 | $8.00 | GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. ... |
| GPT-4.1 (batch) | openai/gpt-4.1:batch | 1.0M | image, text, file | text | Off | $1.00 | $4.00 | GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. ... |
| GPT-4.1 Mini | openai/gpt-4.1-mini | 1.0M | image, text, file | text | Off | $0.4000 | $1.60 | GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token... |
| GPT-4.1 Mini (batch) | openai/gpt-4.1-mini:batch | 1.0M | image, text, file | text | Off | $0.2000 | $0.8000 | GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token... |
| GPT-4.1 Nano | openai/gpt-4.1-nano | 1.0M | image, text, file | text | Off | $0.1000 | $0.4000 | For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a smal... |
| GPT-4.1 Nano (batch) | openai/gpt-4.1-nano:batch | 1.0M | image, text, file | text | Off | $0.0500 | $0.2000 | For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a smal... |
| o3 Mini High | openai/o3-mini-high | 200K | text, file | text | Always on | $1.10 | $4.40 | OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model opti... |
| o3 Mini | openai/o3-mini | 200K | text, file | text | Off | $1.10 | $4.40 | OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This... |
| o3 Mini (batch) | openai/o3-mini:batch | 200K | text, file | text | Off | $0.5500 | $2.20 | OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This... |
| o1 | openai/o1 | 200K | text, image, file | text | Off | $15.0 | $60.0 | The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with l... |
| GPT-4o (2024-11-20) | openai/gpt-4o-2024-11-20 | 128K | text, image, file | text | Off | $2.50 | $10.0 | The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance &... |
| GPT-4o (2024-08-06) | openai/gpt-4o-2024-08-06 | 128K | text, image, file | text | Off | $2.50 | $10.0 | The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Re... |
| GPT-4o-mini | openai/gpt-4o-mini | 128K | text, image, file | text | Off | $0.1500 | $0.6000 | GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most... |
| GPT-4o-mini (2024-07-18) | openai/gpt-4o-mini-2024-07-18 | 128K | text, image, file | text | Off | $0.1500 | $0.6000 | GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most... |
| GPT-4o-mini (batch) | openai/gpt-4o-mini:batch | 128K | text, image, file | text | Off | $0.0750 | $0.3000 | GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most... |
| GPT-4o | openai/gpt-4o | 128K | text, image, file | text | Off | $2.50 | $10.0 | GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [... |
| GPT-4o (2024-05-13) | openai/gpt-4o-2024-05-13 | 128K | text, image, file | text | Off | $5.00 | $15.0 | GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [... |
| GPT-4o (batch) | openai/gpt-4o:batch | 128K | text, image, file | text | Off | $1.25 | $5.00 | GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [... |
| GPT-4 Turbo | openai/gpt-4-turbo | 128K | text, image | text | Off | $10.0 | $30.0 | The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023. |
| GPT-4 Turbo (batch) | openai/gpt-4-turbo:batch | 128K | text, image | text | Off | $5.00 | $15.0 | The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023. |
| GPT-3.5 Turbo (older v0613) | openai/gpt-3.5-turbo-0613 | 4K | text | text | Off | $1.00 | $2.00 | GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion... |
| GPT-3.5 Turbo 16k | openai/gpt-3.5-turbo-16k | 16K | text | text | Off | $3.00 | $4.00 | This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a highe... |
| GPT-3.5 Turbo | openai/gpt-3.5-turbo | 16K | text | text | Off | $0.5000 | $1.50 | GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion... |
| GPT-3.5 Turbo (batch) | openai/gpt-3.5-turbo:batch | 16K | text | text | Off | $0.2500 | $0.7500 | GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion... |
| GPT-4 | openai/gpt-4 | 8K | text | text | Off | $30.0 | $60.0 | OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous mo... |
| GPT-6 Astra Fast | openai/gpt-6-astra-fast | 1.1M | file, image, text | text | Always on | $20.0 | $100.0 | GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scienti... |
| GPT-6 Astra Flex | openai/gpt-6-astra-flex | 1.1M | file, image, text | text | Always on | $5.00 | $25.0 | GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scienti... |
| GPT-6 Astra Pro Fast | openai/gpt-6-astra-pro-fast | 1.1M | file, image, text | text | Always on | $20.0 | $100.0 | GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. ... |
| GPT-6 Astra Pro Flex | openai/gpt-6-astra-pro-flex | 1.1M | file, image, text | text | Always on | $5.00 | $25.0 | GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. ... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Qwen3.8 Omni Flash | qwen/qwen3.8-omni-flash | 1.0M | text, image, audio, video | text | On | $0.1500 | $0.4700 | Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video under... |
| Qwen3.8 Max (0902) | qwen/qwen3.8-max-0902 | 1.0M | text, image, video | text | Always on | $2.00 | $6.00 | Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts ... |
| Qwen3.8 Flash | qwen/qwen3.8-flash | 1.0M | text, image, video | text | On | $0.1500 | $0.4700 | Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and ... |
| Qwen3.8 27B | qwen/qwen3.8-27b | 1.0M | text, image, video | text | On | $0.0990 | $4.40 | Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction... |
| Qwen3.8 2.4T A95B | qwen/qwen3.8-2.4t-a95b | 1.0M | text | text | Always on | $2.00 | $6.00 | Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95... |
| Qwen3.7 Flash | qwen/qwen3.7-flash | 1.0M | text, image, video | text | On | $0.0300 | $0.1300 | Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, ... |
| Qwen3.7 Plus | qwen/qwen3.7-plus | 1.0M | text, image | text | On | $0.3200 | $1.28 | Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text ca... |
| Qwen3.7 Max | qwen/qwen3.7-max | 1.0M | text | text | On | $1.48 | $4.42 | Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with par... |
| Qwen3.5 Plus 2026-04-20 | qwen/qwen3.5-plus-20260420 | 1.0M | text, image, video | text | Off | $0.3000 | $1.80 | Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, w... |
| Qwen3.6 Flash | qwen/qwen3.6-flash | 1.0M | text, image, video | text | Off | $0.1875 | $1.12 | Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context win... |
| Qwen3.6 35B A3B | qwen/qwen3.6-35b-a3b | 262K | text, image, video | text | On | $0.0500 | $0.7000 | Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It u... |
| Qwen3.6 Max Preview | qwen/qwen3.6-max-preview | 262K | text | text | On | $1.03 | $6.16 | Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion... |
| Qwen3.6 27B | qwen/qwen3.6-27b | 262K | text, image, video | text | On | $0.3000 | $2.00 | Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabi... |
| Qwen3.6 Plus | qwen/qwen3.6-plus | 1.0M | text, image, video | text | Off | $0.3250 | $1.95 | Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalabi... |
| Qwen3.5-9B | qwen/qwen3.5-9b | 262K | text, image, video | text | Off | $0.0800 | $0.1300 | Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an effi... |
| Qwen3.5-35B-A3B | qwen/qwen3.5-35b-a3b | 262K | text, image, video | text | Off | $0.0800 | $0.7500 | The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a spa... |
| Qwen3.5-27B | qwen/qwen3.5-27b | 262K | text, image, video | text | Off | $0.1950 | $1.56 | The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference... |
| Qwen3.5-122B-A10B | qwen/qwen3.5-122b-a10b | 262K | text, image, video | text | Off | $0.2600 | $2.08 | The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixtur... |
| Qwen3.5-Flash | qwen/qwen3.5-flash-02-23 | 1.0M | text, image, video | text | Off | $0.0650 | $0.2600 | The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-... |
| Qwen3.5 Plus 2026-02-15 | qwen/qwen3.5-plus-02-15 | 1.0M | text, image, video | text | Off | $0.2600 | $1.56 | The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixtu... |
| Qwen3.5 397B A17B | qwen/qwen3.5-397b-a17b | 262K | text, image, video | text | Off | $0.3900 | $2.34 | The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse... |
| Qwen3 Max Thinking | qwen/qwen3-max-thinking | 262K | text | text | Off | $0.7800 | $3.90 | Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoni... |
| Qwen3 Coder Next | qwen/qwen3-coder-next | 262K | text | text | Off | $0.1200 | $0.8000 | Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with... |
| Qwen3 Embedding 8B | qwen/qwen3-embedding-8b | 33K | text | embeddings | Off | $0.0100 | FREE | The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. This ... |
| Qwen3 Embedding 4B | qwen/qwen3-embedding-4b | 33K | text | embeddings | Off | $0.0200 | FREE | The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. This ... |
| Qwen3 VL 32B Instruct | qwen/qwen3-vl-32b-instruct | 131K | text, image | text | Off | $0.1040 | $0.4160 | Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, a... |
| Qwen3 VL 8B Thinking | qwen/qwen3-vl-8b-thinking | 131K | image, text | text | Always on | $0.1800 | $2.10 | Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across... |
| Qwen3 VL 8B Instruct | qwen/qwen3-vl-8b-instruct | 262K | image, text | text | Off | $0.1170 | $0.4550 | Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, ... |
| Qwen3 VL 30B A3B Thinking | qwen/qwen3-vl-30b-a3b-thinking | 262K | text, image | text | Always on | $0.2000 | $2.40 | Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking vari... |
| Qwen3 VL 30B A3B Instruct | qwen/qwen3-vl-30b-a3b-instruct | 262K | text, image | text | Off | $0.1300 | $0.5200 | Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct vari... |
| Qwen3 VL 235B A22B Thinking | qwen/qwen3-vl-235b-a22b-thinking | 131K | text, image | text | Always on | $0.4000 | $4.00 | Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking ... |
| Qwen3 VL 235B A22B Instruct | qwen/qwen3-vl-235b-a22b-instruct | 262K | text, image | text | Off | $0.2000 | $0.8800 | Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. ... |
| Qwen3 Max | qwen/qwen3-max | 262K | text | text | Off | $0.7800 | $3.90 | Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and ... |
| Qwen3 Coder Plus | qwen/qwen3-coder-plus | 1.0M | text | text | Off | $0.6500 | $3.25 | Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autono... |
| Qwen3 Coder Flash | qwen/qwen3-coder-flash | 1.0M | text | text | Off | $0.1950 | $0.9750 | Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing... |
| Qwen3 Next 80B A3B Thinking | qwen/qwen3-next-80b-a3b-thinking | 262K | text | text | Always on | $0.1500 | $1.20 | Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed... |
| Qwen3 Next 80B A3B Instruct | qwen/qwen3-next-80b-a3b-instruct | 262K | text | text | Off | $0.0900 | $1.10 | Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces... |
| Qwen Plus 0728 | qwen/qwen-plus-2025-07-28 | 1.0M | text | text | Off | $0.2600 | $0.7800 | Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combin... |
| Qwen3 30B A3B Thinking 2507 | qwen/qwen3-30b-a3b-thinking-2507 | 82K | text | text | Always on | $0.2000 | $2.40 | Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. ... |
| Qwen3 Coder 30B A3B Instruct | qwen/qwen3-coder-30b-a3b-instruct | 262K | text | text | Off | $0.0700 | $0.2700 | Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced c... |
| Qwen3 30B A3B Instruct 2507 | qwen/qwen3-30b-a3b-instruct-2507 | 262K | text | text | Off | $0.0481 | $0.1930 | Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates i... |
| Qwen3 235B A22B Thinking 2507 | qwen/qwen3-235b-a22b-thinking-2507 | 131K | text | text | Always on | $0.2300 | $2.30 | Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It act... |
| Qwen3 Coder 480B A35B | qwen/qwen3-coder | 262K | text | text | Off | $0.2200 | $1.80 | Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding task... |
| Qwen3 235B A22B Instruct 2507 | qwen/qwen3-235b-a22b-2507 | 262K | text | text | Off | $0.0875 | $0.3500 | Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B ac... |
| Qwen3 30B A3B | qwen/qwen3-30b-a3b | 131K | text | text | On | $0.1200 | $0.5000 | Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reaso... |
| Qwen3 8B | qwen/qwen3-8b | 131K | text | text | On | $0.1170 | $0.4550 | Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It sup... |
| Qwen3 14B | qwen/qwen3-14b | 131K | text | text | Off | $0.1000 | $0.2200 | Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It suppo... |
| Qwen3 32B | qwen/qwen3-32b | 131K | text | text | Off | $0.0800 | $0.2800 | Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supp... |
| Qwen3 235B A22B | qwen/qwen3-235b-a22b | 131K | text | text | Off | $0.4550 | $1.82 | Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless... |
| Qwen-Plus | qwen/qwen-plus | 1.0M | text | text | Off | $0.2600 | $0.7800 | Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination. |
| Qwen2.5 7B Instruct | qwen/qwen-2.5-7b-instruct | 33K | text | text | Off | $0.1000 | $0.2000 | Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge an... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | google/gemini-3.8-flash | 1.0M | text, image, video, file, audio | text | Always on | $0.7500 | $3.75 | Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-... |
| Gemini 3.8 Flash (batch) | google/gemini-3.8-flash:batch | 1.0M | text, image, video, file, audio | text | Always on | $0.3750 | $1.88 | Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-... |
| Gemini 3.7 Flash | google/gemini-3.7-flash | 1.0M | text, image, video, file, audio | text | Always on | $0.7500 | $3.75 | Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that... |
| Gemini 3.7 Flash (batch) | google/gemini-3.7-flash:batch | 1.0M | text, image, video, file, audio | text | Always on | $0.3750 | $1.88 | Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that... |
| Gemini 3.6 Flash | google/gemini-3.6-flash | 1.0M | text, image, video, file, audio | text | Always on | $0.7500 | $3.75 | Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished... |
| Gemini 3.6 Flash (batch) | google/gemini-3.6-flash:batch | 1.0M | text, image, video, file, audio | text | Always on | $0.3750 | $1.88 | Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished... |
| Gemini 3.5 Flash Lite | google/gemini-3.5-flash-lite | 1.0M | text, image, video, file, audio | text | Always on | $0.3000 | $2.50 | Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks... |
| Gemini 3.5 Flash Lite (batch) | google/gemini-3.5-flash-lite:batch | 1.0M | text, image, video, file, audio | text | Always on | $0.1500 | $1.25 | Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks... |
| Nano Banana Pro (Gemini 3 Pro Image) | google/gemini-3-pro-image | 131K | image, text | image, text | Always on | $2.00 | $12.0 | Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with signific... |
| Gemini Embedding 2 | google/gemini-embedding-2 | 8K | text, image, file, audio, video | embeddings | Off | $0.2000 | FREE | Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic... |
| Gemini Embedding 2 (batch) | google/gemini-embedding-2:batch | 8K | text, image, file, audio, video | embeddings | Off | $0.1000 | FREE | Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic... |
| Gemini 3.5 Flash | google/gemini-3.5-flash | 1.0M | text, image, video, file, audio | text | Always on | $1.50 | $9.00 | Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly... |
| Gemini 3.5 Flash (batch) | google/gemini-3.5-flash:batch | 1.0M | text, image, video, file, audio | text | Always on | $0.7500 | $4.50 | Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly... |
| Gemini 3.1 Flash Lite | google/gemini-3.1-flash-lite | 1.0M | text, image, video, file, audio | text | On | $0.2500 | $1.50 | Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video... |
| Gemini 3.1 Flash Lite (batch) | google/gemini-3.1-flash-lite:batch | 1.0M | text, image, video, file, audio | text | On | $0.1250 | $0.7500 | Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video... |
| Gemini Pro Latest | ~google/gemini-pro-latest | 1.0M | audio, file, image, text, video | text | Always on | $2.00 | $12.0 | This model always redirects to the latest model in the Gemini Pro family. |
| Gemini Flash Latest | ~google/gemini-flash-latest | 1.0M | text, image, video, file, audio | text | Always on | $0.7500 | $3.75 | This model always redirects to the latest model in the Gemini Flash family. |
| Gemini Embedding 2 Preview | google/gemini-embedding-2-preview | 8K | text, image, file, audio, video | embeddings | Off | $0.2000 | FREE | Gemini Embedding 2 Preview is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for ... |
| Gemma 4 26B A4B | google/gemma-4-26b-a4b-it | 262K | image, text, video | text | Off | $0.0420 | $0.2200 | Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per... |
| Gemma 4 31B | google/gemma-4-31b-it | 262K | image, text, video | text | Off | $0.0900 | $0.3400 | Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context... |
| Gemini 3.1 Flash Lite Preview | google/gemini-3.1-flash-lite-preview | 1.0M | text, image, video, file, audio | text | On | $0.2500 | $1.50 | Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall q... |
| Gemini 3.1 Pro Preview Custom Tools | google/gemini-3.1-pro-preview-customtools | 1.0M | text, audio, image, video, file | text | Always on | $2.00 | $12.0 | Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool ... |
| Gemini 3.1 Pro Preview | google/gemini-3.1-pro-preview | 1.0M | audio, file, image, text, video | text | Always on | $2.00 | $12.0 | Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and m... |
| Gemini 3.1 Pro Preview (batch) | google/gemini-3.1-pro-preview:batch | 1.0M | audio, file, image, text, video | text | Always on | $1.00 | $6.00 | Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and m... |
| Gemini 3 Flash Preview | google/gemini-3-flash-preview | 1.0M | text, image, file, audio, video | text | Off | $0.5000 | $3.00 | Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers ... |
| Gemini 3 Flash Preview (batch) | google/gemini-3-flash-preview:batch | 1.0M | text, image, file, audio, video | text | Off | $0.2500 | $1.50 | Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers ... |
| Gemini Embedding 001 | google/gemini-embedding-001 | 20K | text | embeddings | Off | $0.1500 | FREE | gemini-embedding-001 provides a unified cutting edge experience across domains, including science, legal, finance, and coding. This embedding model ha... |
| Gemini 2.5 Flash Lite | google/gemini-2.5-flash-lite | 1.0M | text, image, file, audio, video | text | Off | $0.1000 | $0.4000 | Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improv... |
| Gemini 2.5 Flash Lite (batch) | google/gemini-2.5-flash-lite:batch | 1.0M | text, image, file, audio, video | text | Off | $0.0500 | $0.2000 | Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improv... |
| Gemini 2.5 Flash | google/gemini-2.5-flash | 1.0M | file, image, text, audio, video | text | Off | $0.3000 | $2.50 | Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks... |
| Gemini 2.5 Flash (batch) | google/gemini-2.5-flash:batch | 1.0M | file, image, text, audio, video | text | Off | $0.1500 | $1.25 | Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks... |
| Gemini 2.5 Pro | google/gemini-2.5-pro | 1.0M | text, image, file, audio, video | text | Always on | $1.25 | $10.0 | Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking”... |
| Gemini 2.5 Pro (batch) | google/gemini-2.5-pro:batch | 1.0M | text, image, file, audio, video | text | Always on | $0.6250 | $5.00 | Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking”... |
| Gemini 2.5 Pro Preview 06-05 | google/gemini-2.5-pro-preview | 1.0M | file, image, text, audio | text | Always on | $1.25 | $10.0 | Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking”... |
| Gemma 3 12B | google/gemma-3-12b-it | 131K | text, image | text | Off | $0.0500 | $0.1500 | Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 14... |
| Gemma 3 27B | google/gemma-3-27b-it | 131K | text, image | text | Off | $0.0800 | $0.1600 | Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 14... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | anthropic/claude-opus-5.5 | 1.0M | text, image, file | text | Always on | $4.00 | $20.0 | Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particul... |
| Claude Opus 5.5 (batch) | anthropic/claude-opus-5.5:batch | 1.0M | text, image, file | text | Always on | $2.00 | $10.0 | Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particul... |
| Claude Fable 5.1 | anthropic/claude-fable-5.1 | 1.0M | text, image, file | text | Always on | $10.0 | $50.0 | Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge ... |
| Claude Fable 5.1 (batch) | anthropic/claude-fable-5.1:batch | 1.0M | text, image, file | text | Always on | $5.00 | $25.0 | Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge ... |
| Claude Opus 5 | anthropic/claude-opus-5 | 1.0M | text, image, file | text | On | $5.00 | $25.0 | Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end so... |
| Claude Opus 5 (batch) | anthropic/claude-opus-5:batch | 1.0M | text, image, file | text | On | $2.50 | $12.5 | Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end so... |
| Claude Sonnet 5 (batch) | anthropic/claude-sonnet-5:batch | 1.0M | text, image, file | text | On | $1.00 | $5.00 | Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive ... |
| Claude Fable Latest | ~anthropic/claude-fable-latest | 1.0M | text, image, file | text | Always on | $10.0 | $50.0 | This model always redirects to the latest model in the Claude Fable family. |
| Claude Fable 5 (batch) | anthropic/claude-fable-5:batch | 1.0M | text, image, file | text | Always on | $5.00 | $25.0 | Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with t... |
| Claude Opus 4.8 (batch) | anthropic/claude-opus-4.8:batch | 1.0M | text, image, file | text | Off | $2.50 | $12.5 | Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, w... |
| Claude Haiku Latest | ~anthropic/claude-haiku-latest | 200K | text, image, file | text | Off | $1.00 | $5.00 | This model always redirects to the latest model in the Claude Haiku family. |
| Claude Sonnet Latest | ~anthropic/claude-sonnet-latest | 1.0M | text, image, file | text | On | $2.00 | $10.0 | This model always redirects to the latest model in the Claude Sonnet family. |
| Claude Opus Latest | ~anthropic/claude-opus-latest | 1.0M | text, image, file | text | Always on | $4.00 | $20.0 | This model always redirects to the latest model in the Claude Opus family. |
| Claude Opus 4.7 (batch) | anthropic/claude-opus-4.7:batch | 1.0M | text, image, file | text | Off | $2.50 | $12.5 | Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths ... |
| Claude Sonnet 4.6 (batch) | anthropic/claude-sonnet-4.6:batch | 1.0M | text, image, file | text | Off | $1.50 | $7.50 | Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at it... |
| Claude Opus 4.6 (batch) | anthropic/claude-opus-4.6:batch | 1.0M | text, image, file | text | Off | $2.50 | $12.5 | Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows ra... |
| Claude Opus 4.5 | anthropic/claude-opus-4.5 | 200K | file, image, text | text | Off | $5.00 | $25.0 | Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. ... |
| Claude Opus 4.5 (batch) | anthropic/claude-opus-4.5:batch | 200K | file, image, text | text | Off | $2.50 | $12.5 | Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. ... |
| Claude Haiku 4.5 (batch) | anthropic/claude-haiku-4.5:batch | 200K | text, image, file | text | Off | $0.5000 | $2.50 | Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of large... |
| Claude Sonnet 4.5 | anthropic/claude-sonnet-4.5 | 1.0M | text, image, file | text | Off | $3.00 | $15.0 | Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-ar... |
| Claude Sonnet 4.5 (batch) | anthropic/claude-sonnet-4.5:batch | 1.0M | text, image, file | text | Off | $1.50 | $7.50 | Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-ar... |
| Claude Opus 4.1 | anthropic/claude-opus-4.1 | 200K | image, text, file | text | Off | $15.0 | $75.0 | Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieve... |
| Claude Opus 4.1 (batch) | anthropic/claude-opus-4.1:batch | 200K | image, text, file | text | Off | $7.50 | $37.5 | Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieve... |
| Claude Sonnet 4 | anthropic/claude-sonnet-4 | 200K | image, text, file | text | Off | $3.00 | $15.0 | Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved pre... |
| Claude 3 Haiku | anthropic/claude-3-haiku | 200K | text, image | text | Off | $0.2500 | $1.25 | Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch ... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Mistral Medium 3.5 | mistralai/mistral-medium-3-5 | 262K | text, image, file | text | Off | $1.50 | $7.50 | Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed fo... |
| Mistral Medium 3.5 (batch) | mistralai/mistral-medium-3-5:batch | 262K | text, image, file | text | Off | $0.7500 | $3.75 | Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed fo... |
| Mistral Small 4 | mistralai/mistral-small-2603 | 262K | text, image | text | Off | $0.1500 | $0.6000 | Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single syst... |
| Mistral Small 4 (batch) | mistralai/mistral-small-2603:batch | 262K | text, image | text | Off | $0.0750 | $0.3000 | Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single syst... |
| Devstral 2 2512 | mistralai/devstral-2512 | 262K | text, file | text | Off | $0.4000 | $2.00 | Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model suppor... |
| Ministral 3 14B 2512 | mistralai/ministral-14b-2512 | 262K | text, image | text | Off | $0.2000 | $0.2000 | The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 2... |
| Ministral 3 8B 2512 | mistralai/ministral-8b-2512 | 262K | text, image | text | Off | $0.1500 | $0.1500 | A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities. |
| Ministral 3 8B 2512 (batch) | mistralai/ministral-8b-2512:batch | 262K | text, image | text | Off | $0.0750 | $0.0750 | A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities. |
| Ministral 3 3B 2512 | mistralai/ministral-3b-2512 | 131K | text, image | text | Off | $0.1000 | $0.1000 | The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities. |
| Mistral Large 3 2512 (batch) | mistralai/mistral-large-2512:batch | 262K | text, image, file | text | Off | $0.2500 | $0.7500 | Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B tota... |
| Mistral Embed 2312 | mistralai/mistral-embed-2312 | 8K | text | embeddings | Off | $0.1000 | FREE | Mistral Embed is a specialized embedding model for text data, optimized for semantic search and RAG applications. Developed by Mistral AI in late 2023... |
| Codestral Embed 2505 | mistralai/codestral-embed-2505 | 8K | text | embeddings | Off | $0.1500 | FREE | Mistral Codestral Embed is specially designed for code, perfect for embedding code databases, repositories, and powering coding assistants with state-... |
| Voxtral Small 24B 2507 | mistralai/voxtral-small-24b-2507 | 33K | text, audio, file | text | Off | $0.1000 | $0.3000 | Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text perform... |
| Mistral Medium 3.1 | mistralai/mistral-medium-3.1 | 131K | text, image, file | text | Off | $0.4000 | $2.00 | Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier... |
| Mistral Medium 3.1 (batch) | mistralai/mistral-medium-3.1:batch | 131K | text, image, file | text | Off | $0.2000 | $1.00 | Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier... |
| Codestral 2508 | mistralai/codestral-2508 | 256K | text, file | text | Off | $0.3000 | $0.9000 | Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in... |
| Codestral 2508 (batch) | mistralai/codestral-2508:batch | 256K | text, file | text | Off | $0.1500 | $0.4500 | Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in... |
| Mistral Small 3.2 24B | mistralai/mistral-small-3.2-24b-instruct | 256K | image, text | text | Off | $0.0750 | $0.2000 | Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and impr... |
| Mistral Medium 3 | mistralai/mistral-medium-3 | 131K | text, image, file | text | Off | $0.4000 | $2.00 | Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operat... |
| Mistral Small 3.1 24B | mistralai/mistral-small-3.1-24b-instruct | 128K | text, image | text | Off | $0.3510 | $0.5550 | Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities... |
| Saba | mistralai/mistral-saba | 33K | text, file | text | Off | $0.2000 | $0.6000 | Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually relevant... |
| Mistral Nemo | mistralai/mistral-nemo | 131K | text | text | Off | $0.0180 | $0.0300 | A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, F... |
| Mixtral 8x22B Instruct | mistralai/mixtral-8x22b-instruct | 66K | text, file | text | Off | $2.00 | $6.00 | Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141B, offerin... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| GLM 5.3 FlashX | z-ai/glm-5.3-flashx | 1.0M | text, image, video | text | Always on | $0.3700 | $1.25 | GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built o... |
| GLM Flash Latest | ~z-ai/glm-flash-latest | 1.3M | text, image, video | text | Always on | $0.0750 | $0.2500 | This model always redirects to the latest model in the GLM Flash family. |
| GLM 5.3 Flash | z-ai/glm-5.3-flash | 1.3M | text, image, video | text | Always on | $0.1200 | $0.4000 | GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear att... |
| GLM 5.3 Flash (batch) | z-ai/glm-5.3-flash:batch | 1.0M | text, image, video | text | Always on | $0.0600 | $0.2000 | GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear att... |
| GLM Latest | ~z-ai/glm-latest | 1.3M | text | text | Always on | $0.5625 | $2.50 | This model always redirects to the latest GLM model from Z.ai. |
| GLM 5.3 (batch) | z-ai/glm-5.3:batch | 1.0M | text | text | Always on | $0.7200 | $2.40 | GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and ou... |
| GLM 5.1 | z-ai/glm-5.1 | 205K | text | text | On | $0.9660 | $3.04 | GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built a... |
| GLM 5V Turbo | z-ai/glm-5v-turbo | 203K | image, text, video | text | On | $1.20 | $4.00 | GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image... |
| GLM 5 Turbo | z-ai/glm-5-turbo | 203K | text | text | On | $1.20 | $4.00 | GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is... |
| GLM 5 | z-ai/glm-5 | 205K | text | text | On | $0.6000 | $1.92 | GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert develop... |
| GLM 4.7 Flash | z-ai/glm-4.7-flash | 200K | text | text | On | $0.0600 | $0.4000 | As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use ... |
| GLM 4.7 | z-ai/glm-4.7 | 205K | text | text | On | $0.4000 | $1.75 | GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/e... |
| GLM 4.6V | z-ai/glm-4.6v | 131K | image, text, video | text | Off | $0.3000 | $0.9000 | GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed me... |
| GLM 4.6 | z-ai/glm-4.6 | 205K | text | text | Off | $0.4300 | $1.75 | Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K ... |
| GLM 4.5V | z-ai/glm-4.5v | 66K | text, image | text | Off | $0.6000 | $1.80 | GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameter... |
| GLM 4.5 | z-ai/glm-4.5 | 131K | text | text | Off | $0.6000 | $2.20 | GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and ... |
| GLM 4.5 Air | z-ai/glm-4.5-air | 131K | text | text | Off | $0.1300 | $0.8500 | GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts... |
| GLM 5.2 | z-ai/glm-5.2 | 1.0M | text | text | Off | $0.9000 | $2.83 | GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon a... |
| GLM 5.3 Flash | z-ai/glm-5.3-flash:US | 1.0M | text, image | text | Always on | $0.2475 | $0.8250 | GLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 series (320B total / 18B active). This `:US` id is the Fireworks-hosted US SKU; `... |
| GLM 5.3 | z-ai/glm-5.3:US | 1.0M | text | text | Always on | $2.31 | $7.26 | GLM-5.3 is Z.ai's flagship text model (743B MoE). This `:US` id is the Fireworks-hosted US SKU; `z-ai/glm-5.3` is the Impala-served global id. Text-on... |
| GLM 5.3 | z-ai/glm-5.3 | 1.0M | text | text | Always on | $0.9100 | $2.86 | GLM-5.3 is Z.ai's flagship text model (743B MoE). Served by Impala. The `:US` id remains the Fireworks-hosted US SKU. Text-only — use `z-ai/glm-5.3-fl... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| DeepSeek Pro Latest | ~deepseek/deepseek-pro-latest | 1.0M | text | text | Off | $0.4000 | $4.30 | This model always redirects to the latest model in the DeepSeek Pro family. |
| DeepSeek Flash Latest | ~deepseek/deepseek-flash-latest | 1.0M | text, image | text | On | $0.1000 | $0.5000 | This model always redirects to the latest model in the DeepSeek Flash family. |
| DeepSeek V4.1 Flash | deepseek/deepseek-v4.1-flash | 1.0M | text, image | text | On | $0.1000 | $0.5000 | DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture... |
| DeepSeek V4.1 Flash (batch) | deepseek/deepseek-v4.1-flash:batch | 1.0M | text, image | text | On | $0.1120 | $0.3360 | DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture... |
| DeepSeek V4 Flash Vision Exp | deepseek/deepseek-v4-flash-vision-exp | 1.0M | text, image | text | On | $0.2156 | $0.6468 | DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while match... |
| DeepSeek V4 Pro 0813 | deepseek/deepseek-v4-pro-0813 | 1.0M | text | text | Off | $0.5800 | $1.74 | DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro. |
| DeepSeek V4 Flash Latest | ~deepseek/deepseek-v4-flash-latest | 1.3M | text | text | On | $0.0380 | $0.5500 | This model always redirects to the latest model in the DeepSeek V4 Flash family. |
| DeepSeek V4 Pro 0423 | deepseek/deepseek-v4-pro | 1.0M | text | text | Off | $0.9400 | $1.89 | DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token... |
| DeepSeek V4 Flash 0423 | deepseek/deepseek-v4-flash | 1.0M | text | text | Off | $0.0500 | $0.1400 | DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supportin... |
| DeepSeek V3.2 | deepseek/deepseek-v3.2 | 164K | text | text | Off | $0.2088 | $0.3096 | DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It... |
| DeepSeek V3.2 Exp | deepseek/deepseek-v3.2-exp | 164K | text | text | Off | $0.2700 | $0.4100 | DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It intro... |
| DeepSeek V3.1 Terminus | deepseek/deepseek-v3.1-terminus | 164K | text | text | Off | $0.2700 | $1.00 | DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing... |
| DeepSeek V3.1 | deepseek/deepseek-chat-v3.1 | 164K | text | text | Off | $0.2500 | $0.9500 | DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates.... |
| R1 0528 | deepseek/deepseek-r1-0528 | 164K | text | text | Always on | $0.5000 | $2.15 | May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully ... |
| DeepSeek V3 0324 | deepseek/deepseek-chat-v3-0324 | 164K | text | text | Off | $0.2400 | $0.9000 | DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds... |
| R1 | deepseek/deepseek-r1 | 64K | text | text | Always on | $0.7000 | $2.50 | DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in s... |
| DeepSeek V3 | deepseek/deepseek-chat | 164K | text | text | Off | $0.2574 | $1.03 | DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-tra... |
| DeepSeek V4 Flash | deepseek/deepseek-v4-flash-0731:US | 1.0M | text | text | On | $0.3630 | $1.09 | DeepSeek V4 Flash is a hybrid-reasoning mixture-of-experts model tuned for high-throughput agentic work. This `:US` id is the Fireworks-hosted US SKU;... |
| DeepSeek V4 Flash | deepseek/deepseek-v4-flash-0731 | 1.0M | text | text | On | $0.0700 | $0.1400 | DeepSeek V4 Flash is a hybrid-reasoning mixture-of-experts model tuned for high-throughput agentic work. Thinking is on by default and can be disabled... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Kimi K3 | moonshotai/kimi-k3 | 1.0M | text, image, video | text | On | $1.50 | $10.8 | Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon... |
| Kimi K3 (batch) | moonshotai/kimi-k3:batch | 1.0M | text, image, video | text | On | $2.28 | $11.4 | Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon... |
| Kimi K2.7 Code | moonshotai/kimi-k2.7-code | 262K | text, image | text | Always on | $0.6800 | $3.40 | MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over lon... |
| Kimi Latest | ~moonshotai/kimi-latest | 1.0M | text, image, video | text | On | $1.50 | $10.8 | This model always redirects to the latest model in the Kimi family. |
| Kimi K2.6 | moonshotai/kimi-k2.6 | 262K | text, image | text | On | $0.4972 | $2.97 | Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchest... |
| Kimi K2.5 | moonshotai/kimi-k2.5 | 262K | text, image | text | On | $0.4500 | $2.25 | Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Bui... |
| Kimi K2 Thinking | moonshotai/kimi-k2-thinking | 262K | text | text | Always on | $0.6000 | $2.50 | Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on t... |
| Kimi K2 0905 | moonshotai/kimi-k2-0905 | 262K | text | text | Off | $0.6000 | $2.50 | Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by M... |
| Kimi K2 0711 | moonshotai/kimi-k2 | 131K | text | text | Off | $0.5700 | $2.30 | Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 bill... |
| Kimi K3 | moonshotai/kimi-k3:US | 1.0M | text, image | text | Always on | $3.63 | $18.2 | Kimi K3 is Moonshot's sparse mixture-of-experts model with a 1M token context window. This `:US` id is the Fireworks-hosted US SKU; `moonshotai/kimi-k... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Muse Spark 1.3 | meta/muse-spark-1.3 | 1.0M | text, image, video, file, audio | text | Always on | $1.25 | $4.25 | Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of ... |
| Muse Glimmer 30B | meta/muse-glimmer-30b | 131K | text, image | text | Always on | $0.3000 | $1.10 | Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous age... |
| Muse Spark 1.2 | meta/muse-spark-1.2 | 1.0M | text, image, video, file, audio | text | Always on | $1.25 | $4.25 | Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns t... |
| Llama 4 Maverick | meta-llama/llama-4-maverick | 1.0M | text, image | text | Off | $0.1875 | $0.6525 | Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128... |
| Llama 4 Scout | meta-llama/llama-4-scout | 1.3M | text, image | text | Off | $0.1000 | $0.3000 | Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 10... |
| Llama 3.3 70B Instruct | meta-llama/llama-3.3-70b-instruct | 131K | text | text | Off | $0.1000 | $0.3200 | The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama... |
| Llama 3.1 70B Instruct | meta-llama/llama-3.1-70b-instruct | 131K | text | text | Off | $0.4000 | $0.4000 | Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dia... |
| Llama 3.1 8B Instruct | meta-llama/llama-3.1-8b-instruct | 131K | text | text | Off | $0.0200 | $0.0400 | Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demo... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Seed 2.1 Turbo | bytedance-seed/seed-2-1-turbo | 262K | text, image, video | text | Off | $0.5000 | $2.50 | Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, m... |
| Seed-2.0-Code | bytedance-seed/seed-2.0-code | 262K | text, image, video | text | Off | $0.5000 | $3.00 | Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and ... |
| Seed-2.0-Lite | bytedance-seed/seed-2.0-lite | 262K | text, image, video | text | Off | $0.2500 | $2.00 | Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably low... |
| Seed-2.0-Mini | bytedance-seed/seed-2.0-mini | 262K | text, image, video | text | Off | $0.1000 | $0.4000 | Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. ... |
| Seed 1.6 Flash | bytedance-seed/seed-1.6-flash | 262K | image, text, video | text | Off | $0.0750 | $0.3000 | Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k co... |
| Seed 1.6 | bytedance-seed/seed-1.6 | 262K | image, text, video | text | Off | $0.2500 | $2.00 | Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| MiniMax M3 | minimax/minimax-m3 | 1.0M | text, image, video | text | Off | $0.2300 | $0.9600 | MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and i... |
| MiniMax M2.7 | minimax/minimax-m2.7 | 205K | text | text | Always on | $0.2100 | $0.8400 | MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively ... |
| MiniMax M2.5 | minimax/minimax-m2.5 | 205K | text | text | Always on | $0.2700 | $0.9500 | MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working env... |
| MiniMax M2.1 | minimax/minimax-m2.1 | 205K | text | text | Always on | $0.3000 | $1.20 | MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With... |
| MiniMax M2 | minimax/minimax-m2 | 205K | text | text | Always on | $0.2550 | $1.02 | MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated paramet... |
| MiniMax M1 | minimax/minimax-m1 | 1.0M | text | text | Off | $0.4000 | $2.20 | MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Nova 2 Lite | amazon/nova-2-lite-v1 | 1.0M | text, image, video, file | text | Off | $0.3000 | $2.50 | Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite d... |
| Nova Premier 1.0 | amazon/nova-premier-v1 | 1.0M | text, image | text | Off | $2.50 | $12.5 | Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custo... |
| Nova Lite 1.0 | amazon/nova-lite-v1 | 300K | text, image | text | Off | $0.0600 | $0.2400 | Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text... |
| Nova Micro 1.0 | amazon/nova-micro-v1 | 128K | text | text | Off | $0.0350 | $0.1400 | Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a c... |
| Nova Pro 1.0 | amazon/nova-pro-v1 | 300K | text, image | text | Off | $0.8000 | $3.20 | Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of task... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| paraphrase-MiniLM-L6-v2 | sentence-transformers/paraphrase-minilm-l6-v2 | 512 | text | embeddings | Off | $0.005000 | FREE | The paraphrase-MiniLM-L6-v2 embedding model converts sentences and short paragraphs into a 384-dimensional dense vector space, producing high-quality ... |
| all-MiniLM-L12-v2 | sentence-transformers/all-minilm-l12-v2 | 512 | text | embeddings | Off | $0.005000 | FREE | The all-MiniLM-L12-v2 embedding model maps sentences and short paragraphs into a 384-dimensional dense vector space, producing efficient and high-qual... |
| multi-qa-mpnet-base-dot-v1 | sentence-transformers/multi-qa-mpnet-base-dot-v1 | 512 | text | embeddings | Off | $0.005000 | FREE | The multi-qa-mpnet-base-dot-v1 embedding model transforms sentences and short paragraphs into a 768-dimensional dense vector space, generating high-qu... |
| all-mpnet-base-v2 | sentence-transformers/all-mpnet-base-v2 | 512 | text | embeddings | Off | $0.005000 | FREE | The all-mpnet-base-v2 embedding model encodes sentences and short paragraphs into a 768-dimensional dense vector space, providing high-fidelity semant... |
| all-MiniLM-L6-v2 | sentence-transformers/all-minilm-l6-v2 | 512 | text | embeddings | Off | $0.005000 | FREE | The all-MiniLM-L6-v2 embedding model maps sentences and short paragraphs into a 384-dimensional dense vector space, enabling high-quality semantic rep... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| voyage-code-4 | voyageai/voyage-code-4 | 32K | text | embeddings | Off | $0.1200 | FREE | voyage-code-4 is a code embedding model from Voyage AI, a MongoDB company. It is designed for coding agents and code retrieval, with Matryoshka embedd... |
| voyage-multimodal-3.5 | voyageai/voyage-multimodal-3.5 | 32K | text, image | embeddings | Off | $0.1200 | FREE | voyage-multimodal-3.5 is a state-of-the-art multimodal embedding model capable of vectorizing not only text, images, and video individually, but also ... |
| voyage-4-lite | voyageai/voyage-4-lite | 32K | text | embeddings | Off | $0.0200 | FREE | voyage-4-lite is a lightweight, general-purpose embedding model optimized for low latency and cost. Enabled by Matryoshka learning and quantization-aw... |
| voyage-4 | voyageai/voyage-4 | 32K | text | embeddings | Off | $0.0600 | FREE | voyage-4 is a general-purpose (including multilingual) embedding model optimized for retrieval/search and AI applications. voyage-4 supports embedding... |
| voyage-4-large | voyageai/voyage-4-large | 32K | text | embeddings | Off | $0.1200 | FREE | voyage-4-large is a state-of-the-art general-purpose and multilingual embedding optimized for retrieval quality. Enabled by Matryoshka learning and qu... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| MiMo-V2.6-Pro-UltraSpeed | xiaomi/mimo-v2.6-pro-ultraspeed | 1.0M | text, image, video, audio | text | Off | $4.35 | $8.70 | MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoi... |
| MiMo-V2.6-Flash | xiaomi/mimo-v2.6-flash | 1.0M | text, image, video, audio | text | Off | $0.1400 | $0.2800 | MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B ... |
| MiMo-V2.6-Pro | xiaomi/mimo-v2.6-pro | 1.0M | text, image, video, audio | text | Off | $0.4350 | $0.8700 | MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capa... |
| MiMo-V2.5-Pro | xiaomi/mimo-v2.5-pro | 1.1M | text | text | Off | $0.3045 | $0.6090 | MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizo... |
| MiMo-V2.5 | xiaomi/mimo-v2.5 | 1.1M | text, audio, image, video | text | Off | $0.1190 | $0.2380 | MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Ling 3.0 Flash VL | inclusionai/ling-3.0-flash-vl | 131K | text, image, video | text | On | $0.0600 | $0.1800 | Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while addi... |
| Ling 3.0 Flash Sante (free) | inclusionai/ling-3.0-flash-sante:free | 262K | text | text | On | FREE | FREE | Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters o... |
| Ling 3.0 Flash Fin | inclusionai/ling-3.0-flash-fin | 262K | text | text | On | $0.0600 | $0.1800 | Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B tot... |
| Ling 3.0 Flash Fin | inclusionai/ling-3.0-flash-fin:free | 262K | text | text | On | FREE | FREE | Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B tot... |
| Ling 3.0 Flash | inclusionai/ling-3.0-flash | 262K | text | text | On | $0.0210 | $0.0630 | *Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Nemotron 3.5 Lightning | nvidia/nemotron-3.5-lightning | 262K | text | text | Off | $0.0650 | $0.1800 | NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throu... |
| Nemotron 3 Ultra | nvidia/nemotron-3-ultra-550b-a55b | 262K | text | text | On | $0.5000 | $2.20 | NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built o... |
| Nemotron 3 Super | nvidia/nemotron-3-super-120b-a12b | 262K | text | text | On | $0.0800 | $0.4500 | NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in compl... |
| Nemotron 3 Nano 30B A3B | nvidia/nemotron-3-nano-30b-a3b | 262K | text | text | Off | $0.0500 | $0.2000 | NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic ... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Laguna S 2.1 | poolside/laguna-s-2.1 | 1.0M | text | text | On | $0.0900 | $0.1800 | Laguna S 2.1 is the latest coding agent model from [Poolside](). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2%... |
| Laguna S 2.1 | poolside/laguna-s-2.1:free | 262K | text | text | On | FREE | FREE | Laguna S 2.1 is the latest coding agent model from [Poolside](). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2%... |
| Laguna XS 2.1 | poolside/laguna-xs-2.1 | 262K | text | text | On | $0.0600 | $0.1200 | Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in Apri... |
| Laguna XS 2.1 | poolside/laguna-xs-2.1:free | 262K | text | text | On | FREE | FREE | Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in Apri... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Grok Build 0.1 | x-ai/grok-build-0.1 | 256K | text, image, file | text | Always on | $1.00 | $2.00 | Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with... |
| Grok 4.3 | x-ai/grok-4.3 | 1.0M | text, image, file | text | On | $1.25 | $2.50 | Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-follo... |
| Grok 4.3 (batch) | x-ai/grok-4.3:batch | 1.0M | text, image, file | text | On | $1.00 | $2.00 | Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-follo... |
| Grok 4.20 | x-ai/grok-4.20 | 2.0M | text, image, file | text | Off | $1.25 | $2.50 | Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination r... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Solar Mini 4 | upstage/solar-mini4 | 524K | text | text | Off | $0.0500 | $0.2000 | Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context wind... |
| Solar Pro 4 | upstage/solar-pro4 | 524K | text | text | Off | $0.0900 | $0.3600 | Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflow... |
| Solar Pro 3 | upstage/solar-pro-3 | 131K | text | text | Off | $0.1500 | $0.6000 | Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it d... |
| Solar Pro 4 | upstage/solar-pro4:free | 524K | text | text | Off | FREE | FREE | Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflow... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Grok Latest | ~x-ai/grok-latest | 500K | text, image, file | text | Always on | $1.60 | $4.80 | This model always redirects to the latest Grok model from xAI. |
| Grok 4.5 | x-ai/grok-4.5 | 500K | text, image | text | Off | $2.00 | $6.00 | xAI's previous-generation frontier model for coding, knowledge work, and STEM, with a 500K context window. Reasoning is always on and its depth is con... |
| Grok 4.6 | x-ai/grok-4.6 | 500K | text, image | text | Off | $2.00 | $6.00 | xAI's frontier model for coding, agentic tasks, and knowledge work, with a 500K context window. Reasoning is always on and its depth is controllable (... |
| Grok 4.7 | x-ai/grok-4.7 | 500K | text, image | text | Off | $1.00 | $3.00 | xAI's frontier model for coding, agentic tasks, and knowledge work, with a 500K context window. Reasoning is always on and its depth is controllable (... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Aion-3.0-Mini | aion-labs/aion-3.0-mini | 131K | text | text | Always on | $0.7000 | $1.40 | Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative gene... |
| Aion-3.0 | aion-labs/aion-3.0 | 131K | text | text | Always on | $3.00 | $6.00 | Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation pro... |
| Aion-2.0 | aion-labs/aion-2.0 | 131K | text | text | Always on | $0.8000 | $1.60 | Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises,... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| bge-base-en-v1.5 | baai/bge-base-en-v1.5 | 512 | text | embeddings | Off | $0.005000 | FREE | The bge-base-en-v1.5 embedding model converts English sentences and paragraphs into 768-dimensional dense vectors, delivering efficient, high-quality ... |
| bge-large-en-v1.5 | baai/bge-large-en-v1.5 | 512 | text | embeddings | Off | $0.0100 | FREE | The bge-large-en-v1.5 embedding model maps English sentences, paragraphs, and documents into a 1024-dimensional dense vector space, delivering high-fi... |
| bge-m3 | baai/bge-m3 | 8K | text | embeddings | Off | $0.0100 | FREE | The bge-m3 embedding model encodes sentences, paragraphs, and long documents into a 1024-dimensional dense vector space, delivering high-quality seman... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Command A+ | cohere/command-a-plus | 192K | text, image | text | Off | $0.3000 | $1.50 | Command A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports native t... |
| Command R (08-2024) | cohere/command-r-08-2024 | 128K | text | text | Off | $0.1500 | $0.6000 | command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented generation ... |
| Command R+ (08-2024) | cohere/command-r-plus-08-2024 | 128K | text | text | Off | $2.50 | $10.0 | command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower latencies as c... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| E5-Large-v2 | intfloat/e5-large-v2 | 512 | text | embeddings | Off | $0.0100 | FREE | The e5-large-v2 embedding model maps English sentences, paragraphs, and documents into a 1024-dimensional dense vector space, delivering high-accuracy... |
| E5-Base-v2 | intfloat/e5-base-v2 | 512 | text | embeddings | Off | $0.005000 | FREE | The e5-base-v2 embedding model encodes English sentences and paragraphs into a 768-dimensional dense vector space, producing efficient and high-qualit... |
| Multilingual-E5-Large | intfloat/multilingual-e5-large | 512 | text | embeddings | Off | $0.0100 | FREE | The multilingual-e5-large embedding model encodes sentences, paragraphs, and documents across over 90 languages into a 1024-dimensional dense vector s... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Fugu Ultra v2 | sakana/fugu-ultra-v2 | 1.0M | text, image, file | text | Always on | $5.00 | $30.0 | Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchest... |
| Fugu Max | sakana/fugu-max | 1.0M | text, image, file | text | Always on | $2.00 | $6.00 | Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration ... |
| Fugu Ultra | sakana/fugu-ultra | 1.0M | text, image | text | Always on | $5.00 | $30.0 | Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestrat... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Step 3.7 Flash | stepfun/step-3.7-flash | 262K | text, image, video | text | Always on | $0.1600 | $0.9200 | Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision enco... |
| Step 3.5 Flash | stepfun/step-3.5-flash | 262K | text | text | Always on | $0.1000 | $0.3000 | Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activat... |
| Step 3.7 Flash | stepfun/step-3.7-flash:free | 262K | text, image, video | text | Always on | FREE | FREE | Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision enco... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Hy4 preview | tencent/hy4-preview | 1.0M | text | text | On | $0.4170 | $1.25 | Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, compl... |
| Hy3 | tencent/hy3 | 262K | text | text | On | $0.1300 | $0.5300 | Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and... |
| Hy3 preview | tencent/hy3-preview | 262K | text | text | On | $0.1800 | $0.6000 | Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable rea... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | inception/mercury-2.5 | 260K | text | text | On | $0.0400 | $0.1500 | Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 p... |
| Mercury 2 | inception/mercury-2 | 128K | text | text | On | $0.2500 | $0.7500 | Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produ... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| LongCat 2.0 | meituan/longcat-2.0 | 1.0M | text | text | On | $0.3000 | $1.20 | LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, reposit... |
| LongCat 2.0 | meituan/longcat-2.0:free | 1.0M | text | text | Off | FREE | FREE | LongCat 2.0 is a sparse mixture-of-experts model from Meituan with 48B active parameters out of 1.6T total, and a 1M token context window. It targets ... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Embed V1 4B | perplexity/pplx-embed-v1-4b | 32K | text | embeddings | Off | $0.0300 | FREE | pplx-embed-v1 -4B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized ... |
| Embed V1 0.6B | perplexity/pplx-embed-v1-0.6b | 32K | text | embeddings | Off | $0.004000 | FREE | pplx-embed-v1-0.6B is one of Perplexity's state-of-the-art text embedding models built for real-world, web-scale retrieval. pplx-embed-v1 is optimized... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| GTE-Base | thenlper/gte-base | 512 | text | embeddings | Off | $0.005000 | FREE | The gte-base embedding model encodes English sentences and paragraphs into a 768-dimensional dense vector space, delivering efficient and effective se... |
| GTE-Large | thenlper/gte-large | 512 | text | embeddings | Off | $0.0100 | FREE | The gte-large embedding model converts English sentences, paragraphs and moderate-length documents into a 1024-dimensional dense vector space, deliver... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Inkling Small | thinkingmachines/inkling-small | 1.0M | text, image, audio | text | On | $0.4500 | $1.20 | Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is po... |
| Inkling | thinkingmachines/inkling | 1.0M | text, image, audio | text | On | $0.9500 | $4.05 | Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Trinity Large Thinking | arcee-ai/trinity-large-thinking | 262K | text | text | Always on | $0.2500 | $0.8000 | Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloa... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | anthropic/claude-fable-5 | 1.0M | text | text | Off | $10.0 | $50.0 |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | anthropic/claude-haiku-4.5 | 410K | text | text | Off | $1.00 | $5.00 |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.6 | anthropic/claude-opus-4.6 | 1.0M | text | text | Off | $5.00 | $25.0 |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.7 | anthropic/claude-opus-4.7 | 1.0M | text | text | Off | $5.00 | $25.0 |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.8 | anthropic/claude-opus-4.8 | 1.0M | text | text | Off | $5.00 | $25.0 |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 4.6 | anthropic/claude-sonnet-4.6 | 1.0M | text | text | Off | $3.00 | $15.0 |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | anthropic/claude-sonnet-5 | 1.0M | text | text | Off | $2.00 | $10.0 |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Granite 4.2 8B | ibm-granite/granite-4.2-8b | 131K | text | text | On | $0.0600 | $0.2500 | Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that n... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| KAT-Coder-Pro V2.5 | kwaipilot/kat-coder-pro-v2.5 | 262K | text | text | Off | $0.7400 | $2.96 | KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Mistral Large | mistralai/mistral-large | 128K | text, file | text | Off | $2.00 | $6.00 | This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasonin... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Mistral Large 2407 | mistralai/mistral-large-2407 | 131K | text, file | text | Off | $2.00 | $6.00 | This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning,... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Nex-N2.5-Pro | nex-agi/nex-n2.5-pro | 262K | text, image | text | Off | $0.0750 | $0.2500 | Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: i... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Pareto | unbiased/pareto | 262K | text, image | text | Off | $2.50 | $7.50 | Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad r... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Ternary Bonsai 2 27B | prism-ml/ternary-bonsai-2-27b | 262K | text, image | text | On | $0.0750 | $0.5000 | Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image unders... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Qwen2.5 72B Instruct | qwen/qwen-2.5-72b-instruct | 33K | text | text | Off | $0.3600 | $0.4000 | Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge a... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Reka Edge | rekaai/reka-edge | 16K | image, text, video | text | Off | $0.1000 | $0.1000 | Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Relace Search | relace/relace-search | 256K | text | text | Off | $1.00 | $3.00 | The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In con... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| Llama 3.1 Euryale 70B v2.2 | sao10k/l3.1-euryale-70b | 131K | text | text | Off | $0.8500 | $0.8500 | Euryale L3.1 70B v2.2 is a model focused on creative roleplay from Sao10k. It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b)... |
| Model | ID | Context | Input | Output | Reasoning | Prompt/1M | Completion/1M | Description |
|---|---|---|---|---|---|---|---|---|
| meta/muse-spark-1.1 | meta/muse-spark-1.1 | 1.0M | text | text | Off | $1.25 | $4.25 |