Gemini-3-Pro-Preview
GeminiGemini

Gemini-3-Pro-Preview

Google’s latest multimodal AI model that combines advanced reasoning, coding, and image-video understanding with a massive 1-million-token context window

Manta-Flash-1.0
MegaNovaMegaNova

Manta-Flash-1.0

Balanced tier for general-purpose AI agent, good tradeoff between cost and latency. (Mid-tier)

Web Search
MegaNovaMegaNova

Web Search

Real-time web search with intelligent multi-provider fallback and LLM-based quality gating. 50 free queries per user per day, $0.002 per overage query.

DeepSeek-V3-0324 (free)
DeepseekDeepseek

DeepSeek-V3-0324 (free)

The most powerful AI-driven LLM with 685B parameters released by Deepseek. Community-shared access, daily limits, great for testing and exploration

L3.3-MS-Nevoria-70b
AI Agent ModelsAI Agent Models

L3.3-MS-Nevoria-70b

70B narrative generation, LoRA tuning & story-rich outputs

MN-Violet-Lotus-12B
AI Agent ModelsAI Agent Models

MN-Violet-Lotus-12B

12B parameter model for creative, emotionally intelligent conversations

Search
Try:
Recently Added
In $0.09/1M
Out $0.18/1M
DeepSeek-V4-Flash-0731DeepseekDeepseek

DeepSeek-V4-Flash-0731

DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding performance, at 284B total / 13B active parameters and a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5-Enterprise

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

Featured
Free
Gemini-3.6-FlashGoogleGoogle

Gemini-3.6-Flash

Available for Creator Tier+

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
In $1.50/1M
Out $7.50/1M
Gemini-3.6-Flash-EnterpriseGoogleGoogle

Gemini-3.6-Flash-Enterprise

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
Free
Gemini-3.5-Flash-LiteGoogleGoogle

Gemini-3.5-Flash-Lite

Available for Creator Tier+

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

Featured
In $0.30/1M
Out $2.50/1M
Gemini-3.5-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Lite-Enterprise

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

In $1.50/1M
Out $5.00/1M
Qwen3.8-Max-PreviewQwenQwen

Qwen3.8-Max-Preview

Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data analysis, and office workflows; always uses thinking mode and supports long-horizon tool workflows and structured output

In $3.00/1M
Out $15.00/1M
Kimi-K3Moonshot AIMoonshot AI

Kimi-K3

Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning

$7.00/1M
Seedance-2.0BytedanceBytedance

Seedance-2.0

BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video

$5.60/1M
Seedance-2.0-FastBytedanceBytedance

Seedance-2.0-Fast

BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens

$3.50/1M
Seedance-2.0-MiniBytedanceBytedance

Seedance-2.0-Mini

BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens

In $5.00/1M
Out $30.00/1M
GPT-5.6-SolOpen AIOpen AI

GPT-5.6-Sol

OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding

In $2.50/1M
Out $15.00/1M
GPT-5.6-TerraOpen AIOpen AI

GPT-5.6-Terra

OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers

In $2.00/1M
Out $6.00/1M
xAI

Grok-4.5

xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input

In $1.60/1M
Out $8.00/1M
Anthropic

Claude-Sonnet-5

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $1.00/1M
Out $3.00/1M
GLM-5.2Z.aiZ.ai

GLM-5.2

Z.ai's open-source MoE model for long-context reasoning and agentic software engineering, with a 1M-token context window

In $0.75/1M
Out $3.50/1M
Kimi-K2.7-CodeMoonshot AIMoonshot AI

Kimi-K2.7-Code

Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering

In $10.00/1M
Out $50.00/1M
Anthropic

Claude-Fable-5

Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M3

A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling

In $0.12/1M
Out $0.76/1M
Qwen3.6-35B-A3BQwenQwen

Qwen3.6-35B-A3B

Qwen3.6 35B (3B-active) MoE with built-in NEXTN speculative decoding, reasoning and tool calling — fast low-cost generation

In $4.25/1M
Out $21.25/1M
Anthropic

Claude-Opus-4.8

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

Featured
In $1.20/1M
Out $7.20/1M
Gemini-3.5-FlashGoogleGoogle

Gemini-3.5-Flash

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-LiteGoogleGoogle

Gemini-3.1-Flash-Lite

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

In $1.25/1M
Out $2.50/1M
xAI

Grok-4.3

xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input

In $4.00/1M
Out $24.00/1M
GPT-5.5Open AIOpen AI

GPT-5.5

OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks

In $1.30/1M
Out $2.60/1M
DeepSeek-V4-ProDeepseekDeepseek

DeepSeek-V4-Pro

DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic performance with a 1M context window

In $0.10/1M
Out $0.20/1M
DeepSeek-V4-FlashDeepseekDeepseek

DeepSeek-V4-Flash

DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M context window

In $1.00/1M
Out $3.00/1M
Xiaomi

MiMo-V2.5-Pro

Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon tasks with a 1M context window

In $0.40/1M
Out $2.00/1M
Xiaomi

MiMo-V2.5

Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.7

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $1.05/1M
Out $3.50/1M
GLM-5.1Z.aiZ.ai

GLM-5.1

Z.ai's open-source flagship MoE model built for agentic engineering, delivering state-of-the-art coding and long-horizon task performance that improves the longer it runs

In $0.50/1M
Out $3.00/1M
Qwen3.6-PlusQwenQwen

Qwen3.6-Plus

Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks

In $0.13/1M
Out $0.38/1M
Gemma-4-31BGoogleGoogle

Gemma-4-31B

Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows

In $0.15/1M
Out $0.60/1M
Mistral-Small-4MistralMistral

Mistral-Small-4

A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models

Featured
In $0.20/1M
Out $1.20/1M
Gemini-3.1-Flash-Lite-PreviewGoogleGoogle

Gemini-3.1-Flash-Lite-Preview

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $1.60/1M
Out $9.60/1M
Token-length tiered pricing
Gemini-3.1-Pro-PreviewGoogleGoogle

Gemini-3.1-Pro-Preview

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M2.7

A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity tasks like coding, tool use, and multi-step reasoning — at a highly competitive price point

Free
GLM-4.7-FlashZ.aiZ.ai

GLM-4.7-Flash

Available for Creator Tier+

A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks

In $0.40/1M
In $0.28/1M
Qwen3.5-PlusQwenQwen

Qwen3.5-Plus

Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing

In $0.80/1M
Out $2.56/1M
GLM-5Z.aiZ.ai

GLM-5

Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding, and agentic tasks

In $2.40/1M
Out $12.00/1M
Anthropic

Claude-Sonnet-4.6

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.6

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.5

An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads

In $0.80/1M
Out $4.00/1M
Anthropic

Claude-Haiku-4.5

A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications

In $2.40/1M
Out $12.00/1M
Anthropic

Claude-Sonnet-4.5

A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks

In $12.00/1M
Out $60.00/1M
Anthropic

Claude-Opus-4

Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks

In $0.28/1M
Out $1.20/1M
MiniMax

MiniMax-M2.1

A lightweight, open-source text-to-text large language model optimized for coding, agent-oriented workflows, and instruction following, delivering faster, cleaner outputs and strong reasoning and problem-solving performance compared to its predecessor

In $0.09/1M
Out $0.10/1M
Qwen3-235B-A22B-Instruct-2507QwenQwen

Qwen3-235B-A22B-Instruct-2507

A 235B-parameter MoE instruction-tuned language model for general, multilingual, and coding tasks

In $0.45/1M
Out $1.90/1M
GLM-4.6Z.aiZ.ai

GLM-4.6

Flagship model from Z.ai with 200K token context, superior reasoning, advanced tool use, and coding excellence

In $2.00/1M
Out $12.00/1M
GPT-5.4Open AIOpen AI

GPT-5.4

A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding, and tool-use capabilities with very long context support

In $1.40/1M
Out $11.20/1M
GPT-5.3-ChatOpen AIOpen AI

GPT-5.3-Chat

A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality dialogue, instruction following, and everyday chat applications

In $1.40/1M
Out $11.20/1M
GPT-5.3-CodexOpen AIOpen AI

GPT-5.3-Codex

A high-performance agentic coding model designed to autonomously write, debug, and manage complex software development tasks while collaborating interactively with developers

In $0.30/1M
Out $0.50/1M
AI Agent ModelsAI Agent Models

Cydonia-24B-v4.3

A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence

In $0.30/1M
Out $0.50/1M
AI Agent ModelsAI Agent Models

Cydonia-24B-v4.1

A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence

In $0.85/1M
Out $0.85/1M
AI Agent ModelsAI Agent Models

L3.3-70B-Euryale-v2.3

70B model focused on creative roleplay from Sao10k

In $0.85/1M
Out $0.85/1M
AI Agent ModelsAI Agent Models

L3.1-70B-Euryale-v2.2

70B model focused on creative roleplay from Sao10k

Free
AI Agent ModelsAI Agent Models

Sapphira-L3.3-70B-0.1

70B parameter language model optimized for storytelling, dialogue, and immersive role-play

Free
AI Agent ModelsAI Agent Models

MN-Violet-Lotus-12B

12B parameter model for creative, emotionally intelligent conversations

Free
AI Agent ModelsAI Agent Models

L3.3-MS-Nevoria-70B

70B narrative generation, LoRA tuning & story-rich outputs

In $0.02/1M
Out $0.04/1M
Mistral-Nemo-Instruct-2407MistralMistral

Mistral-Nemo-Instruct-2407

24B instruction model, long context, fewer errors

$48/M
Nano-Banana-2GoogleGoogle

Nano-Banana-2

Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.

$0.015/image
Alibaba-Z-Image-TurboZ-ImageZ-Image

Alibaba-Z-Image-Turbo

Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.

$96/M
Nano-Banana-Pro-EditGoogleGoogle

Nano-Banana-Pro-Edit

Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).

$24/M
Nano-Banana-EditGoogleGoogle

Nano-Banana-Edit

AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.

$0.035/image
Bytedance-Seedream-5.0-LiteBytedanceBytedance

Bytedance-Seedream-5.0-Lite

Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.

$0.10/second
Alibaba-Wan-2.6-T2VWanWan

Alibaba-Wan-2.6-T2V

Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.10/second
Alibaba-Wan-2.6-I2VWanWan

Alibaba-Wan-2.6-I2V

Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.025/second
Alibaba-Wan2.6-I2V-FlashWanWan

Alibaba-Wan2.6-I2V-Flash

Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing

$0.10/second
Wan2.6-R2VWanWan

Wan2.6-R2V

Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)

$0.025/second
Alibaba-Wan-2.6-R2V-FlashWanWan

Alibaba-Wan-2.6-R2V-Flash

Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.

$0.06/second
Veo-3.1-FastGoogleGoogle

Veo-3.1-Fast

Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p. Optional AI-generated audio (+50% cost).

$0.06/second
Veo-3.1-Fast-I2VGoogleGoogle

Veo-3.1-Fast-I2V

Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at 720p/1080p.

$0.12/second
Veo-3.1GoogleGoogle

Veo-3.1

Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds at 720p/1080p.

$0.12/second
Veo-3.1-I2VGoogleGoogle

Veo-3.1-I2V

Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.

In $0.20/1M
Out $1.60/1M
Context-length tiered pricing
Qwen3-VL-PlusQwenQwen

Qwen3-VL-Plus

Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual reasoning and agent capabilities

In $0.05/1M
Out $0.40/1M
Context-length tiered pricing
Qwen3-VL-FlashQwenQwen

Qwen3-VL-Flash

Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on text, image, and video tasks with integrated thinking and non-thinking modes

50 / day free
then $0.002/query
Web SearchMegaNovaMegaNova

Web Search

Real-time managed natural-language web search that returns ranked results, optionally enriched with full extracted page text in a single call—includes 50 free queries per user per day, with additional queries priced at $0.002 each

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-5-Enterprise

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.7-Enterprise

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.6-Enterprise

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.5-Enterprise

An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads

In $15.00/1M
Out $75.00/1M
Anthropic

Claude-Opus-4-Enterprise

Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-4.6-Enterprise

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-4.5-Enterprise

A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks

In $1.00/1M
Out $5.00/1M
Anthropic

Claude-Haiku-4.5-Enterprise

A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications

Featured
In $1.50/1M
Out $9.00/1M
Gemini-3.5-Flash-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Enterprise

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Enterprise

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Preview-Enterprise

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $2.00/1M
Out $12.00/1M
Token-length tiered pricing
Gemini-3.1-Pro-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Pro-Preview-Enterprise

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

Featured
In $0.75/1M
Out $12.00/1M
Gemini-3.1-Flash-LiveGoogleGoogle

Gemini-3.1-Flash-Live

Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket

Best AI Agent Models
Featured
Free
Manta-Mini-1.0MegaNovaMegaNova

Manta-Mini-1.0

Lightweight tier optimized for speed and cost. (Free)

Featured
Free
Manta-Flash-1.0MegaNovaMegaNova

Manta-Flash-1.0

Balanced tier for general-purpose roleplay, good tradeoff between cost and latency. (Mid-tier)

Featured
Free
Manta-Pro-1.0MegaNovaMegaNova

Manta-Pro-1.0

Available for Creator Tier+

Premium tier with advanced reasoning and long context, best for complex roleplay.(Top-tier)

In $0.26/1M
Out $0.38/1M
DeepSeek-V3.2DeepseekDeepseek

DeepSeek-V3.2

A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and general chat capabilities with efficient inference

In $0.30/1M
Out $0.50/1M
AI Agent ModelsAI Agent Models

Cydonia-24B-v4.3

A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence

In $0.30/1M
Out $0.50/1M
AI Agent ModelsAI Agent Models

Cydonia-24B-v4.1

A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence

In $0.85/1M
Out $0.85/1M
AI Agent ModelsAI Agent Models

L3.3-70B-Euryale-v2.3

70B model focused on creative roleplay from Sao10k

In $0.85/1M
Out $0.85/1M
AI Agent ModelsAI Agent Models

L3.1-70B-Euryale-v2.2

70B model focused on creative roleplay from Sao10k

Free
AI Agent ModelsAI Agent Models

Sapphira-L3.3-70B-0.1

70B parameter language model optimized for storytelling, dialogue, and immersive role-play

Free
AI Agent ModelsAI Agent Models

MN-Violet-Lotus-12B

12B parameter model for creative, emotionally intelligent conversations

Free
AI Agent ModelsAI Agent Models

L3.3-MS-Nevoria-70B

70B narrative generation, LoRA tuning & story-rich outputs

Free
AI Agent ModelsAI Agent Models

L3-70B-Euryale-v2.1

70B tuned for storytelling, roleplay, creative writing

Free
AI Agent ModelsAI Agent Models

L3-8B-Stheno-v3.2

8B LLaMA-3 model for roleplay & assistant tasks

In $0.26/1M
Out $0.38/1M
DeepSeek-V3.2-ExpDeepseekDeepseek

DeepSeek-V3.2-Exp

Experimental version with sparse-attention architecture for long-context efficiency

In $0.20/1M
Out $0.77/1M
DeepSeek-V3-0324DeepseekDeepseek

DeepSeek-V3-0324

High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation

In $0.19/1M
Out $0.79/1M
DeepSeek-V3.1DeepseekDeepseek

DeepSeek-V3.1

Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent

In $0.50/1M
Out $2.15/1M
DeepSeek-R1-0528DeepseekDeepseek

DeepSeek-R1-0528

An open-source next-generation reasoning-optimized language model with enhanced logic, math, and code performance, larger context handling, and improved function-calling capabilities

Text Generation
Featured
Free
Manta-Mini-1.0MegaNovaMegaNova

Manta-Mini-1.0

Lightweight tier optimized for speed and cost. (Free)

Featured
Free
Manta-Flash-1.0MegaNovaMegaNova

Manta-Flash-1.0

Balanced tier for general-purpose roleplay, good tradeoff between cost and latency. (Mid-tier)

Featured
Free
Manta-Pro-1.0MegaNovaMegaNova

Manta-Pro-1.0

Available for Creator Tier+

Premium tier with advanced reasoning and long context, best for complex roleplay.(Top-tier)

In $0.09/1M
Out $0.18/1M
DeepSeek-V4-Flash-0731DeepseekDeepseek

DeepSeek-V4-Flash-0731

DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding performance, at 284B total / 13B active parameters and a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5-Enterprise

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

Featured
Free
Gemini-3.6-FlashGoogleGoogle

Gemini-3.6-Flash

Available for Creator Tier+

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
In $1.50/1M
Out $7.50/1M
Gemini-3.6-Flash-EnterpriseGoogleGoogle

Gemini-3.6-Flash-Enterprise

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
Free
Gemini-3.5-Flash-LiteGoogleGoogle

Gemini-3.5-Flash-Lite

Available for Creator Tier+

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

Featured
In $0.30/1M
Out $2.50/1M
Gemini-3.5-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Lite-Enterprise

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

In $1.50/1M
Out $5.00/1M
Qwen3.8-Max-PreviewQwenQwen

Qwen3.8-Max-Preview

Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data analysis, and office workflows; always uses thinking mode and supports long-horizon tool workflows and structured output

In $3.00/1M
Out $15.00/1M
Kimi-K3Moonshot AIMoonshot AI

Kimi-K3

Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning

In $5.00/1M
Out $30.00/1M
GPT-5.6-SolOpen AIOpen AI

GPT-5.6-Sol

OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding

In $2.50/1M
Out $15.00/1M
GPT-5.6-TerraOpen AIOpen AI

GPT-5.6-Terra

OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers

In $2.00/1M
Out $6.00/1M
xAI

Grok-4.5

xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input

In $1.60/1M
Out $8.00/1M
Anthropic

Claude-Sonnet-5

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $1.00/1M
Out $3.00/1M
GLM-5.2Z.aiZ.ai

GLM-5.2

Z.ai's open-source MoE model for long-context reasoning and agentic software engineering, with a 1M-token context window

In $0.75/1M
Out $3.50/1M
Kimi-K2.7-CodeMoonshot AIMoonshot AI

Kimi-K2.7-Code

Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering

In $10.00/1M
Out $50.00/1M
Anthropic

Claude-Fable-5

Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M3

A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling

In $0.12/1M
Out $0.76/1M
Qwen3.6-35B-A3BQwenQwen

Qwen3.6-35B-A3B

Qwen3.6 35B (3B-active) MoE with built-in NEXTN speculative decoding, reasoning and tool calling — fast low-cost generation

In $4.25/1M
Out $21.25/1M
Anthropic

Claude-Opus-4.8

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

Featured
In $1.20/1M
Out $7.20/1M
Gemini-3.5-FlashGoogleGoogle

Gemini-3.5-Flash

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-LiteGoogleGoogle

Gemini-3.1-Flash-Lite

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

In $1.25/1M
Out $2.50/1M
xAI

Grok-4.3

xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input

In $4.00/1M
Out $24.00/1M
GPT-5.5Open AIOpen AI

GPT-5.5

OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks

In $1.30/1M
Out $2.60/1M
DeepSeek-V4-ProDeepseekDeepseek

DeepSeek-V4-Pro

DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic performance with a 1M context window

In $0.10/1M
Out $0.20/1M
DeepSeek-V4-FlashDeepseekDeepseek

DeepSeek-V4-Flash

DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M context window

In $1.00/1M
Out $3.00/1M
Xiaomi

MiMo-V2.5-Pro

Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon tasks with a 1M context window

In $0.40/1M
Out $2.00/1M
Xiaomi

MiMo-V2.5

Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning

In $0.75/1M
Out $3.50/1M
Kimi-K2.6Moonshot AIMoonshot AI

Kimi-K2.6

An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context window, native vision input, and scaling up to 300 parallel sub-agents

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.7

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $1.05/1M
Out $3.50/1M
GLM-5.1Z.aiZ.ai

GLM-5.1

Z.ai's open-source flagship MoE model built for agentic engineering, delivering state-of-the-art coding and long-horizon task performance that improves the longer it runs

In $0.50/1M
Out $3.00/1M
Qwen3.6-PlusQwenQwen

Qwen3.6-Plus

Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks

In $0.13/1M
Out $0.38/1M
Gemma-4-31BGoogleGoogle

Gemma-4-31B

Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows

In $0.40/1M
Out $2.00/1M
AI Agent ModelsAI Agent Models

MiMo-V2-Omni

Xiaomi's omni-modal agent model that natively understands text, images, video, and audio in a unified architecture, designed for complex real-world multimodal tasks

In $0.15/1M
Out $0.60/1M
Mistral-Small-4MistralMistral

Mistral-Small-4

A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models

Featured
In $0.20/1M
Out $1.20/1M
Gemini-3.1-Flash-Lite-PreviewGoogleGoogle

Gemini-3.1-Flash-Lite-Preview

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $1.60/1M
Out $9.60/1M
Token-length tiered pricing
Gemini-3.1-Pro-PreviewGoogleGoogle

Gemini-3.1-Pro-Preview

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

In $1.00/1M
Out $3.00/1M
Xiaomi

MiMo-V2-Pro

A trillion-parameter text-only flagship agent model from Xiaomi, built for complex coding, long-horizon planning, and multi-step tool use with a 1M context window

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M2.7

A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity tasks like coding, tool use, and multi-step reasoning — at a highly competitive price point

Free
GLM-4.7-FlashZ.aiZ.ai

GLM-4.7-Flash

Available for Creator Tier+

A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks

In $0.40/1M
In $0.28/1M
Qwen3.5-PlusQwenQwen

Qwen3.5-Plus

Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M2.5

MiniMax-M2.5 is a cost-efficient frontier model that achieves SOTA in coding, agentic tool use, and search tasks through extensive reinforcement learning in real-world environments

In $0.80/1M
Out $2.56/1M
GLM-5Z.aiZ.ai

GLM-5

Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding, and agentic tasks

In $0.45/1M
Out $2.25/1M
Kimi-K2.5Moonshot AIMoonshot AI

Kimi-K2.5

An open-source native multimodal AI model that combines vision and language understanding with advanced agent-oriented capabilities, including self-directed agent swarm and visual coding, for versatile reasoning and tool use

In $2.40/1M
Out $12.00/1M
Anthropic

Claude-Sonnet-4.6

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.6

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.5

An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads

In $0.80/1M
Out $4.00/1M
Anthropic

Claude-Haiku-4.5

A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications

In $2.40/1M
Out $12.00/1M
Anthropic

Claude-Sonnet-4.5

A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks

In $12.00/1M
Out $60.00/1M
Anthropic

Claude-Opus-4

Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks

In $0.26/1M
Out $0.38/1M
DeepSeek-V3.2DeepseekDeepseek

DeepSeek-V3.2

A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and general chat capabilities with efficient inference

In $0.28/1M
Out $1.20/1M
MiniMax

MiniMax-M2.1

A lightweight, open-source text-to-text large language model optimized for coding, agent-oriented workflows, and instruction following, delivering faster, cleaner outputs and strong reasoning and problem-solving performance compared to its predecessor

In $0.20/1M
Out $0.80/1M
GLM-4.7Z.aiZ.ai

GLM-4.7

An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents

In $0.20/1M
Out $0.80/1M
GLM-4.7-EnterpriseZ.aiZ.ai

GLM-4.7-Enterprise

An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents

In $0.09/1M
Out $0.10/1M
Qwen3-235B-A22B-Instruct-2507QwenQwen

Qwen3-235B-A22B-Instruct-2507

A 235B-parameter MoE instruction-tuned language model for general, multilingual, and coding tasks

In $0.60/1M
Out $2.60/1M
Kimi-K2-ThinkingMoonshot AIMoonshot AI

Kimi-K2-Thinking

1T-parameter reasoning-focused language model designed for complex, multi-step problem solving and tool use

In $0.45/1M
Out $1.90/1M
GLM-4.6Z.aiZ.ai

GLM-4.6

Flagship model from Z.ai with 200K token context, superior reasoning, advanced tool use, and coding excellence

Featured
In $0.40/1M
Out $2.40/1M
Gemini-3-Flash-PreviewGoogleGoogle

Gemini-3-Flash-Preview

Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed

Featured
In $1.00/1M
Out $8.00/1M
Token-length tiered pricing
Gemini-2.5-ProGoogleGoogle

Gemini-2.5-Pro

Strongest Gemini, 1M context, great for code & knowledge

Featured
In $0.24/1M
Out $2.00/1M
Gemini-2.5-FlashGoogleGoogle

Gemini-2.5-Flash

Fast reasoning with improved latency & accuracy

Featured
In $0.08/1M
Out $0.32/1M
Gemini-2.5-Flash-LiteGoogleGoogle

Gemini-2.5-Flash-Lite

Ultra-low latency Gemini variant

In $2.00/1M
Out $12.00/1M
GPT-5.4Open AIOpen AI

GPT-5.4

A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding, and tool-use capabilities with very long context support

In $1.40/1M
Out $11.20/1M
GPT-5.3-ChatOpen AIOpen AI

GPT-5.3-Chat

A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality dialogue, instruction following, and everyday chat applications

In $1.40/1M
Out $11.20/1M
GPT-5.3-CodexOpen AIOpen AI

GPT-5.3-Codex

A high-performance agentic coding model designed to autonomously write, debug, and manage complex software development tasks while collaborating interactively with developers

In $1.40/1M
Out $11.20/1M
GPT-5.2Open AIOpen AI

GPT-5.2

OpenAI’s flagship multimodal (text + image) GPT-5.2 model for top-tier coding and agentic tasks, producing text output with the highest reasoning capability in the GPT-5 family

In $1.00/1M
Out $8.00/1M
GPT-5.1Open AIOpen AI

GPT-5.1

OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs

In $1.00/1M
Out $8.00/1M
GPT-5Open AIOpen AI

GPT-5

OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs

In $0.20/1M
Out $1.60/1M
GPT-5-MiniOpen AIOpen AI

GPT-5-Mini

OpenAI’s faster, more cost-efficient GPT-5 model for well-defined tasks, supporting text and image input with text output

In $0.04/1M
Out $0.32/1M
GPT-5-NanoOpen AIOpen AI

GPT-5-Nano

OpenAI’s fastest, most cost-efficient GPT-5 model for lightweight tasks like summarization and classification, supporting text and image input with text output

In $0.12/1M
Out $0.48/1M
GPT-4o-miniOpen AIOpen AI

GPT-4o-mini

Compact GPT-4 Omni, supports text & image inputs

In $0.30/1M
Out $0.50/1M
AI Agent ModelsAI Agent Models

Cydonia-24B-v4.3

A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence

In $0.30/1M
Out $0.50/1M
AI Agent ModelsAI Agent Models

Cydonia-24B-v4.1

A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence

In $0.85/1M
Out $0.85/1M
AI Agent ModelsAI Agent Models

L3.3-70B-Euryale-v2.3

70B model focused on creative roleplay from Sao10k

In $0.85/1M
Out $0.85/1M
AI Agent ModelsAI Agent Models

L3.1-70B-Euryale-v2.2

70B model focused on creative roleplay from Sao10k

Free
AI Agent ModelsAI Agent Models

Sapphira-L3.3-70B-0.1

70B parameter language model optimized for storytelling, dialogue, and immersive role-play

Free
AI Agent ModelsAI Agent Models

MN-Violet-Lotus-12B

12B parameter model for creative, emotionally intelligent conversations

Free
AI Agent ModelsAI Agent Models

L3.3-MS-Nevoria-70B

70B narrative generation, LoRA tuning & story-rich outputs

In $0.02/1M
Out $0.04/1M
Mistral-Nemo-Instruct-2407MistralMistral

Mistral-Nemo-Instruct-2407

24B instruction model, long context, fewer errors

Free
Mistral-Small-3.2-24B-Instruct-2506MistralMistral

Mistral-Small-3.2-24B-Instruct-2506

24B instruction model, long context, fewer errors

Free
AI Agent ModelsAI Agent Models

L3-70B-Euryale-v2.1

70B tuned for storytelling, roleplay, creative writing

Free
AI Agent ModelsAI Agent Models

L3-8B-Stheno-v3.2

8B LLaMA-3 model for roleplay & assistant tasks

In $0.26/1M
Out $0.38/1M
DeepSeek-V3.2-ExpDeepseekDeepseek

DeepSeek-V3.2-Exp

Experimental version with sparse-attention architecture for long-context efficiency

In $0.20/1M
Out $0.77/1M
DeepSeek-V3-0324DeepseekDeepseek

DeepSeek-V3-0324

High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation

In $0.20/1M
Out $0.77/1M
DeepSeek-V3-0324-EnterpriseDeepseekDeepseek

DeepSeek-V3-0324-Enterprise

High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation

In $0.19/1M
Out $0.79/1M
DeepSeek-V3.1DeepseekDeepseek

DeepSeek-V3.1

Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent

In $0.17/1M
Out $0.54/1M
DeepSeek-V3.1-EnterpriseDeepseekDeepseek

DeepSeek-V3.1-Enterprise

Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent

In $0.50/1M
Out $2.15/1M
DeepSeek-R1-0528DeepseekDeepseek

DeepSeek-R1-0528

An open-source next-generation reasoning-optimized language model with enhanced logic, math, and code performance, larger context handling, and improved function-calling capabilities

In $0.10/1M
Out $0.30/1M
Llama3.3-70BLlamaLlama

Llama3.3-70B

Multilingual 70B delivering 405B-level performance

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-5-Enterprise

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.7-Enterprise

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.6-Enterprise

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.5-Enterprise

An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads

In $15.00/1M
Out $75.00/1M
Anthropic

Claude-Opus-4-Enterprise

Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-4.6-Enterprise

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-4.5-Enterprise

A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks

In $1.00/1M
Out $5.00/1M
Anthropic

Claude-Haiku-4.5-Enterprise

A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications

Featured
In $1.50/1M
Out $9.00/1M
Gemini-3.5-Flash-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Enterprise

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Enterprise

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Preview-Enterprise

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $2.00/1M
Out $12.00/1M
Token-length tiered pricing
Gemini-3.1-Pro-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Pro-Preview-Enterprise

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

Featured
In $0.50/1M
Out $3.00/1M
Gemini-3-Flash-Preview-EnterpriseGoogleGoogle

Gemini-3-Flash-Preview-Enterprise

Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed

Featured
In $1.25/1M
Out $10.00/1M
Token-length tiered pricing
Gemini-2.5-Pro-EnterpriseGoogleGoogle

Gemini-2.5-Pro-Enterprise

Strongest Gemini, 1M context, great for code & knowledge

Featured
In $0.30/1M
Out $2.50/1M
Gemini-2.5-Flash-EnterpriseGoogleGoogle

Gemini-2.5-Flash-Enterprise

Fast reasoning with improved latency & accuracy

Featured
In $0.10/1M
Out $0.40/1M
Gemini-2.5-Flash-Lite-EnterpriseGoogleGoogle

Gemini-2.5-Flash-Lite-Enterprise

Ultra-low latency Gemini variant

Featured
In $0.09/1M
Out $0.18/1M
DeepSeek-V4-Flash-0731DeepseekDeepseek

DeepSeek-V4-Flash-0731

DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding performance, at 284B total / 13B active parameters and a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

Featured
Free
Gemini-3.6-FlashGoogleGoogle

Gemini-3.6-Flash

Available for Creator Tier+

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
In $1.50/1M
Out $7.50/1M
Gemini-3.6-Flash-EnterpriseGoogleGoogle

Gemini-3.6-Flash-Enterprise

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
Free
Gemini-3.5-Flash-LiteGoogleGoogle

Gemini-3.5-Flash-Lite

Available for Creator Tier+

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

Featured
In $0.30/1M
Out $2.50/1M
Gemini-3.5-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Lite-Enterprise

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

In $1.50/1M
Out $5.00/1M
Qwen3.8-Max-PreviewQwenQwen

Qwen3.8-Max-Preview

Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data analysis, and office workflows; always uses thinking mode and supports long-horizon tool workflows and structured output

In $3.00/1M
Out $15.00/1M
Kimi-K3Moonshot AIMoonshot AI

Kimi-K3

Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning

$7.00/1M
Seedance-2.0BytedanceBytedance

Seedance-2.0

BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video

$5.60/1M
Seedance-2.0-FastBytedanceBytedance

Seedance-2.0-Fast

BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens

$3.50/1M
Seedance-2.0-MiniBytedanceBytedance

Seedance-2.0-Mini

BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens

In $5.00/1M
Out $30.00/1M
GPT-5.6-SolOpen AIOpen AI

GPT-5.6-Sol

OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding

In $2.50/1M
Out $15.00/1M
GPT-5.6-TerraOpen AIOpen AI

GPT-5.6-Terra

OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers

In $2.00/1M
Out $6.00/1M
xAI

Grok-4.5

xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input

In $1.60/1M
Out $8.00/1M
Anthropic

Claude-Sonnet-5

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $1.00/1M
Out $3.00/1M
GLM-5.2Z.aiZ.ai

GLM-5.2

Z.ai's open-source MoE model for long-context reasoning and agentic software engineering, with a 1M-token context window

In $0.75/1M
Out $3.50/1M
Kimi-K2.7-CodeMoonshot AIMoonshot AI

Kimi-K2.7-Code

Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering

In $10.00/1M
Out $50.00/1M
Anthropic

Claude-Fable-5

Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M3

A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling

In $4.25/1M
Out $21.25/1M
Anthropic

Claude-Opus-4.8

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

Featured
In $1.20/1M
Out $7.20/1M
Gemini-3.5-FlashGoogleGoogle

Gemini-3.5-Flash

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-LiteGoogleGoogle

Gemini-3.1-Flash-Lite

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

In $1.25/1M
Out $2.50/1M
xAI

Grok-4.3

xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input

In $4.00/1M
Out $24.00/1M
GPT-5.5Open AIOpen AI

GPT-5.5

OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks

In $1.30/1M
Out $2.60/1M
DeepSeek-V4-ProDeepseekDeepseek

DeepSeek-V4-Pro

DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic performance with a 1M context window

In $0.10/1M
Out $0.20/1M
DeepSeek-V4-FlashDeepseekDeepseek

DeepSeek-V4-Flash

DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M context window

In $1.00/1M
Out $3.00/1M
Xiaomi

MiMo-V2.5-Pro

Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon tasks with a 1M context window

In $0.40/1M
Out $2.00/1M
Xiaomi

MiMo-V2.5

Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning

In $0.75/1M
Out $3.50/1M
Kimi-K2.6Moonshot AIMoonshot AI

Kimi-K2.6

An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context window, native vision input, and scaling up to 300 parallel sub-agents

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.7

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $1.05/1M
Out $3.50/1M
GLM-5.1Z.aiZ.ai

GLM-5.1

Z.ai's open-source flagship MoE model built for agentic engineering, delivering state-of-the-art coding and long-horizon task performance that improves the longer it runs

In $0.50/1M
Out $3.00/1M
Qwen3.6-PlusQwenQwen

Qwen3.6-Plus

Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks

In $0.13/1M
Out $0.38/1M
Gemma-4-31BGoogleGoogle

Gemma-4-31B

Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows

In $0.15/1M
Out $0.60/1M
Mistral-Small-4MistralMistral

Mistral-Small-4

A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models

Featured
In $0.20/1M
Out $1.20/1M
Gemini-3.1-Flash-Lite-PreviewGoogleGoogle

Gemini-3.1-Flash-Lite-Preview

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $1.60/1M
Out $9.60/1M
Token-length tiered pricing
Gemini-3.1-Pro-PreviewGoogleGoogle

Gemini-3.1-Pro-Preview

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M2.7

A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity tasks like coding, tool use, and multi-step reasoning — at a highly competitive price point

Free
GLM-4.7-FlashZ.aiZ.ai

GLM-4.7-Flash

Available for Creator Tier+

A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks

In $0.40/1M
In $0.28/1M
Qwen3.5-PlusQwenQwen

Qwen3.5-Plus

Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M2.5

MiniMax-M2.5 is a cost-efficient frontier model that achieves SOTA in coding, agentic tool use, and search tasks through extensive reinforcement learning in real-world environments

In $0.80/1M
Out $2.56/1M
GLM-5Z.aiZ.ai

GLM-5

Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding, and agentic tasks

In $0.45/1M
Out $2.25/1M
Kimi-K2.5Moonshot AIMoonshot AI

Kimi-K2.5

An open-source native multimodal AI model that combines vision and language understanding with advanced agent-oriented capabilities, including self-directed agent swarm and visual coding, for versatile reasoning and tool use

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.6

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $0.26/1M
Out $0.38/1M
DeepSeek-V3.2DeepseekDeepseek

DeepSeek-V3.2

A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and general chat capabilities with efficient inference

In $0.20/1M
Out $0.80/1M
GLM-4.7Z.aiZ.ai

GLM-4.7

An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents

In $0.20/1M
Out $0.77/1M
DeepSeek-V3-0324-EnterpriseDeepseekDeepseek

DeepSeek-V3-0324-Enterprise

High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation

In $0.17/1M
Out $0.54/1M
DeepSeek-V3.1-EnterpriseDeepseekDeepseek

DeepSeek-V3.1-Enterprise

Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent

$48/M
Nano-Banana-2GoogleGoogle

Nano-Banana-2

Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.

$0.015/image
Alibaba-Z-Image-TurboZ-ImageZ-Image

Alibaba-Z-Image-Turbo

Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.

$0.035/image
Bytedance-Seedream-5.0-LiteBytedanceBytedance

Bytedance-Seedream-5.0-Lite

Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.

$0.04/image
Bytedance-Seedream-4.5BytedanceBytedance

Bytedance-Seedream-4.5

The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals

$0.10/second
Alibaba-Wan-2.6-T2VWanWan

Alibaba-Wan-2.6-T2V

Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.10/second
Alibaba-Wan-2.6-I2VWanWan

Alibaba-Wan-2.6-I2V

Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.025/second
Alibaba-Wan2.6-I2V-FlashWanWan

Alibaba-Wan2.6-I2V-Flash

Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing

$0.10/second
Wan2.6-R2VWanWan

Wan2.6-R2V

Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)

$0.025/second
Alibaba-Wan-2.6-R2V-FlashWanWan

Alibaba-Wan-2.6-R2V-Flash

Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.

In $0.20/1M
Out $1.60/1M
Context-length tiered pricing
Qwen3-VL-PlusQwenQwen

Qwen3-VL-Plus

Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual reasoning and agent capabilities

In $0.05/1M
Out $0.40/1M
Context-length tiered pricing
Qwen3-VL-FlashQwenQwen

Qwen3-VL-Flash

Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on text, image, and video tasks with integrated thinking and non-thinking modes

50 / day free
then $0.002/query
Web SearchMegaNovaMegaNova

Web Search

Real-time managed natural-language web search that returns ranked results, optionally enriched with full extracted page text in a single call—includes 50 free queries per user per day, with additional queries priced at $0.002 each

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.7-Enterprise

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.6-Enterprise

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

Featured
In $1.50/1M
Out $9.00/1M
Gemini-3.5-Flash-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Enterprise

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Enterprise

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Preview-Enterprise

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $2.00/1M
Out $12.00/1M
Token-length tiered pricing
Gemini-3.1-Pro-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Pro-Preview-Enterprise

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

Featured
In $0.75/1M
Out $12.00/1M
Gemini-3.1-Flash-LiveGoogleGoogle

Gemini-3.1-Flash-Live

Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket

DeepSeek
In $0.09/1M
Out $0.18/1M
DeepSeek-V4-Flash-0731DeepseekDeepseek

DeepSeek-V4-Flash-0731

DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding performance, at 284B total / 13B active parameters and a 1M-token context window

In $1.30/1M
Out $2.60/1M
DeepSeek-V4-ProDeepseekDeepseek

DeepSeek-V4-Pro

DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic performance with a 1M context window

In $0.10/1M
Out $0.20/1M
DeepSeek-V4-FlashDeepseekDeepseek

DeepSeek-V4-Flash

DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M context window

In $0.26/1M
Out $0.38/1M
DeepSeek-V3.2DeepseekDeepseek

DeepSeek-V3.2

A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and general chat capabilities with efficient inference

In $0.26/1M
Out $0.38/1M
DeepSeek-V3.2-ExpDeepseekDeepseek

DeepSeek-V3.2-Exp

Experimental version with sparse-attention architecture for long-context efficiency

In $0.20/1M
Out $0.77/1M
DeepSeek-V3-0324DeepseekDeepseek

DeepSeek-V3-0324

High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation

In $0.20/1M
Out $0.77/1M
DeepSeek-V3-0324-EnterpriseDeepseekDeepseek

DeepSeek-V3-0324-Enterprise

High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation

In $0.19/1M
Out $0.79/1M
DeepSeek-V3.1DeepseekDeepseek

DeepSeek-V3.1

Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent

In $0.17/1M
Out $0.54/1M
DeepSeek-V3.1-EnterpriseDeepseekDeepseek

DeepSeek-V3.1-Enterprise

Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent

In $0.50/1M
Out $2.15/1M
DeepSeek-R1-0528DeepseekDeepseek

DeepSeek-R1-0528

An open-source next-generation reasoning-optimized language model with enhanced logic, math, and code performance, larger context handling, and improved function-calling capabilities

Multimodal
In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5-Enterprise

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

Featured
Free
Gemini-3.6-FlashGoogleGoogle

Gemini-3.6-Flash

Available for Creator Tier+

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
In $1.50/1M
Out $7.50/1M
Gemini-3.6-Flash-EnterpriseGoogleGoogle

Gemini-3.6-Flash-Enterprise

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
Free
Gemini-3.5-Flash-LiteGoogleGoogle

Gemini-3.5-Flash-Lite

Available for Creator Tier+

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

Featured
In $0.30/1M
Out $2.50/1M
Gemini-3.5-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Lite-Enterprise

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

In $3.00/1M
Out $15.00/1M
Kimi-K3Moonshot AIMoonshot AI

Kimi-K3

Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning

In $5.00/1M
Out $30.00/1M
GPT-5.6-SolOpen AIOpen AI

GPT-5.6-Sol

OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding

In $2.50/1M
Out $15.00/1M
GPT-5.6-TerraOpen AIOpen AI

GPT-5.6-Terra

OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers

In $2.00/1M
Out $6.00/1M
xAI

Grok-4.5

xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input

In $1.60/1M
Out $8.00/1M
Anthropic

Claude-Sonnet-5

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $0.75/1M
Out $3.50/1M
Kimi-K2.7-CodeMoonshot AIMoonshot AI

Kimi-K2.7-Code

Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering

In $10.00/1M
Out $50.00/1M
Anthropic

Claude-Fable-5

Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M3

A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling

In $4.25/1M
Out $21.25/1M
Anthropic

Claude-Opus-4.8

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

Featured
In $1.20/1M
Out $7.20/1M
Gemini-3.5-FlashGoogleGoogle

Gemini-3.5-Flash

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-LiteGoogleGoogle

Gemini-3.1-Flash-Lite

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

In $1.25/1M
Out $2.50/1M
xAI

Grok-4.3

xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input

In $4.00/1M
Out $24.00/1M
GPT-5.5Open AIOpen AI

GPT-5.5

OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks

In $0.40/1M
Out $2.00/1M
Xiaomi

MiMo-V2.5

Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning

In $0.75/1M
Out $3.50/1M
Kimi-K2.6Moonshot AIMoonshot AI

Kimi-K2.6

An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context window, native vision input, and scaling up to 300 parallel sub-agents

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.7

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $0.50/1M
Out $3.00/1M
Qwen3.6-PlusQwenQwen

Qwen3.6-Plus

Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks

In $0.13/1M
Out $0.38/1M
Gemma-4-31BGoogleGoogle

Gemma-4-31B

Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows

In $0.40/1M
Out $2.00/1M
AI Agent ModelsAI Agent Models

MiMo-V2-Omni

Xiaomi's omni-modal agent model that natively understands text, images, video, and audio in a unified architecture, designed for complex real-world multimodal tasks

In $0.15/1M
Out $0.60/1M
Mistral-Small-4MistralMistral

Mistral-Small-4

A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models

Featured
In $0.20/1M
Out $1.20/1M
Gemini-3.1-Flash-Lite-PreviewGoogleGoogle

Gemini-3.1-Flash-Lite-Preview

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $1.60/1M
Out $9.60/1M
Token-length tiered pricing
Gemini-3.1-Pro-PreviewGoogleGoogle

Gemini-3.1-Pro-Preview

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

In $0.40/1M
In $0.28/1M
Qwen3.5-PlusQwenQwen

Qwen3.5-Plus

Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing

In $0.45/1M
Out $2.25/1M
Kimi-K2.5Moonshot AIMoonshot AI

Kimi-K2.5

An open-source native multimodal AI model that combines vision and language understanding with advanced agent-oriented capabilities, including self-directed agent swarm and visual coding, for versatile reasoning and tool use

In $2.40/1M
Out $12.00/1M
Anthropic

Claude-Sonnet-4.6

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.6

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.5

An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads

In $0.80/1M
Out $4.00/1M
Anthropic

Claude-Haiku-4.5

A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications

In $2.40/1M
Out $12.00/1M
Anthropic

Claude-Sonnet-4.5

A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks

In $12.00/1M
Out $60.00/1M
Anthropic

Claude-Opus-4

Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks

Featured
In $0.40/1M
Out $2.40/1M
Gemini-3-Flash-PreviewGoogleGoogle

Gemini-3-Flash-Preview

Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed

Featured
In $1.00/1M
Out $8.00/1M
Token-length tiered pricing
Gemini-2.5-ProGoogleGoogle

Gemini-2.5-Pro

Strongest Gemini, 1M context, great for code & knowledge

Featured
In $0.24/1M
Out $2.00/1M
Gemini-2.5-FlashGoogleGoogle

Gemini-2.5-Flash

Fast reasoning with improved latency & accuracy

Featured
In $0.08/1M
Out $0.32/1M
Gemini-2.5-Flash-LiteGoogleGoogle

Gemini-2.5-Flash-Lite

Ultra-low latency Gemini variant

In $2.00/1M
Out $12.00/1M
GPT-5.4Open AIOpen AI

GPT-5.4

A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding, and tool-use capabilities with very long context support

In $1.40/1M
Out $11.20/1M
GPT-5.3-ChatOpen AIOpen AI

GPT-5.3-Chat

A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality dialogue, instruction following, and everyday chat applications

In $1.40/1M
Out $11.20/1M
GPT-5.3-CodexOpen AIOpen AI

GPT-5.3-Codex

A high-performance agentic coding model designed to autonomously write, debug, and manage complex software development tasks while collaborating interactively with developers

In $1.40/1M
Out $11.20/1M
GPT-5.2Open AIOpen AI

GPT-5.2

OpenAI’s flagship multimodal (text + image) GPT-5.2 model for top-tier coding and agentic tasks, producing text output with the highest reasoning capability in the GPT-5 family

In $1.00/1M
Out $8.00/1M
GPT-5.1Open AIOpen AI

GPT-5.1

OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs

In $1.00/1M
Out $8.00/1M
GPT-5Open AIOpen AI

GPT-5

OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs

In $0.20/1M
Out $1.60/1M
GPT-5-MiniOpen AIOpen AI

GPT-5-Mini

OpenAI’s faster, more cost-efficient GPT-5 model for well-defined tasks, supporting text and image input with text output

In $0.04/1M
Out $0.32/1M
GPT-5-NanoOpen AIOpen AI

GPT-5-Nano

OpenAI’s fastest, most cost-efficient GPT-5 model for lightweight tasks like summarization and classification, supporting text and image input with text output

In $0.12/1M
Out $0.48/1M
GPT-4o-miniOpen AIOpen AI

GPT-4o-mini

Compact GPT-4 Omni, supports text & image inputs

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-5-Enterprise

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.7-Enterprise

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.6-Enterprise

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.5-Enterprise

An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads

In $15.00/1M
Out $75.00/1M
Anthropic

Claude-Opus-4-Enterprise

Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-4.6-Enterprise

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-4.5-Enterprise

A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks

In $1.00/1M
Out $5.00/1M
Anthropic

Claude-Haiku-4.5-Enterprise

A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications

Featured
In $1.50/1M
Out $9.00/1M
Gemini-3.5-Flash-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Enterprise

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Enterprise

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Preview-Enterprise

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $2.00/1M
Out $12.00/1M
Token-length tiered pricing
Gemini-3.1-Pro-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Pro-Preview-Enterprise

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

Featured
In $0.50/1M
Out $3.00/1M
Gemini-3-Flash-Preview-EnterpriseGoogleGoogle

Gemini-3-Flash-Preview-Enterprise

Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed

Featured
In $1.25/1M
Out $10.00/1M
Token-length tiered pricing
Gemini-2.5-Pro-EnterpriseGoogleGoogle

Gemini-2.5-Pro-Enterprise

Strongest Gemini, 1M context, great for code & knowledge

Featured
In $0.30/1M
Out $2.50/1M
Gemini-2.5-Flash-EnterpriseGoogleGoogle

Gemini-2.5-Flash-Enterprise

Fast reasoning with improved latency & accuracy

Featured
In $0.10/1M
Out $0.40/1M
Gemini-2.5-Flash-Lite-EnterpriseGoogleGoogle

Gemini-2.5-Flash-Lite-Enterprise

Ultra-low latency Gemini variant

Claude
In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-5-Enterprise

Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window

In $1.60/1M
Out $8.00/1M
Anthropic

Claude-Sonnet-5

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $10.00/1M
Out $50.00/1M
Anthropic

Claude-Fable-5

Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity

In $4.25/1M
Out $21.25/1M
Anthropic

Claude-Opus-4.8

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.7

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $2.40/1M
Out $12.00/1M
Anthropic

Claude-Sonnet-4.6

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.6

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $4.00/1M
Out $20.00/1M
Anthropic

Claude-Opus-4.5

An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads

In $0.80/1M
Out $4.00/1M
Anthropic

Claude-Haiku-4.5

A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications

In $2.40/1M
Out $12.00/1M
Anthropic

Claude-Sonnet-4.5

A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks

In $12.00/1M
Out $60.00/1M
Anthropic

Claude-Opus-4

Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-5-Enterprise

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.7-Enterprise

Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.6-Enterprise

Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks

In $5.00/1M
Out $25.00/1M
Anthropic

Claude-Opus-4.5-Enterprise

An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads

In $15.00/1M
Out $75.00/1M
Anthropic

Claude-Opus-4-Enterprise

Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-4.6-Enterprise

Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window

In $3.00/1M
Out $15.00/1M
Anthropic

Claude-Sonnet-4.5-Enterprise

A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks

In $1.00/1M
Out $5.00/1M
Anthropic

Claude-Haiku-4.5-Enterprise

A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications

Free
Featured
Free
Manta-Mini-1.0MegaNovaMegaNova

Manta-Mini-1.0

Lightweight tier optimized for speed and cost. (Free)

Featured
Free
Manta-Flash-1.0MegaNovaMegaNova

Manta-Flash-1.0

Balanced tier for general-purpose roleplay, good tradeoff between cost and latency. (Mid-tier)

Featured
Free
Manta-Pro-1.0MegaNovaMegaNova

Manta-Pro-1.0

Available for Creator Tier+

Premium tier with advanced reasoning and long context, best for complex roleplay.(Top-tier)

Featured
Free
Gemini-3.6-FlashGoogleGoogle

Gemini-3.6-Flash

Available for Creator Tier+

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
Free
Gemini-3.5-Flash-LiteGoogleGoogle

Gemini-3.5-Flash-Lite

Available for Creator Tier+

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

Free
GLM-4.7-FlashZ.aiZ.ai

GLM-4.7-Flash

Available for Creator Tier+

A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks

Free
AI Agent ModelsAI Agent Models

Sapphira-L3.3-70B-0.1

70B parameter language model optimized for storytelling, dialogue, and immersive role-play

Free
AI Agent ModelsAI Agent Models

MN-Violet-Lotus-12B

12B parameter model for creative, emotionally intelligent conversations

Free
AI Agent ModelsAI Agent Models

L3.3-MS-Nevoria-70B

70B narrative generation, LoRA tuning & story-rich outputs

Free
Mistral-Small-3.2-24B-Instruct-2506MistralMistral

Mistral-Small-3.2-24B-Instruct-2506

24B instruction model, long context, fewer errors

Free
AI Agent ModelsAI Agent Models

L3-70B-Euryale-v2.1

70B tuned for storytelling, roleplay, creative writing

Free
AI Agent ModelsAI Agent Models

L3-8B-Stheno-v3.2

8B LLaMA-3 model for roleplay & assistant tasks

Gemini
Featured
Free
Gemini-3.6-FlashGoogleGoogle

Gemini-3.6-Flash

Available for Creator Tier+

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
In $1.50/1M
Out $7.50/1M
Gemini-3.6-Flash-EnterpriseGoogleGoogle

Gemini-3.6-Flash-Enterprise

Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window

Featured
Free
Gemini-3.5-Flash-LiteGoogleGoogle

Gemini-3.5-Flash-Lite

Available for Creator Tier+

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

Featured
In $0.30/1M
Out $2.50/1M
Gemini-3.5-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Lite-Enterprise

Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads

Featured
In $1.20/1M
Out $7.20/1M
Gemini-3.5-FlashGoogleGoogle

Gemini-3.5-Flash

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-LiteGoogleGoogle

Gemini-3.1-Flash-Lite

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

Featured
In $0.20/1M
Out $1.20/1M
Gemini-3.1-Flash-Lite-PreviewGoogleGoogle

Gemini-3.1-Flash-Lite-Preview

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $1.60/1M
Out $9.60/1M
Token-length tiered pricing
Gemini-3.1-Pro-PreviewGoogleGoogle

Gemini-3.1-Pro-Preview

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

Featured
In $0.40/1M
Out $2.40/1M
Gemini-3-Flash-PreviewGoogleGoogle

Gemini-3-Flash-Preview

Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed

Featured
In $1.00/1M
Out $8.00/1M
Token-length tiered pricing
Gemini-2.5-ProGoogleGoogle

Gemini-2.5-Pro

Strongest Gemini, 1M context, great for code & knowledge

Featured
In $0.24/1M
Out $2.00/1M
Gemini-2.5-FlashGoogleGoogle

Gemini-2.5-Flash

Fast reasoning with improved latency & accuracy

Featured
In $0.08/1M
Out $0.32/1M
Gemini-2.5-Flash-LiteGoogleGoogle

Gemini-2.5-Flash-Lite

Ultra-low latency Gemini variant

$48/M
Nano-Banana-2GoogleGoogle

Nano-Banana-2

Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.

$96/M
Nano-Banana-Pro-EditGoogleGoogle

Nano-Banana-Pro-Edit

Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).

$96/M
Nano-Banana-ProGoogleGoogle

Nano-Banana-Pro

High-quality image generation with better text rendering, character consistency, and advanced composition. Powered by Gemini 3 Pro Image.

$24/M
Nano-Banana-EditGoogleGoogle

Nano-Banana-Edit

AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.

$24/M
Nano-BananaGoogleGoogle

Nano-Banana

Gemini 2.5 Flash Image - High-quality image generation with multiple aspect ratios

In $0.30/1M
Out $6.00/1M
Gemini-2.5-Flash-TTSGoogleGoogle

Gemini-2.5-Flash-TTS

A low-latency, cost-efficient text-to-speech model that converts text into expressive single or multi-speaker audio with controllable tone, style, and pacing

In $0.60/1M
Out $12.00/1M
Gemini-2.5-Pro-TTSGoogleGoogle

Gemini-2.5-Pro-TTS

A high-quality text-to-speech model designed to generate natural, expressive audio with fine-grained control over tone, pacing, and emotion for narration, assistants, and conversational voice applications

Featured
In $1.50/1M
Out $9.00/1M
Gemini-3.5-Flash-EnterpriseGoogleGoogle

Gemini-3.5-Flash-Enterprise

Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Enterprise

Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs

Featured
In $0.25/1M
Out $1.50/1M
Gemini-3.1-Flash-Lite-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Flash-Lite-Preview-Enterprise

Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications

Featured
In $2.00/1M
Out $12.00/1M
Token-length tiered pricing
Gemini-3.1-Pro-Preview-EnterpriseGoogleGoogle

Gemini-3.1-Pro-Preview-Enterprise

Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window

Featured
In $0.50/1M
Out $3.00/1M
Gemini-3-Flash-Preview-EnterpriseGoogleGoogle

Gemini-3-Flash-Preview-Enterprise

Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed

Featured
In $1.25/1M
Out $10.00/1M
Token-length tiered pricing
Gemini-2.5-Pro-EnterpriseGoogleGoogle

Gemini-2.5-Pro-Enterprise

Strongest Gemini, 1M context, great for code & knowledge

Featured
In $0.30/1M
Out $2.50/1M
Gemini-2.5-Flash-EnterpriseGoogleGoogle

Gemini-2.5-Flash-Enterprise

Fast reasoning with improved latency & accuracy

Featured
In $0.10/1M
Out $0.40/1M
Gemini-2.5-Flash-Lite-EnterpriseGoogleGoogle

Gemini-2.5-Flash-Lite-Enterprise

Ultra-low latency Gemini variant

Featured
In $0.75/1M
Out $12.00/1M
Gemini-3.1-Flash-LiveGoogleGoogle

Gemini-3.1-Flash-Live

Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket

Qwen
In $1.50/1M
Out $5.00/1M
Qwen3.8-Max-PreviewQwenQwen

Qwen3.8-Max-Preview

Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data analysis, and office workflows; always uses thinking mode and supports long-horizon tool workflows and structured output

In $0.12/1M
Out $0.76/1M
Qwen3.6-35B-A3BQwenQwen

Qwen3.6-35B-A3B

Qwen3.6 35B (3B-active) MoE with built-in NEXTN speculative decoding, reasoning and tool calling — fast low-cost generation

In $0.50/1M
Out $3.00/1M
Qwen3.6-PlusQwenQwen

Qwen3.6-Plus

Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks

In $0.40/1M
In $0.28/1M
Qwen3.5-PlusQwenQwen

Qwen3.5-Plus

Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing

In $0.20/1M
Out $1.60/1M
Context-length tiered pricing
Qwen3-VL-PlusQwenQwen

Qwen3-VL-Plus

Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual reasoning and agent capabilities

In $0.05/1M
Out $0.40/1M
Context-length tiered pricing
Qwen3-VL-FlashQwenQwen

Qwen3-VL-Flash

Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on text, image, and video tasks with integrated thinking and non-thinking modes

In $0.20/1M
Out $0.60/1M
Qwen2.5-VL-32B-InstructQwenQwen

Qwen2.5-VL-32B-Instruct

A 32B-parameter vision-language model tuned to follow instructions and reason over images and text

Free
Qwen3-Embedding-8BQwenQwen

Qwen3-Embedding-8B

8B embedding model, strong multilingual & code support, top MTEB scorer

Kimi
In $3.00/1M
Out $15.00/1M
Kimi-K3Moonshot AIMoonshot AI

Kimi-K3

Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning

In $0.75/1M
Out $3.50/1M
Kimi-K2.7-CodeMoonshot AIMoonshot AI

Kimi-K2.7-Code

Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering

In $0.75/1M
Out $3.50/1M
Kimi-K2.6Moonshot AIMoonshot AI

Kimi-K2.6

An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context window, native vision input, and scaling up to 300 parallel sub-agents

In $0.45/1M
Out $2.25/1M
Kimi-K2.5Moonshot AIMoonshot AI

Kimi-K2.5

An open-source native multimodal AI model that combines vision and language understanding with advanced agent-oriented capabilities, including self-directed agent swarm and visual coding, for versatile reasoning and tool use

In $0.60/1M
Out $2.60/1M
Kimi-K2-ThinkingMoonshot AIMoonshot AI

Kimi-K2-Thinking

1T-parameter reasoning-focused language model designed for complex, multi-step problem solving and tool use

Text To Video
$7.00/1M
Seedance-2.0BytedanceBytedance

Seedance-2.0

BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video

$5.60/1M
Seedance-2.0-FastBytedanceBytedance

Seedance-2.0-Fast

BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens

$3.50/1M
Seedance-2.0-MiniBytedanceBytedance

Seedance-2.0-Mini

BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens

$0.10/second
Alibaba-Wan-2.6-T2VWanWan

Alibaba-Wan-2.6-T2V

Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.06/second
Veo-3.1-FastGoogleGoogle

Veo-3.1-Fast

Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p. Optional AI-generated audio (+50% cost).

$0.12/second
Veo-3.1GoogleGoogle

Veo-3.1

Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds at 720p/1080p.

$2.50/1M
Seedance-1.0-Pro-Text-to-VideoBytedanceBytedance

Seedance-1.0-Pro-Text-to-Video

Pro-tier text-to-video, generates multi-shot 1080p videos with narrative flow

$1.80/1M
Seedance-1.0-Lite-Text-to-VideoBytedanceBytedance

Seedance-1.0-Lite-Text-to-Video

Lite text-to-video, faster, lower compute, shorter outputs

Image To Video
$7.00/1M
Seedance-2.0BytedanceBytedance

Seedance-2.0

BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video

$5.60/1M
Seedance-2.0-FastBytedanceBytedance

Seedance-2.0-Fast

BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens

$3.50/1M
Seedance-2.0-MiniBytedanceBytedance

Seedance-2.0-Mini

BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens

$0.10/second
Alibaba-Wan-2.6-I2VWanWan

Alibaba-Wan-2.6-I2V

Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.025/second
Alibaba-Wan2.6-I2V-FlashWanWan

Alibaba-Wan2.6-I2V-Flash

Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing

$0.06/second
Veo-3.1-Fast-I2VGoogleGoogle

Veo-3.1-Fast-I2V

Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at 720p/1080p.

$0.12/second
Veo-3.1-I2VGoogleGoogle

Veo-3.1-I2V

Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.

$2.50/1M
Seedance-1.0-Pro-Image-to-VideoBytedanceBytedance

Seedance-1.0-Pro-Image-to-Video

Pro-tier model, turns images into cinematic 1080p videos with smooth motion

$1.80/1M
Seedance-1.0-Lite-Image-to-VideoBytedanceBytedance

Seedance-1.0-Lite-Image-to-Video

Lite version, quick image-to-video at up to 1080p, shorter sequences

Video To Video
$7.00/1M
Seedance-2.0BytedanceBytedance

Seedance-2.0

BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video

$5.60/1M
Seedance-2.0-FastBytedanceBytedance

Seedance-2.0-Fast

BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens

$3.50/1M
Seedance-2.0-MiniBytedanceBytedance

Seedance-2.0-Mini

BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens

Best Video Models
$7.00/1M
Seedance-2.0BytedanceBytedance

Seedance-2.0

BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video

$5.60/1M
Seedance-2.0-FastBytedanceBytedance

Seedance-2.0-Fast

BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens

$3.50/1M
Seedance-2.0-MiniBytedanceBytedance

Seedance-2.0-Mini

BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens

$0.10/second
Alibaba-Wan-2.6-T2VWanWan

Alibaba-Wan-2.6-T2V

Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.10/second
Alibaba-Wan-2.6-I2VWanWan

Alibaba-Wan-2.6-I2V

Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.025/second
Alibaba-Wan2.6-I2V-FlashWanWan

Alibaba-Wan2.6-I2V-Flash

Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing

$0.10/second
Wan2.6-R2VWanWan

Wan2.6-R2V

Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)

$0.025/second
Alibaba-Wan-2.6-R2V-FlashWanWan

Alibaba-Wan-2.6-R2V-Flash

Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.

$0.06/second
Veo-3.1-FastGoogleGoogle

Veo-3.1-Fast

Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p. Optional AI-generated audio (+50% cost).

$0.06/second
Veo-3.1-Fast-I2VGoogleGoogle

Veo-3.1-Fast-I2V

Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at 720p/1080p.

$0.12/second
Veo-3.1GoogleGoogle

Veo-3.1

Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds at 720p/1080p.

$0.12/second
Veo-3.1-I2VGoogleGoogle

Veo-3.1-I2V

Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.

$2.50/1M
Seedance-1.0-Pro-Image-to-VideoBytedanceBytedance

Seedance-1.0-Pro-Image-to-Video

Pro-tier model, turns images into cinematic 1080p videos with smooth motion

$2.50/1M
Seedance-1.0-Pro-Text-to-VideoBytedanceBytedance

Seedance-1.0-Pro-Text-to-Video

Pro-tier text-to-video, generates multi-shot 1080p videos with narrative flow

$1.80/1M
Seedance-1.0-Lite-Image-to-VideoBytedanceBytedance

Seedance-1.0-Lite-Image-to-Video

Lite version, quick image-to-video at up to 1080p, shorter sequences

$1.80/1M
Seedance-1.0-Lite-Text-to-VideoBytedanceBytedance

Seedance-1.0-Lite-Text-to-Video

Lite text-to-video, faster, lower compute, shorter outputs

Bytedance
$7.00/1M
Seedance-2.0BytedanceBytedance

Seedance-2.0

BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video

$5.60/1M
Seedance-2.0-FastBytedanceBytedance

Seedance-2.0-Fast

BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens

$3.50/1M
Seedance-2.0-MiniBytedanceBytedance

Seedance-2.0-Mini

BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens

$0.035/image
Bytedance-Seedream-5.0-LiteBytedanceBytedance

Bytedance-Seedream-5.0-Lite

Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.

$0.04/image
Bytedance-Seedream-4.5BytedanceBytedance

Bytedance-Seedream-4.5

The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals

$0.03/image
Bytedance-Seedream-4.0BytedanceBytedance

Bytedance-Seedream-4.0

Next-generation image creation model that unifies image generation and editing to handle complex multimodal tasks with faster inference and produce high-definition images up to 4K

$0.021/image
Bytedance-Seedream-3.0BytedanceBytedance

Bytedance-Seedream-3.0

Bilingual text-to-image, 2K resolution, accurate text & artistic layouts

$2.50/1M
Seedance-1.0-Pro-Image-to-VideoBytedanceBytedance

Seedance-1.0-Pro-Image-to-Video

Pro-tier model, turns images into cinematic 1080p videos with smooth motion

$2.50/1M
Seedance-1.0-Pro-Text-to-VideoBytedanceBytedance

Seedance-1.0-Pro-Text-to-Video

Pro-tier text-to-video, generates multi-shot 1080p videos with narrative flow

$1.80/1M
Seedance-1.0-Lite-Image-to-VideoBytedanceBytedance

Seedance-1.0-Lite-Image-to-Video

Lite version, quick image-to-video at up to 1080p, shorter sequences

$1.80/1M
Seedance-1.0-Lite-Text-to-VideoBytedanceBytedance

Seedance-1.0-Lite-Text-to-Video

Lite text-to-video, faster, lower compute, shorter outputs

OpenAI
In $5.00/1M
Out $30.00/1M
GPT-5.6-SolOpen AIOpen AI

GPT-5.6-Sol

OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding

In $2.50/1M
Out $15.00/1M
GPT-5.6-TerraOpen AIOpen AI

GPT-5.6-Terra

OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers

In $4.00/1M
Out $24.00/1M
GPT-5.5Open AIOpen AI

GPT-5.5

OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks

In $2.00/1M
Out $12.00/1M
GPT-5.4Open AIOpen AI

GPT-5.4

A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding, and tool-use capabilities with very long context support

In $1.40/1M
Out $11.20/1M
GPT-5.3-ChatOpen AIOpen AI

GPT-5.3-Chat

A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality dialogue, instruction following, and everyday chat applications

In $1.40/1M
Out $11.20/1M
GPT-5.3-CodexOpen AIOpen AI

GPT-5.3-Codex

A high-performance agentic coding model designed to autonomously write, debug, and manage complex software development tasks while collaborating interactively with developers

In $1.40/1M
Out $11.20/1M
GPT-5.2Open AIOpen AI

GPT-5.2

OpenAI’s flagship multimodal (text + image) GPT-5.2 model for top-tier coding and agentic tasks, producing text output with the highest reasoning capability in the GPT-5 family

In $1.00/1M
Out $8.00/1M
GPT-5.1Open AIOpen AI

GPT-5.1

OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs

In $1.00/1M
Out $8.00/1M
GPT-5Open AIOpen AI

GPT-5

OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs

In $0.20/1M
Out $1.60/1M
GPT-5-MiniOpen AIOpen AI

GPT-5-Mini

OpenAI’s faster, more cost-efficient GPT-5 model for well-defined tasks, supporting text and image input with text output

In $0.04/1M
Out $0.32/1M
GPT-5-NanoOpen AIOpen AI

GPT-5-Nano

OpenAI’s fastest, most cost-efficient GPT-5 model for lightweight tasks like summarization and classification, supporting text and image input with text output

In $0.12/1M
Out $0.48/1M
GPT-4o-miniOpen AIOpen AI

GPT-4o-mini

Compact GPT-4 Omni, supports text & image inputs

Grok
In $2.00/1M
Out $6.00/1M
xAI

Grok-4.5

xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input

In $1.25/1M
Out $2.50/1M
xAI

Grok-4.3

xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input

GLM
In $1.00/1M
Out $3.00/1M
GLM-5.2Z.aiZ.ai

GLM-5.2

Z.ai's open-source MoE model for long-context reasoning and agentic software engineering, with a 1M-token context window

In $1.05/1M
Out $3.50/1M
GLM-5.1Z.aiZ.ai

GLM-5.1

Z.ai's open-source flagship MoE model built for agentic engineering, delivering state-of-the-art coding and long-horizon task performance that improves the longer it runs

Free
GLM-4.7-FlashZ.aiZ.ai

GLM-4.7-Flash

Available for Creator Tier+

A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks

In $0.80/1M
Out $2.56/1M
GLM-5Z.aiZ.ai

GLM-5

Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding, and agentic tasks

In $0.20/1M
Out $0.80/1M
GLM-4.7Z.aiZ.ai

GLM-4.7

An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents

In $0.20/1M
Out $0.80/1M
GLM-4.7-EnterpriseZ.aiZ.ai

GLM-4.7-Enterprise

An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents

In $0.45/1M
Out $1.90/1M
GLM-4.6Z.aiZ.ai

GLM-4.6

Flagship model from Z.ai with 200K token context, superior reasoning, advanced tool use, and coding excellence

MiniMax
In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M3

A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M2.7

A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity tasks like coding, tool use, and multi-step reasoning — at a highly competitive price point

In $0.30/1M
Out $1.20/1M
MiniMax

MiniMax-M2.5

MiniMax-M2.5 is a cost-efficient frontier model that achieves SOTA in coding, agentic tool use, and search tasks through extensive reinforcement learning in real-world environments

In $0.28/1M
Out $1.20/1M
MiniMax

MiniMax-M2.1

A lightweight, open-source text-to-text large language model optimized for coding, agent-oriented workflows, and instruction following, delivering faster, cleaner outputs and strong reasoning and problem-solving performance compared to its predecessor

Xiaomi
In $1.00/1M
Out $3.00/1M
Xiaomi

MiMo-V2.5-Pro

Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon tasks with a 1M context window

In $0.40/1M
Out $2.00/1M
Xiaomi

MiMo-V2.5

Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning

In $0.40/1M
Out $2.00/1M
AI Agent ModelsAI Agent Models

MiMo-V2-Omni

Xiaomi's omni-modal agent model that natively understands text, images, video, and audio in a unified architecture, designed for complex real-world multimodal tasks

In $1.00/1M
Out $3.00/1M
Xiaomi

MiMo-V2-Pro

A trillion-parameter text-only flagship agent model from Xiaomi, built for complex coding, long-horizon planning, and multi-step tool use with a 1M context window

Google
In $0.13/1M
Out $0.38/1M
Gemma-4-31BGoogleGoogle

Gemma-4-31B

Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows

Mistral
In $0.15/1M
Out $0.60/1M
Mistral-Small-4MistralMistral

Mistral-Small-4

A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models

In $0.02/1M
Out $0.04/1M
Mistral-Nemo-Instruct-2407MistralMistral

Mistral-Nemo-Instruct-2407

24B instruction model, long context, fewer errors

Free
Mistral-Small-3.2-24B-Instruct-2506MistralMistral

Mistral-Small-3.2-24B-Instruct-2506

24B instruction model, long context, fewer errors

Roleplay Models
In $0.09/1M
Out $0.10/1M
Qwen3-235B-A22B-Instruct-2507QwenQwen

Qwen3-235B-A22B-Instruct-2507

A 235B-parameter MoE instruction-tuned language model for general, multilingual, and coding tasks

In $0.30/1M
Out $0.50/1M
AI Agent ModelsAI Agent Models

Cydonia-24B-v4.3

A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence

In $0.30/1M
Out $0.50/1M
AI Agent ModelsAI Agent Models

Cydonia-24B-v4.1

A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence

In $0.85/1M
Out $0.85/1M
AI Agent ModelsAI Agent Models

L3.3-70B-Euryale-v2.3

70B model focused on creative roleplay from Sao10k

In $0.85/1M
Out $0.85/1M
AI Agent ModelsAI Agent Models

L3.1-70B-Euryale-v2.2

70B model focused on creative roleplay from Sao10k

Free
AI Agent ModelsAI Agent Models

Sapphira-L3.3-70B-0.1

70B parameter language model optimized for storytelling, dialogue, and immersive role-play

Free
AI Agent ModelsAI Agent Models

MN-Violet-Lotus-12B

12B parameter model for creative, emotionally intelligent conversations

Free
AI Agent ModelsAI Agent Models

L3.3-MS-Nevoria-70B

70B narrative generation, LoRA tuning & story-rich outputs

Free
AI Agent ModelsAI Agent Models

L3-70B-Euryale-v2.1

70B tuned for storytelling, roleplay, creative writing

Free
AI Agent ModelsAI Agent Models

L3-8B-Stheno-v3.2

8B LLaMA-3 model for roleplay & assistant tasks

Manta
Featured
Free
Manta-Mini-1.0MegaNovaMegaNova

Manta-Mini-1.0

Lightweight tier optimized for speed and cost. (Free)

Featured
Free
Manta-Flash-1.0MegaNovaMegaNova

Manta-Flash-1.0

Balanced tier for general-purpose roleplay, good tradeoff between cost and latency. (Mid-tier)

Featured
Free
Manta-Pro-1.0MegaNovaMegaNova

Manta-Pro-1.0

Available for Creator Tier+

Premium tier with advanced reasoning and long context, best for complex roleplay.(Top-tier)

Llama
In $0.10/1M
Out $0.30/1M
Llama3.3-70BLlamaLlama

Llama3.3-70B

Multilingual 70B delivering 405B-level performance

Text To Image
$48/M
Nano-Banana-2GoogleGoogle

Nano-Banana-2

Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.

$0.015/image
Alibaba-Z-Image-TurboZ-ImageZ-Image

Alibaba-Z-Image-Turbo

Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.

$96/M
Nano-Banana-Pro-EditGoogleGoogle

Nano-Banana-Pro-Edit

Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).

$96/M
Nano-Banana-ProGoogleGoogle

Nano-Banana-Pro

High-quality image generation with better text rendering, character consistency, and advanced composition. Powered by Gemini 3 Pro Image.

$24/M
Nano-Banana-EditGoogleGoogle

Nano-Banana-Edit

AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.

$24/M
Nano-BananaGoogleGoogle

Nano-Banana

Gemini 2.5 Flash Image - High-quality image generation with multiple aspect ratios

$0.002/image
AI Agent ModelsAI Agent Models

DreamShaper-8

A Stable Diffusion v1.5-based text-to-image model, balancing photorealistic and stylized (including anime) generation with strong LoRA compatibility

$0.035/image
Bytedance-Seedream-5.0-LiteBytedanceBytedance

Bytedance-Seedream-5.0-Lite

Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.

$0.04/image
Bytedance-Seedream-4.5BytedanceBytedance

Bytedance-Seedream-4.5

The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals

$0.03/image
Bytedance-Seedream-4.0BytedanceBytedance

Bytedance-Seedream-4.0

Next-generation image creation model that unifies image generation and editing to handle complex multimodal tasks with faster inference and produce high-definition images up to 4K

$0.021/image
Bytedance-Seedream-3.0BytedanceBytedance

Bytedance-Seedream-3.0

Bilingual text-to-image, 2K resolution, accurate text & artistic layouts

Image Generation
$48/M
Nano-Banana-2GoogleGoogle

Nano-Banana-2

Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.

Best Image Models
$48/M
Nano-Banana-2GoogleGoogle

Nano-Banana-2

Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.

$0.015/image
Alibaba-Z-Image-TurboZ-ImageZ-Image

Alibaba-Z-Image-Turbo

Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.

$96/M
Nano-Banana-Pro-EditGoogleGoogle

Nano-Banana-Pro-Edit

Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).

$96/M
Nano-Banana-ProGoogleGoogle

Nano-Banana-Pro

High-quality image generation with better text rendering, character consistency, and advanced composition. Powered by Gemini 3 Pro Image.

$24/M
Nano-Banana-EditGoogleGoogle

Nano-Banana-Edit

AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.

$24/M
Nano-BananaGoogleGoogle

Nano-Banana

Gemini 2.5 Flash Image - High-quality image generation with multiple aspect ratios

$0.002/image
AI Agent ModelsAI Agent Models

DreamShaper-8

A Stable Diffusion v1.5-based text-to-image model, balancing photorealistic and stylized (including anime) generation with strong LoRA compatibility

$0.035/image
Bytedance-Seedream-5.0-LiteBytedanceBytedance

Bytedance-Seedream-5.0-Lite

Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.

$0.04/image
Bytedance-Seedream-4.5BytedanceBytedance

Bytedance-Seedream-4.5

The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals

$0.03/image
Bytedance-Seedream-4.0BytedanceBytedance

Bytedance-Seedream-4.0

Next-generation image creation model that unifies image generation and editing to handle complex multimodal tasks with faster inference and produce high-definition images up to 4K

$0.021/image
Bytedance-Seedream-3.0BytedanceBytedance

Bytedance-Seedream-3.0

Bilingual text-to-image, 2K resolution, accurate text & artistic layouts

Z-Image
$0.015/image
Alibaba-Z-Image-TurboZ-ImageZ-Image

Alibaba-Z-Image-Turbo

Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.

Image To Image
$96/M
Nano-Banana-Pro-EditGoogleGoogle

Nano-Banana-Pro-Edit

Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).

$24/M
Nano-Banana-EditGoogleGoogle

Nano-Banana-Edit

AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.

$0.035/image
Bytedance-Seedream-5.0-LiteBytedanceBytedance

Bytedance-Seedream-5.0-Lite

Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.

$0.04/image
Bytedance-Seedream-4.5BytedanceBytedance

Bytedance-Seedream-4.5

The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals

$0.03/image
Bytedance-Seedream-4.0BytedanceBytedance

Bytedance-Seedream-4.0

Next-generation image creation model that unifies image generation and editing to handle complex multimodal tasks with faster inference and produce high-definition images up to 4K

MegaNova
$0.002/image
AI Agent ModelsAI Agent Models

DreamShaper-8

A Stable Diffusion v1.5-based text-to-image model, balancing photorealistic and stylized (including anime) generation with strong LoRA compatibility

50 / day free
then $0.002/query
Web SearchMegaNovaMegaNova

Web Search

Real-time managed natural-language web search that returns ranked results, optionally enriched with full extracted page text in a single call—includes 50 free queries per user per day, with additional queries priced at $0.002 each

Wan
$0.10/second
Alibaba-Wan-2.6-T2VWanWan

Alibaba-Wan-2.6-T2V

Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.10/second
Alibaba-Wan-2.6-I2VWanWan

Alibaba-Wan-2.6-I2V

Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.

$0.025/second
Alibaba-Wan2.6-I2V-FlashWanWan

Alibaba-Wan2.6-I2V-Flash

Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing

$0.10/second
Wan2.6-R2VWanWan

Wan2.6-R2V

Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)

$0.025/second
Alibaba-Wan-2.6-R2V-FlashWanWan

Alibaba-Wan-2.6-R2V-Flash

Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.

Reference To Video
$0.10/second
Wan2.6-R2VWanWan

Wan2.6-R2V

Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)

$0.025/second
Alibaba-Wan-2.6-R2V-FlashWanWan

Alibaba-Wan-2.6-R2V-Flash

Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.

Veo
$0.06/second
Veo-3.1-FastGoogleGoogle

Veo-3.1-Fast

Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p. Optional AI-generated audio (+50% cost).

$0.06/second
Veo-3.1-Fast-I2VGoogleGoogle

Veo-3.1-Fast-I2V

Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at 720p/1080p.

$0.12/second
Veo-3.1GoogleGoogle

Veo-3.1

Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds at 720p/1080p.

$0.12/second
Veo-3.1-I2VGoogleGoogle

Veo-3.1-I2V

Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.

Vision
In $0.20/1M
Out $1.60/1M
Context-length tiered pricing
Qwen3-VL-PlusQwenQwen

Qwen3-VL-Plus

Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual reasoning and agent capabilities

In $0.05/1M
Out $0.40/1M
Context-length tiered pricing
Qwen3-VL-FlashQwenQwen

Qwen3-VL-Flash

Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on text, image, and video tasks with integrated thinking and non-thinking modes

In $0.20/1M
Out $0.60/1M
Qwen2.5-VL-32B-InstructQwenQwen

Qwen2.5-VL-32B-Instruct

A 32B-parameter vision-language model tuned to follow instructions and reason over images and text

Audio
In $0.30/1M
Out $6.00/1M
Gemini-2.5-Flash-TTSGoogleGoogle

Gemini-2.5-Flash-TTS

A low-latency, cost-efficient text-to-speech model that converts text into expressive single or multi-speaker audio with controllable tone, style, and pacing

In $0.60/1M
Out $12.00/1M
Gemini-2.5-Pro-TTSGoogleGoogle

Gemini-2.5-Pro-TTS

A high-quality text-to-speech model designed to generate natural, expressive audio with fine-grained control over tone, pacing, and emotion for narration, assistants, and conversational voice applications

$0.006/min

Faster-Whisper-Large-V3

High-performance speech-to-text model optimized for fast, accurate transcription using an efficient Whisper-based architecture

Featured
In $0.75/1M
Out $12.00/1M
Gemini-3.1-Flash-LiveGoogleGoogle

Gemini-3.1-Flash-Live

Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket

Text To Speech
In $0.30/1M
Out $6.00/1M
Gemini-2.5-Flash-TTSGoogleGoogle

Gemini-2.5-Flash-TTS

A low-latency, cost-efficient text-to-speech model that converts text into expressive single or multi-speaker audio with controllable tone, style, and pacing

In $0.60/1M
Out $12.00/1M
Gemini-2.5-Pro-TTSGoogleGoogle

Gemini-2.5-Pro-TTS

A high-quality text-to-speech model designed to generate natural, expressive audio with fine-grained control over tone, pacing, and emotion for narration, assistants, and conversational voice applications

Speech To Text
$0.006/min

Faster-Whisper-Large-V3

High-performance speech-to-text model optimized for fast, accurate transcription using an efficient Whisper-based architecture

Whisper
$0.006/min

Faster-Whisper-Large-V3

High-performance speech-to-text model optimized for fast, accurate transcription using an efficient Whisper-based architecture

Embeddings
Free
Qwen3-Embedding-8BQwenQwen

Qwen3-Embedding-8B

8B embedding model, strong multilingual & code support, top MTEB scorer

Reranker
Free
BAAI

BGE-reranker-v2-m3

Multilingual reranker, query+passage → relevance score, lightweight & fast

Web Search
50 / day free
then $0.002/query
Web SearchMegaNovaMegaNova

Web Search

Real-time managed natural-language web search that returns ranked results, optionally enriched with full extracted page text in a single call—includes 50 free queries per user per day, with additional queries priced at $0.002 each

Realtime
Featured
In $0.75/1M
Out $12.00/1M
Gemini-3.1-Flash-LiveGoogleGoogle

Gemini-3.1-Flash-Live

Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket