DeepSeek-V4-Flash-0731
DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding performance, at 284B total / 13B active parameters and a 1M-token context window
Claude-Opus-5
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Claude-Opus-5-Enterprise
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Gemini-3.6-Flash
Available for Creator Tier+
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.6-Flash-Enterprise
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.5-Flash-Lite
Available for Creator Tier+
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Gemini-3.5-Flash-Lite-Enterprise
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Qwen3.8-Max-Preview
Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data analysis, and office workflows; always uses thinking mode and supports long-horizon tool workflows and structured output
Kimi-K3
Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning
Seedance-2.0
BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video
Seedance-2.0-Fast
BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens
Seedance-2.0-Mini
BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens
GPT-5.6-Sol
OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding
GPT-5.6-Terra
OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers
Grok-4.5
xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input
Claude-Sonnet-5
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
GLM-5.2
Z.ai's open-source MoE model for long-context reasoning and agentic software engineering, with a 1M-token context window
Kimi-K2.7-Code
Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering
Claude-Fable-5
Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity
MiniMax-M3
A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling
Qwen3.6-35B-A3B
Qwen3.6 35B (3B-active) MoE with built-in NEXTN speculative decoding, reasoning and tool calling — fast low-cost generation
Claude-Opus-4.8
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Gemini-3.5-Flash
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Grok-4.3
xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input
GPT-5.5
OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks
DeepSeek-V4-Pro
DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic performance with a 1M context window
DeepSeek-V4-Flash
DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M context window
MiMo-V2.5-Pro
Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon tasks with a 1M context window
MiMo-V2.5
Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning
Claude-Opus-4.7
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
GLM-5.1
Z.ai's open-source flagship MoE model built for agentic engineering, delivering state-of-the-art coding and long-horizon task performance that improves the longer it runs
Qwen3.6-Plus
Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks
Gemma-4-31B
Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows
Mistral-Small-4
A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models
Gemini-3.1-Flash-Lite-Preview
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
MiniMax-M2.7
A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity tasks like coding, tool use, and multi-step reasoning — at a highly competitive price point
GLM-4.7-Flash
Available for Creator Tier+
A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks
Qwen3.5-Plus
Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing
GLM-5
Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding, and agentic tasks
Claude-Sonnet-4.6
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Opus-4.6
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Claude-Opus-4.5
An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads
Claude-Haiku-4.5
A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications
Claude-Sonnet-4.5
A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks
Claude-Opus-4
Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks
MiniMax-M2.1
A lightweight, open-source text-to-text large language model optimized for coding, agent-oriented workflows, and instruction following, delivering faster, cleaner outputs and strong reasoning and problem-solving performance compared to its predecessor
Qwen3-235B-A22B-Instruct-2507
A 235B-parameter MoE instruction-tuned language model for general, multilingual, and coding tasks
GLM-4.6
Flagship model from Z.ai with 200K token context, superior reasoning, advanced tool use, and coding excellence
GPT-5.4
A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding, and tool-use capabilities with very long context support
GPT-5.3-Chat
A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality dialogue, instruction following, and everyday chat applications
GPT-5.3-Codex
A high-performance agentic coding model designed to autonomously write, debug, and manage complex software development tasks while collaborating interactively with developers
Cydonia-24B-v4.3
A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence
Cydonia-24B-v4.1
A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence
L3.3-70B-Euryale-v2.3
70B model focused on creative roleplay from Sao10k
L3.1-70B-Euryale-v2.2
70B model focused on creative roleplay from Sao10k
Sapphira-L3.3-70B-0.1
70B parameter language model optimized for storytelling, dialogue, and immersive role-play
MN-Violet-Lotus-12B
12B parameter model for creative, emotionally intelligent conversations
L3.3-MS-Nevoria-70B
70B narrative generation, LoRA tuning & story-rich outputs
Mistral-Nemo-Instruct-2407
24B instruction model, long context, fewer errors
Nano-Banana-2
Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.
Alibaba-Z-Image-Turbo
Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.
Nano-Banana-Pro-Edit
Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).
Nano-Banana-Edit
AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.
Bytedance-Seedream-5.0-Lite
Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.
Alibaba-Wan-2.6-T2V
Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan-2.6-I2V
Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan2.6-I2V-Flash
Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing
Wan2.6-R2V
Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)
Alibaba-Wan-2.6-R2V-Flash
Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.
Veo-3.1-Fast
Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p. Optional AI-generated audio (+50% cost).
Veo-3.1-Fast-I2V
Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at 720p/1080p.
Veo-3.1
Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds at 720p/1080p.
Veo-3.1-I2V
Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.
Qwen3-VL-Plus
Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual reasoning and agent capabilities
Qwen3-VL-Flash
Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on text, image, and video tasks with integrated thinking and non-thinking modes
Web Search
Real-time managed natural-language web search that returns ranked results, optionally enriched with full extracted page text in a single call—includes 50 free queries per user per day, with additional queries priced at $0.002 each
Claude-Sonnet-5-Enterprise
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Opus-4.7-Enterprise
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Claude-Opus-4.6-Enterprise
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Claude-Opus-4.5-Enterprise
An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads
Claude-Opus-4-Enterprise
Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks
Claude-Sonnet-4.6-Enterprise
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Sonnet-4.5-Enterprise
A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks
Claude-Haiku-4.5-Enterprise
A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications
Gemini-3.5-Flash-Enterprise
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite-Enterprise
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Gemini-3.1-Flash-Lite-Preview-Enterprise
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview-Enterprise
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
Gemini-3.1-Flash-Live
Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket
Manta-Mini-1.0
Lightweight tier optimized for speed and cost. (Free)
Manta-Flash-1.0
Balanced tier for general-purpose roleplay, good tradeoff between cost and latency. (Mid-tier)
Manta-Pro-1.0
Available for Creator Tier+
Premium tier with advanced reasoning and long context, best for complex roleplay.(Top-tier)
DeepSeek-V3.2
A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and general chat capabilities with efficient inference
Cydonia-24B-v4.3
A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence
Cydonia-24B-v4.1
A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence
L3.3-70B-Euryale-v2.3
70B model focused on creative roleplay from Sao10k
L3.1-70B-Euryale-v2.2
70B model focused on creative roleplay from Sao10k
Sapphira-L3.3-70B-0.1
70B parameter language model optimized for storytelling, dialogue, and immersive role-play
MN-Violet-Lotus-12B
12B parameter model for creative, emotionally intelligent conversations
L3.3-MS-Nevoria-70B
70B narrative generation, LoRA tuning & story-rich outputs
L3-70B-Euryale-v2.1
70B tuned for storytelling, roleplay, creative writing
L3-8B-Stheno-v3.2
8B LLaMA-3 model for roleplay & assistant tasks
DeepSeek-V3.2-Exp
Experimental version with sparse-attention architecture for long-context efficiency
DeepSeek-V3-0324
High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation
DeepSeek-V3.1
Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent
DeepSeek-R1-0528
An open-source next-generation reasoning-optimized language model with enhanced logic, math, and code performance, larger context handling, and improved function-calling capabilities
Manta-Mini-1.0
Lightweight tier optimized for speed and cost. (Free)
Manta-Flash-1.0
Balanced tier for general-purpose roleplay, good tradeoff between cost and latency. (Mid-tier)
Manta-Pro-1.0
Available for Creator Tier+
Premium tier with advanced reasoning and long context, best for complex roleplay.(Top-tier)
DeepSeek-V4-Flash-0731
DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding performance, at 284B total / 13B active parameters and a 1M-token context window
Claude-Opus-5
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Claude-Opus-5-Enterprise
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Gemini-3.6-Flash
Available for Creator Tier+
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.6-Flash-Enterprise
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.5-Flash-Lite
Available for Creator Tier+
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Gemini-3.5-Flash-Lite-Enterprise
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Qwen3.8-Max-Preview
Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data analysis, and office workflows; always uses thinking mode and supports long-horizon tool workflows and structured output
Kimi-K3
Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning
GPT-5.6-Sol
OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding
GPT-5.6-Terra
OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers
Grok-4.5
xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input
Claude-Sonnet-5
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
GLM-5.2
Z.ai's open-source MoE model for long-context reasoning and agentic software engineering, with a 1M-token context window
Kimi-K2.7-Code
Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering
Claude-Fable-5
Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity
MiniMax-M3
A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling
Qwen3.6-35B-A3B
Qwen3.6 35B (3B-active) MoE with built-in NEXTN speculative decoding, reasoning and tool calling — fast low-cost generation
Claude-Opus-4.8
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Gemini-3.5-Flash
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Grok-4.3
xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input
GPT-5.5
OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks
DeepSeek-V4-Pro
DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic performance with a 1M context window
DeepSeek-V4-Flash
DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M context window
MiMo-V2.5-Pro
Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon tasks with a 1M context window
MiMo-V2.5
Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning
Kimi-K2.6
An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context window, native vision input, and scaling up to 300 parallel sub-agents
Claude-Opus-4.7
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
GLM-5.1
Z.ai's open-source flagship MoE model built for agentic engineering, delivering state-of-the-art coding and long-horizon task performance that improves the longer it runs
Qwen3.6-Plus
Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks
Gemma-4-31B
Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows
MiMo-V2-Omni
Xiaomi's omni-modal agent model that natively understands text, images, video, and audio in a unified architecture, designed for complex real-world multimodal tasks
Mistral-Small-4
A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models
Gemini-3.1-Flash-Lite-Preview
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
MiMo-V2-Pro
A trillion-parameter text-only flagship agent model from Xiaomi, built for complex coding, long-horizon planning, and multi-step tool use with a 1M context window
MiniMax-M2.7
A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity tasks like coding, tool use, and multi-step reasoning — at a highly competitive price point
GLM-4.7-Flash
Available for Creator Tier+
A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks
Qwen3.5-Plus
Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing
MiniMax-M2.5
MiniMax-M2.5 is a cost-efficient frontier model that achieves SOTA in coding, agentic tool use, and search tasks through extensive reinforcement learning in real-world environments
GLM-5
Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding, and agentic tasks
Kimi-K2.5
An open-source native multimodal AI model that combines vision and language understanding with advanced agent-oriented capabilities, including self-directed agent swarm and visual coding, for versatile reasoning and tool use
Claude-Sonnet-4.6
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Opus-4.6
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Claude-Opus-4.5
An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads
Claude-Haiku-4.5
A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications
Claude-Sonnet-4.5
A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks
Claude-Opus-4
Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks
DeepSeek-V3.2
A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and general chat capabilities with efficient inference
MiniMax-M2.1
A lightweight, open-source text-to-text large language model optimized for coding, agent-oriented workflows, and instruction following, delivering faster, cleaner outputs and strong reasoning and problem-solving performance compared to its predecessor
GLM-4.7
An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents
GLM-4.7-Enterprise
An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents
Qwen3-235B-A22B-Instruct-2507
A 235B-parameter MoE instruction-tuned language model for general, multilingual, and coding tasks
Kimi-K2-Thinking
1T-parameter reasoning-focused language model designed for complex, multi-step problem solving and tool use
GLM-4.6
Flagship model from Z.ai with 200K token context, superior reasoning, advanced tool use, and coding excellence
Gemini-3-Flash-Preview
Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed
Gemini-2.5-Pro
Strongest Gemini, 1M context, great for code & knowledge
Gemini-2.5-Flash
Fast reasoning with improved latency & accuracy
Gemini-2.5-Flash-Lite
Ultra-low latency Gemini variant
GPT-5.4
A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding, and tool-use capabilities with very long context support
GPT-5.3-Chat
A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality dialogue, instruction following, and everyday chat applications
GPT-5.3-Codex
A high-performance agentic coding model designed to autonomously write, debug, and manage complex software development tasks while collaborating interactively with developers
GPT-5.2
OpenAI’s flagship multimodal (text + image) GPT-5.2 model for top-tier coding and agentic tasks, producing text output with the highest reasoning capability in the GPT-5 family
GPT-5.1
OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs
GPT-5
OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs
GPT-5-Mini
OpenAI’s faster, more cost-efficient GPT-5 model for well-defined tasks, supporting text and image input with text output
GPT-5-Nano
OpenAI’s fastest, most cost-efficient GPT-5 model for lightweight tasks like summarization and classification, supporting text and image input with text output
GPT-4o-mini
Compact GPT-4 Omni, supports text & image inputs
Cydonia-24B-v4.3
A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence
Cydonia-24B-v4.1
A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence
L3.3-70B-Euryale-v2.3
70B model focused on creative roleplay from Sao10k
L3.1-70B-Euryale-v2.2
70B model focused on creative roleplay from Sao10k
Sapphira-L3.3-70B-0.1
70B parameter language model optimized for storytelling, dialogue, and immersive role-play
MN-Violet-Lotus-12B
12B parameter model for creative, emotionally intelligent conversations
L3.3-MS-Nevoria-70B
70B narrative generation, LoRA tuning & story-rich outputs
Mistral-Nemo-Instruct-2407
24B instruction model, long context, fewer errors
Mistral-Small-3.2-24B-Instruct-2506
24B instruction model, long context, fewer errors
L3-70B-Euryale-v2.1
70B tuned for storytelling, roleplay, creative writing
L3-8B-Stheno-v3.2
8B LLaMA-3 model for roleplay & assistant tasks
DeepSeek-V3.2-Exp
Experimental version with sparse-attention architecture for long-context efficiency
DeepSeek-V3-0324
High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation
DeepSeek-V3-0324-Enterprise
High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation
DeepSeek-V3.1
Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent
DeepSeek-V3.1-Enterprise
Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent
DeepSeek-R1-0528
An open-source next-generation reasoning-optimized language model with enhanced logic, math, and code performance, larger context handling, and improved function-calling capabilities
Llama3.3-70B
Multilingual 70B delivering 405B-level performance
Claude-Sonnet-5-Enterprise
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Opus-4.7-Enterprise
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Claude-Opus-4.6-Enterprise
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Claude-Opus-4.5-Enterprise
An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads
Claude-Opus-4-Enterprise
Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks
Claude-Sonnet-4.6-Enterprise
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Sonnet-4.5-Enterprise
A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks
Claude-Haiku-4.5-Enterprise
A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications
Gemini-3.5-Flash-Enterprise
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite-Enterprise
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Gemini-3.1-Flash-Lite-Preview-Enterprise
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview-Enterprise
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
Gemini-3-Flash-Preview-Enterprise
Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed
Gemini-2.5-Pro-Enterprise
Strongest Gemini, 1M context, great for code & knowledge
Gemini-2.5-Flash-Enterprise
Fast reasoning with improved latency & accuracy
Gemini-2.5-Flash-Lite-Enterprise
Ultra-low latency Gemini variant
DeepSeek-V4-Flash-0731
DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding performance, at 284B total / 13B active parameters and a 1M-token context window
Claude-Opus-5
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Gemini-3.6-Flash
Available for Creator Tier+
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.6-Flash-Enterprise
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.5-Flash-Lite
Available for Creator Tier+
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Gemini-3.5-Flash-Lite-Enterprise
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Qwen3.8-Max-Preview
Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data analysis, and office workflows; always uses thinking mode and supports long-horizon tool workflows and structured output
Kimi-K3
Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning
Seedance-2.0
BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video
Seedance-2.0-Fast
BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens
Seedance-2.0-Mini
BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens
GPT-5.6-Sol
OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding
GPT-5.6-Terra
OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers
Grok-4.5
xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input
Claude-Sonnet-5
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
GLM-5.2
Z.ai's open-source MoE model for long-context reasoning and agentic software engineering, with a 1M-token context window
Kimi-K2.7-Code
Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering
Claude-Fable-5
Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity
MiniMax-M3
A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling
Claude-Opus-4.8
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Gemini-3.5-Flash
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Grok-4.3
xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input
GPT-5.5
OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks
DeepSeek-V4-Pro
DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic performance with a 1M context window
DeepSeek-V4-Flash
DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M context window
MiMo-V2.5-Pro
Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon tasks with a 1M context window
MiMo-V2.5
Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning
Kimi-K2.6
An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context window, native vision input, and scaling up to 300 parallel sub-agents
Claude-Opus-4.7
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
GLM-5.1
Z.ai's open-source flagship MoE model built for agentic engineering, delivering state-of-the-art coding and long-horizon task performance that improves the longer it runs
Qwen3.6-Plus
Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks
Gemma-4-31B
Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows
Mistral-Small-4
A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models
Gemini-3.1-Flash-Lite-Preview
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
MiniMax-M2.7
A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity tasks like coding, tool use, and multi-step reasoning — at a highly competitive price point
GLM-4.7-Flash
Available for Creator Tier+
A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks
Qwen3.5-Plus
Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing
MiniMax-M2.5
MiniMax-M2.5 is a cost-efficient frontier model that achieves SOTA in coding, agentic tool use, and search tasks through extensive reinforcement learning in real-world environments
GLM-5
Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding, and agentic tasks
Kimi-K2.5
An open-source native multimodal AI model that combines vision and language understanding with advanced agent-oriented capabilities, including self-directed agent swarm and visual coding, for versatile reasoning and tool use
Claude-Opus-4.6
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
DeepSeek-V3.2
A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and general chat capabilities with efficient inference
GLM-4.7
An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents
DeepSeek-V3-0324-Enterprise
High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation
DeepSeek-V3.1-Enterprise
Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent
Nano-Banana-2
Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.
Alibaba-Z-Image-Turbo
Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.
Bytedance-Seedream-5.0-Lite
Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.
Bytedance-Seedream-4.5
The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals
Alibaba-Wan-2.6-T2V
Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan-2.6-I2V
Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan2.6-I2V-Flash
Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing
Wan2.6-R2V
Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)
Alibaba-Wan-2.6-R2V-Flash
Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.
Qwen3-VL-Plus
Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual reasoning and agent capabilities
Qwen3-VL-Flash
Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on text, image, and video tasks with integrated thinking and non-thinking modes
Web Search
Real-time managed natural-language web search that returns ranked results, optionally enriched with full extracted page text in a single call—includes 50 free queries per user per day, with additional queries priced at $0.002 each
Claude-Opus-4.7-Enterprise
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Claude-Opus-4.6-Enterprise
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Gemini-3.5-Flash-Enterprise
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite-Enterprise
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Gemini-3.1-Flash-Lite-Preview-Enterprise
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview-Enterprise
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
Gemini-3.1-Flash-Live
Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket
DeepSeek-V4-Flash-0731
DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding performance, at 284B total / 13B active parameters and a 1M-token context window
DeepSeek-V4-Pro
DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic performance with a 1M context window
DeepSeek-V4-Flash
DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M context window
DeepSeek-V3.2
A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and general chat capabilities with efficient inference
DeepSeek-V3.2-Exp
Experimental version with sparse-attention architecture for long-context efficiency
DeepSeek-V3-0324
High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation
DeepSeek-V3-0324-Enterprise
High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation
DeepSeek-V3.1
Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent
DeepSeek-V3.1-Enterprise
Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent
DeepSeek-R1-0528
An open-source next-generation reasoning-optimized language model with enhanced logic, math, and code performance, larger context handling, and improved function-calling capabilities
Claude-Opus-5
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Claude-Opus-5-Enterprise
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Gemini-3.6-Flash
Available for Creator Tier+
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.6-Flash-Enterprise
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.5-Flash-Lite
Available for Creator Tier+
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Gemini-3.5-Flash-Lite-Enterprise
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Kimi-K3
Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning
GPT-5.6-Sol
OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding
GPT-5.6-Terra
OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers
Grok-4.5
xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input
Claude-Sonnet-5
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Kimi-K2.7-Code
Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering
Claude-Fable-5
Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity
MiniMax-M3
A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling
Claude-Opus-4.8
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Gemini-3.5-Flash
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Grok-4.3
xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input
GPT-5.5
OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks
MiMo-V2.5
Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning
Kimi-K2.6
An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context window, native vision input, and scaling up to 300 parallel sub-agents
Claude-Opus-4.7
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Qwen3.6-Plus
Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks
Gemma-4-31B
Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows
MiMo-V2-Omni
Xiaomi's omni-modal agent model that natively understands text, images, video, and audio in a unified architecture, designed for complex real-world multimodal tasks
Mistral-Small-4
A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models
Gemini-3.1-Flash-Lite-Preview
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
Qwen3.5-Plus
Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing
Kimi-K2.5
An open-source native multimodal AI model that combines vision and language understanding with advanced agent-oriented capabilities, including self-directed agent swarm and visual coding, for versatile reasoning and tool use
Claude-Sonnet-4.6
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Opus-4.6
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Claude-Opus-4.5
An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads
Claude-Haiku-4.5
A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications
Claude-Sonnet-4.5
A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks
Claude-Opus-4
Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks
Gemini-3-Flash-Preview
Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed
Gemini-2.5-Pro
Strongest Gemini, 1M context, great for code & knowledge
Gemini-2.5-Flash
Fast reasoning with improved latency & accuracy
Gemini-2.5-Flash-Lite
Ultra-low latency Gemini variant
GPT-5.4
A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding, and tool-use capabilities with very long context support
GPT-5.3-Chat
A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality dialogue, instruction following, and everyday chat applications
GPT-5.3-Codex
A high-performance agentic coding model designed to autonomously write, debug, and manage complex software development tasks while collaborating interactively with developers
GPT-5.2
OpenAI’s flagship multimodal (text + image) GPT-5.2 model for top-tier coding and agentic tasks, producing text output with the highest reasoning capability in the GPT-5 family
GPT-5.1
OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs
GPT-5
OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs
GPT-5-Mini
OpenAI’s faster, more cost-efficient GPT-5 model for well-defined tasks, supporting text and image input with text output
GPT-5-Nano
OpenAI’s fastest, most cost-efficient GPT-5 model for lightweight tasks like summarization and classification, supporting text and image input with text output
GPT-4o-mini
Compact GPT-4 Omni, supports text & image inputs
Claude-Sonnet-5-Enterprise
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Opus-4.7-Enterprise
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Claude-Opus-4.6-Enterprise
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Claude-Opus-4.5-Enterprise
An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads
Claude-Opus-4-Enterprise
Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks
Claude-Sonnet-4.6-Enterprise
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Sonnet-4.5-Enterprise
A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks
Claude-Haiku-4.5-Enterprise
A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications
Gemini-3.5-Flash-Enterprise
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite-Enterprise
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Gemini-3.1-Flash-Lite-Preview-Enterprise
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview-Enterprise
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
Gemini-3-Flash-Preview-Enterprise
Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed
Gemini-2.5-Pro-Enterprise
Strongest Gemini, 1M context, great for code & knowledge
Gemini-2.5-Flash-Enterprise
Fast reasoning with improved latency & accuracy
Gemini-2.5-Flash-Lite-Enterprise
Ultra-low latency Gemini variant
Claude-Opus-5
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Claude-Opus-5-Enterprise
Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in deep reasoning and long-horizon autonomy over Claude Opus 4.8, with a 1M-token context window
Claude-Sonnet-5
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Fable-5
Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on adaptive thinking that adjusts reasoning depth based on task complexity
Claude-Opus-4.8
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Claude-Opus-4.7
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Claude-Sonnet-4.6
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Opus-4.6
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Claude-Opus-4.5
An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads
Claude-Haiku-4.5
A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications
Claude-Sonnet-4.5
A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks
Claude-Opus-4
Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks
Claude-Sonnet-5-Enterprise
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Opus-4.7-Enterprise
Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and complex professional work with adaptive thinking that adjusts reasoning depth based on task complexity
Claude-Opus-4.6-Enterprise
Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks
Claude-Opus-4.5-Enterprise
An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the most demanding enterprise and research workloads
Claude-Opus-4-Enterprise
Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and high-accuracy enterprise-grade tasks
Claude-Sonnet-4.6-Enterprise
Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use, agentic workflows, and long-context reasoning with a 1M-token context window
Claude-Sonnet-4.5-Enterprise
A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks
Claude-Haiku-4.5-Enterprise
A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications
Manta-Mini-1.0
Lightweight tier optimized for speed and cost. (Free)
Manta-Flash-1.0
Balanced tier for general-purpose roleplay, good tradeoff between cost and latency. (Mid-tier)
Manta-Pro-1.0
Available for Creator Tier+
Premium tier with advanced reasoning and long context, best for complex roleplay.(Top-tier)
Gemini-3.6-Flash
Available for Creator Tier+
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.5-Flash-Lite
Available for Creator Tier+
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
GLM-4.7-Flash
Available for Creator Tier+
A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks
Sapphira-L3.3-70B-0.1
70B parameter language model optimized for storytelling, dialogue, and immersive role-play
MN-Violet-Lotus-12B
12B parameter model for creative, emotionally intelligent conversations
L3.3-MS-Nevoria-70B
70B narrative generation, LoRA tuning & story-rich outputs
Mistral-Small-3.2-24B-Instruct-2506
24B instruction model, long context, fewer errors
L3-70B-Euryale-v2.1
70B tuned for storytelling, roleplay, creative writing
L3-8B-Stheno-v3.2
8B LLaMA-3 model for roleplay & assistant tasks
Gemini-3.6-Flash
Available for Creator Tier+
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.6-Flash-Enterprise
Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed and cost, with a 1M-token context window
Gemini-3.5-Flash-Lite
Available for Creator Tier+
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Gemini-3.5-Flash-Lite-Enterprise
Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume classification, extraction and translation workloads
Gemini-3.5-Flash
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Gemini-3.1-Flash-Lite-Preview
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
Gemini-3-Flash-Preview
Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed
Gemini-2.5-Pro
Strongest Gemini, 1M context, great for code & knowledge
Gemini-2.5-Flash
Fast reasoning with improved latency & accuracy
Gemini-2.5-Flash-Lite
Ultra-low latency Gemini variant
Nano-Banana-2
Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.
Nano-Banana-Pro-Edit
Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).
Nano-Banana-Pro
High-quality image generation with better text rendering, character consistency, and advanced composition. Powered by Gemini 3 Pro Image.
Nano-Banana-Edit
AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.
Nano-Banana
Gemini 2.5 Flash Image - High-quality image generation with multiple aspect ratios
Gemini-2.5-Flash-TTS
A low-latency, cost-efficient text-to-speech model that converts text into expressive single or multi-speaker audio with controllable tone, style, and pacing
Gemini-2.5-Pro-TTS
A high-quality text-to-speech model designed to generate natural, expressive audio with fine-grained control over tone, pacing, and emotion for narration, assistants, and conversational voice applications
Gemini-3.5-Flash-Enterprise
Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context
Gemini-3.1-Flash-Lite-Enterprise
Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume, latency-sensitive tasks like translation, classification, and simple data extraction with support for text, image, video, audio, and PDF inputs
Gemini-3.1-Flash-Lite-Preview-Enterprise
Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable real-time applications
Gemini-3.1-Pro-Preview-Enterprise
Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities, improved token efficiency, and stronger performance in software engineering and agentic workflows, featuring a 1M-token context window
Gemini-3-Flash-Preview-Enterprise
Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more balanced cost and speed
Gemini-2.5-Pro-Enterprise
Strongest Gemini, 1M context, great for code & knowledge
Gemini-2.5-Flash-Enterprise
Fast reasoning with improved latency & accuracy
Gemini-2.5-Flash-Lite-Enterprise
Ultra-low latency Gemini variant
Gemini-3.1-Flash-Live
Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket
Qwen3.8-Max-Preview
Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data analysis, and office workflows; always uses thinking mode and supports long-horizon tool workflows and structured output
Qwen3.6-35B-A3B
Qwen3.6 35B (3B-active) MoE with built-in NEXTN speculative decoding, reasoning and tool calling — fast low-cost generation
Qwen3.6-Plus
Alibaba's flagship hosted model with a 1M-token context window, excelling at agentic coding, multimodal reasoning, and long-horizon planning tasks
Qwen3.5-Plus
Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in tool use, offering frontier-class multimodal reasoning at competitive pricing
Qwen3-VL-Plus
Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual reasoning and agent capabilities
Qwen3-VL-Flash
Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on text, image, and video tasks with integrated thinking and non-thinking modes
Qwen2.5-VL-32B-Instruct
A 32B-parameter vision-language model tuned to follow instructions and reason over images and text
Qwen3-Embedding-8B
8B embedding model, strong multilingual & code support, top MTEB scorer
Kimi-K3
Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool use, with native vision input and always-on reasoning
Kimi-K2.7-Code
Moonshot AI's trillion-parameter open-source MoE coding model, built for long-horizon, high-efficiency agentic software engineering
Kimi-K2.6
An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context window, native vision input, and scaling up to 300 parallel sub-agents
Kimi-K2.5
An open-source native multimodal AI model that combines vision and language understanding with advanced agent-oriented capabilities, including self-directed agent swarm and visual coding, for versatile reasoning and tool use
Kimi-K2-Thinking
1T-parameter reasoning-focused language model designed for complex, multi-step problem solving and tool use
Seedance-2.0
BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video
Seedance-2.0-Fast
BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens
Seedance-2.0-Mini
BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens
Alibaba-Wan-2.6-T2V
Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.
Veo-3.1-Fast
Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p. Optional AI-generated audio (+50% cost).
Veo-3.1
Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds at 720p/1080p.
Seedance-1.0-Pro-Text-to-Video
Pro-tier text-to-video, generates multi-shot 1080p videos with narrative flow
Seedance-1.0-Lite-Text-to-Video
Lite text-to-video, faster, lower compute, shorter outputs
Seedance-2.0
BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video
Seedance-2.0-Fast
BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens
Seedance-2.0-Mini
BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens
Alibaba-Wan-2.6-I2V
Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan2.6-I2V-Flash
Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing
Veo-3.1-Fast-I2V
Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at 720p/1080p.
Veo-3.1-I2V
Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.
Seedance-1.0-Pro-Image-to-Video
Pro-tier model, turns images into cinematic 1080p videos with smooth motion
Seedance-1.0-Lite-Image-to-Video
Lite version, quick image-to-video at up to 1080p, shorter sequences
Seedance-2.0
BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video
Seedance-2.0-Fast
BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens
Seedance-2.0-Mini
BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens
Seedance-2.0
BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video
Seedance-2.0-Fast
BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens
Seedance-2.0-Mini
BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens
Alibaba-Wan-2.6-T2V
Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan-2.6-I2V
Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan2.6-I2V-Flash
Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing
Wan2.6-R2V
Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)
Alibaba-Wan-2.6-R2V-Flash
Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.
Veo-3.1-Fast
Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p. Optional AI-generated audio (+50% cost).
Veo-3.1-Fast-I2V
Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at 720p/1080p.
Veo-3.1
Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds at 720p/1080p.
Veo-3.1-I2V
Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.
Seedance-1.0-Pro-Image-to-Video
Pro-tier model, turns images into cinematic 1080p videos with smooth motion
Seedance-1.0-Pro-Text-to-Video
Pro-tier text-to-video, generates multi-shot 1080p videos with narrative flow
Seedance-1.0-Lite-Image-to-Video
Lite version, quick image-to-video at up to 1080p, shorter sequences
Seedance-1.0-Lite-Text-to-Video
Lite text-to-video, faster, lower compute, shorter outputs
Seedance-2.0
BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on actual tokens; rate varies by resolution & reference video
Seedance-2.0-Fast
BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens
Seedance-2.0-Mini
BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens
Bytedance-Seedream-5.0-Lite
Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.
Bytedance-Seedream-4.5
The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals
Bytedance-Seedream-4.0
Next-generation image creation model that unifies image generation and editing to handle complex multimodal tasks with faster inference and produce high-definition images up to 4K
Bytedance-Seedream-3.0
Bilingual text-to-image, 2K resolution, accurate text & artistic layouts
Seedance-1.0-Pro-Image-to-Video
Pro-tier model, turns images into cinematic 1080p videos with smooth motion
Seedance-1.0-Pro-Text-to-Video
Pro-tier text-to-video, generates multi-shot 1080p videos with narrative flow
Seedance-1.0-Lite-Image-to-Video
Lite version, quick image-to-video at up to 1080p, shorter sequences
Seedance-1.0-Lite-Text-to-Video
Lite text-to-video, faster, lower compute, shorter outputs
GPT-5.6-Sol
OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic coding
GPT-5.6-Terra
OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned between the flagship Sol and cost-efficient Luna tiers
GPT-5.5
OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding, and agentic tasks
GPT-5.4
A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding, and tool-use capabilities with very long context support
GPT-5.3-Chat
A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality dialogue, instruction following, and everyday chat applications
GPT-5.3-Codex
A high-performance agentic coding model designed to autonomously write, debug, and manage complex software development tasks while collaborating interactively with developers
GPT-5.2
OpenAI’s flagship multimodal (text + image) GPT-5.2 model for top-tier coding and agentic tasks, producing text output with the highest reasoning capability in the GPT-5 family
GPT-5.1
OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs
GPT-5
OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across domains, producing text outputs
GPT-5-Mini
OpenAI’s faster, more cost-efficient GPT-5 model for well-defined tasks, supporting text and image input with text output
GPT-5-Nano
OpenAI’s fastest, most cost-efficient GPT-5 model for lightweight tasks like summarization and classification, supporting text and image input with text output
GPT-4o-mini
Compact GPT-4 Omni, supports text & image inputs
Grok-4.5
xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM. 500K-token context window and multimodal text/image input
Grok-4.3
xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context window, and multimodal text/image input
GLM-5.2
Z.ai's open-source MoE model for long-context reasoning and agentic software engineering, with a 1M-token context window
GLM-5.1
Z.ai's open-source flagship MoE model built for agentic engineering, delivering state-of-the-art coding and long-horizon task performance that improves the longer it runs
GLM-4.7-Flash
Available for Creator Tier+
A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and long-context reasoning tasks
GLM-5
Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding, and agentic tasks
GLM-4.7
An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents
GLM-4.7-Enterprise
An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents
GLM-4.6
Flagship model from Z.ai with 200K token context, superior reasoning, advanced tool use, and coding excellence
MiniMax-M3
A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native text/image/video understanding and tool calling
MiniMax-M2.7
A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity tasks like coding, tool use, and multi-step reasoning — at a highly competitive price point
MiniMax-M2.5
MiniMax-M2.5 is a cost-efficient frontier model that achieves SOTA in coding, agentic tool use, and search tasks through extensive reinforcement learning in real-world environments
MiniMax-M2.1
A lightweight, open-source text-to-text large language model optimized for coding, agent-oriented workflows, and instruction following, delivering faster, cleaner outputs and strong reasoning and problem-solving performance compared to its predecessor
MiMo-V2.5-Pro
Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon tasks with a 1M context window
MiMo-V2.5
Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for agentic and long-horizon reasoning
MiMo-V2-Omni
Xiaomi's omni-modal agent model that natively understands text, images, video, and audio in a unified architecture, designed for complex real-world multimodal tasks
MiMo-V2-Pro
A trillion-parameter text-only flagship agent model from Xiaomi, built for complex coding, long-horizon planning, and multi-step tool use with a 1M context window
Gemma-4-31B
Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and agentic workflows
Mistral-Small-4
A large (~119B parameter) dense language model designed for efficient, high-quality text generation and reasoning with a focus on strong performance-to-cost efficiency compared to similarly sized models
Mistral-Nemo-Instruct-2407
24B instruction model, long context, fewer errors
Mistral-Small-3.2-24B-Instruct-2506
24B instruction model, long context, fewer errors
Qwen3-235B-A22B-Instruct-2507
A 235B-parameter MoE instruction-tuned language model for general, multilingual, and coding tasks
Cydonia-24B-v4.3
A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence
Cydonia-24B-v4.1
A 24B-parameter uncensored instruction-tuned model focused on creative writing with strong recall and prompt adherence
L3.3-70B-Euryale-v2.3
70B model focused on creative roleplay from Sao10k
L3.1-70B-Euryale-v2.2
70B model focused on creative roleplay from Sao10k
Sapphira-L3.3-70B-0.1
70B parameter language model optimized for storytelling, dialogue, and immersive role-play
MN-Violet-Lotus-12B
12B parameter model for creative, emotionally intelligent conversations
L3.3-MS-Nevoria-70B
70B narrative generation, LoRA tuning & story-rich outputs
L3-70B-Euryale-v2.1
70B tuned for storytelling, roleplay, creative writing
L3-8B-Stheno-v3.2
8B LLaMA-3 model for roleplay & assistant tasks
Manta-Mini-1.0
Lightweight tier optimized for speed and cost. (Free)
Manta-Flash-1.0
Balanced tier for general-purpose roleplay, good tradeoff between cost and latency. (Mid-tier)
Manta-Pro-1.0
Available for Creator Tier+
Premium tier with advanced reasoning and long context, best for complex roleplay.(Top-tier)
Llama3.3-70B
Multilingual 70B delivering 405B-level performance
Nano-Banana-2
Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.
Alibaba-Z-Image-Turbo
Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.
Nano-Banana-Pro-Edit
Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).
Nano-Banana-Pro
High-quality image generation with better text rendering, character consistency, and advanced composition. Powered by Gemini 3 Pro Image.
Nano-Banana-Edit
AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.
Nano-Banana
Gemini 2.5 Flash Image - High-quality image generation with multiple aspect ratios
DreamShaper-8
A Stable Diffusion v1.5-based text-to-image model, balancing photorealistic and stylized (including anime) generation with strong LoRA compatibility
Bytedance-Seedream-5.0-Lite
Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.
Bytedance-Seedream-4.5
The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals
Bytedance-Seedream-4.0
Next-generation image creation model that unifies image generation and editing to handle complex multimodal tasks with faster inference and produce high-definition images up to 4K
Bytedance-Seedream-3.0
Bilingual text-to-image, 2K resolution, accurate text & artistic layouts
Nano-Banana-2
Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.
Nano-Banana-2
Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect_ratio, image_size, person_generation, output_mime_type, compression_quality, thinking_level.
Alibaba-Z-Image-Turbo
Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.
Nano-Banana-Pro-Edit
Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).
Nano-Banana-Pro
High-quality image generation with better text rendering, character consistency, and advanced composition. Powered by Gemini 3 Pro Image.
Nano-Banana-Edit
AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.
Nano-Banana
Gemini 2.5 Flash Image - High-quality image generation with multiple aspect ratios
DreamShaper-8
A Stable Diffusion v1.5-based text-to-image model, balancing photorealistic and stylized (including anime) generation with strong LoRA compatibility
Bytedance-Seedream-5.0-Lite
Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.
Bytedance-Seedream-4.5
The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals
Bytedance-Seedream-4.0
Next-generation image creation model that unifies image generation and editing to handle complex multimodal tasks with faster inference and produce high-definition images up to 4K
Bytedance-Seedream-3.0
Bilingual text-to-image, 2K resolution, accurate text & artistic layouts
Alibaba-Z-Image-Turbo
Alibaba Z-Image Turbo - Fast text-to-image with Chinese/English text rendering support.
Nano-Banana-Pro-Edit
Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality, supports multiple resolution outputs (1K/2K/4K).
Nano-Banana-Edit
AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove backgrounds, change colors, apply styles, and more.
Bytedance-Seedream-5.0-Lite
Advanced text-to-image generation with web search, deep reasoning and smart editing capabilities. Supports up to 3K resolution and seed parameters.
Bytedance-Seedream-4.5
The latest Bytedance image model, delivering better editing consistency, improved multi-image fusion, finer detail control, natural small text and faces, and harmonious, aesthetic visuals
Bytedance-Seedream-4.0
Next-generation image creation model that unifies image generation and editing to handle complex multimodal tasks with faster inference and produce high-definition images up to 4K
DreamShaper-8
A Stable Diffusion v1.5-based text-to-image model, balancing photorealistic and stylized (including anime) generation with strong LoRA compatibility
Web Search
Real-time managed natural-language web search that returns ranked results, optionally enriched with full extracted page text in a single call—includes 50 free queries per user per day, with additional queries priced at $0.002 each
Alibaba-Wan-2.6-T2V
Alibaba Wan 2.6 Flash - Fast text-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan-2.6-I2V
Alibaba Wan 2.6 I2V - Image-to-video generation. Supports 2-15 seconds at 720p/1080p.
Alibaba-Wan2.6-I2V-Flash
Alibaba Wan 2.6 I2V Flash - Fast lightweight I2V with 480P/720P/1080P, audio/silent pricing
Wan2.6-R2V
Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)
Alibaba-Wan-2.6-R2V-Flash
Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.
Wan2.6-R2V
Alibaba Wan 2.6 R2V - Mixed image+video references (up to 5 combined)
Alibaba-Wan-2.6-R2V-Flash
Alibaba Wan 2.6 R2V Flash - Reference-to-video generation from 1-3 reference videos. Supports 5 or 10 seconds.
Veo-3.1-Fast
Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p. Optional AI-generated audio (+50% cost).
Veo-3.1-Fast-I2V
Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at 720p/1080p.
Veo-3.1
Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds at 720p/1080p.
Veo-3.1-I2V
Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.
Qwen3-VL-Plus
Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual reasoning and agent capabilities
Qwen3-VL-Flash
Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on text, image, and video tasks with integrated thinking and non-thinking modes
Qwen2.5-VL-32B-Instruct
A 32B-parameter vision-language model tuned to follow instructions and reason over images and text
Gemini-2.5-Flash-TTS
A low-latency, cost-efficient text-to-speech model that converts text into expressive single or multi-speaker audio with controllable tone, style, and pacing
Gemini-2.5-Pro-TTS
A high-quality text-to-speech model designed to generate natural, expressive audio with fine-grained control over tone, pacing, and emotion for narration, assistants, and conversational voice applications
Faster-Whisper-Large-V3
High-performance speech-to-text model optimized for fast, accurate transcription using an efficient Whisper-based architecture
Gemini-3.1-Flash-Live
Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket
Gemini-2.5-Flash-TTS
A low-latency, cost-efficient text-to-speech model that converts text into expressive single or multi-speaker audio with controllable tone, style, and pacing
Gemini-2.5-Pro-TTS
A high-quality text-to-speech model designed to generate natural, expressive audio with fine-grained control over tone, pacing, and emotion for narration, assistants, and conversational voice applications
Faster-Whisper-Large-V3
High-performance speech-to-text model optimized for fast, accurate transcription using an efficient Whisper-based architecture
Faster-Whisper-Large-V3
High-performance speech-to-text model optimized for fast, accurate transcription using an efficient Whisper-based architecture
Qwen3-Embedding-8B
8B embedding model, strong multilingual & code support, top MTEB scorer
BGE-reranker-v2-m3
Multilingual reranker, query+passage → relevance score, lightweight & fast
Web Search
Real-time managed natural-language web search that returns ranked results, optionally enriched with full extracted page text in a single call—includes 50 free queries per user per day, with additional queries priced at $0.002 each
Gemini-3.1-Flash-Live
Gemini Live speech-to-speech model for real-time voice conversations — bidirectional audio streaming over the OpenAI-compatible /v1/realtime WebSocket





