Infer

88 models

Claude Fable 5

Anthropic's high-capability Fable 5 model for advanced reasoning, repository-scale coding, long-context analysis, and agentic workflows.

Anthropic·1M context·700 Credits / 1M in·3,500 Credits / 1M out·70 Credits / 1M cache read·875 Credits / 1M cache write

Claude Haiku 4.5

Default Claude Haiku route for fast background work, summaries, and lightweight agent loops.

Anthropic·1M context·70 Credits / 1M in·350 Credits / 1M out·7 Credits / 1M cache read·87.5 Credits / 1M cache write

Claude Opus 4.6

Default Claude Opus route for complex planning and high-quality reasoning.

Anthropic·1M context·350 Credits / 1M in·1,750 Credits / 1M out·35 Credits / 1M cache read·437.5 Credits / 1M cache write

Claude Opus 5

Default Claude Opus 5 route for the strongest planning, coding, and long-running agent work.

Anthropic·1M context·350 Credits / 1M in·1,750 Credits / 1M out·35 Credits / 1M cache read·437.5 Credits / 1M cache write

Claude Sonnet 4.6

Default Claude Sonnet route for balanced coding, analysis, and agent traffic.

Anthropic·1M context·210 Credits / 1M in·1,050 Credits / 1M out·21 Credits / 1M cache read·262.5 Credits / 1M cache write

DeepSeek Chat V3 0324

DeepSeek V3 chat snapshot for dependable coding, math, and general reasoning.

DeepSeek·164K context·20 Credits / 1M in·77 Credits / 1M out·13.5 Credits / 1M cache read·20 Credits / 1M cache write

DeepSeek Chat V3.1

DeepSeek V3.1 chat model for balanced quality, cost, and coding support.

DeepSeek·33K context·15 Credits / 1M in·75 Credits / 1M out·15 Credits / 1M cache read·15 Credits / 1M cache write

DeepSeek V3.1 Terminus

DeepSeek V3.1 Terminus model for stronger chat and technical reasoning workloads.

DeepSeek·164K context·21 Credits / 1M in·79 Credits / 1M out·13 Credits / 1M cache read·21 Credits / 1M cache write

DeepSeek V3.2

DeepSeek V3.2 served through Qianfan International for high-ROI chat, coding, math, and reasoning workloads.

DeepSeek·144K context·35 Credits / 1M in·105 Credits / 1M out·30 Credits / 1M cache read

DeepSeek V3.2 Exp

Experimental DeepSeek V3.2 model for previewing newer reasoning and chat behavior.

DeepSeek·164K context·27 Credits / 1M in·41 Credits / 1M out·27 Credits / 1M cache read·27 Credits / 1M cache write

DeepSeek V4 Flash

Lower-latency DeepSeek V4 route for cost-sensitive chat and structured generation.

DeepSeek·1M context·Official time-window pricing (UTC):Off-peak:22 Credits / 1M in·66 Credits / 1M out·0.7 Credits / 1M cache read·22 Credits / 1M cache write·Peak:44 Credits / 1M in·132 Credits / 1M out·1.4 Credits / 1M cache read·44 Credits / 1M cache write

DeepSeek V4 Pro

Infer DeepSeek Pro route for coding, math, and complex analysis with official DeepSeek primary routing and provider fallbacks.

DeepSeek·1M context·Official time-window pricing (UTC):Off-peak:66 Credits / 1M in·198 Credits / 1M out·2.2 Credits / 1M cache read·66 Credits / 1M cache write·Peak:132 Credits / 1M in·396 Credits / 1M out·4.4 Credits / 1M cache read·132 Credits / 1M cache write

ERNIE 5.0

Baidu's flagship Qianfan International model for text chat plus verified vision input with image URLs, Base64 images, multi-turn image conversations, and detail controls.

Baidu·128K context·98 Credits / 1M in·392 Credits / 1M out

Gemini 2.5 Flash

Discounted Gemini Flash route for fast multimodal-capable assistant traffic.

Google·1M context·7.5 Credits / 1M in·62.55 Credits / 1M out·1.875 Credits / 1M cache read

Gemini 2.5 Flash Image

Gemini image-capable route for visual prompts and image-token billing.

Google·1M context·7.5 Credits / 1M in·750 Credits / 1M out

Gemini 3 Pro Image Preview

Higher-fidelity Gemini image-generation route for richer visual outputs and prompt-following.

Google·66K context·80 Credits / 1M in·4,800 Credits / 1M out

Gemini 3.1 Flash Image Preview

Low-latency Gemini 3.1 image-preview route for quick visual generation with discounted output-image pricing.

Google·66K context·20 Credits / 1M in·600 Credits / 1M out

Gemma 3 12B IT

Small Gemma 3 instruction model for very low-cost chat and classification tasks.

Google·131K context·4 Credits / 1M in·13 Credits / 1M out·4 Credits / 1M cache read·4 Credits / 1M cache write

Gemma 3 27B IT

Instruction-tuned Gemma 3 model for lightweight assistants and text generation.

Google·131K context·8 Credits / 1M in·16 Credits / 1M out·8 Credits / 1M cache read·8 Credits / 1M cache write

Gemma 4 26B A4B IT

Compact instruction-tuned Gemma model for fast and low-cost production use.

Google·262K context·13 Credits / 1M in·40 Credits / 1M out·13 Credits / 1M cache read·13 Credits / 1M cache write

Gemma 4 31B IT

Instruction-tuned Gemma model for economical chat, coding support, and analysis.

Google·262K context·13 Credits / 1M in·38 Credits / 1M out·13 Credits / 1M cache read·13 Credits / 1M cache write

GLM-4.5 Air

Efficient GLM model for low-cost bilingual chat, summarization, and extraction.

Zhipu·131K context·13 Credits / 1M in·85 Credits / 1M out·2.5 Credits / 1M cache read·13 Credits / 1M cache write

GLM-4.6

Earlier GLM generation with dependable bilingual chat and instruction following.

Zhipu·205K context·39 Credits / 1M in·190 Credits / 1M out·39 Credits / 1M cache read·39 Credits / 1M cache write

GLM-4.7

Zhipu GLM model for balanced reasoning, coding support, and Chinese-English tasks.

Zhipu·203K context·40 Credits / 1M in·175 Credits / 1M out·8 Credits / 1M cache read·40 Credits / 1M cache write

GLM-4.7 Flash

Fast GLM flash model for inexpensive chat, classification, and structured output.

Zhipu·203K context·6 Credits / 1M in·40 Credits / 1M out·1 Credits / 1M cache read·6 Credits / 1M cache write

GLM-5

Zhipu AI's latest foundation model with strong bilingual capabilities and advanced reasoning for enterprise applications.

Zhipu·200K context·55 Credits / 1M in·176 Credits / 1M out·11 Credits / 1M cache read·55 Credits / 1M cache write

GLM-5.1

Newer Zhipu flagship with stronger bilingual reasoning and enterprise-oriented instruction following.

Zhipu·200K context·77 Credits / 1M in·242 Credits / 1M out·14.3 Credits / 1M cache read·77 Credits / 1M cache write

GLM-5.2

Z.ai flagship route for long-context bilingual reasoning and agent workloads.

Zhipu·1M context·98 Credits / 1M in·308 Credits / 1M out·18.2 Credits / 1M cache read·98 Credits / 1M cache write

GPT Image 2

OpenAI image-generation route for synchronous text-to-image requests over the standard Images API.

OpenAI·33K context·320 Credits / 1M in·1,200 Credits / 1M out

GPT OSS 120B

Large open-weight GPT OSS model for economical reasoning, coding, and general chat.

OpenAI·131K context·3.9 Credits / 1M in·19 Credits / 1M out·3.9 Credits / 1M cache read·3.9 Credits / 1M cache write

GPT OSS 20B

Smaller open-weight GPT OSS model for lightweight chat and high-volume routing.

OpenAI·131K context·3 Credits / 1M in·14 Credits / 1M out·3 Credits / 1M cache read·3 Credits / 1M cache write

GPT-5.3 Codex

Default Codex-oriented GPT route for developer agents and code tasks.

OpenAI·400K context·140 Credits / 1M in·1,120 Credits / 1M out·14 Credits / 1M cache read·175 Credits / 1M cache write

GPT-5.4

Default GPT-5.4 route for cost-sensitive production agents.

OpenAI·1M context·200 Credits / 1M in·1,200 Credits / 1M out·20 Credits / 1M cache read·250 Credits / 1M cache write

GPT-5.5

Frontier GPT route for demanding reasoning, tool use, and agentic workloads.

OpenAI·1M context·Tiered pricing:0-272K:350 Credits / 1M in·2,100 Credits / 1M out·35 Credits / 1M cache read·437.5 Credits / 1M cache write·272K-1M:700 Credits / 1M in·3,150 Credits / 1M out·70 Credits / 1M cache read·875 Credits / 1M cache write

GPT-5.6 Luna

Default GPT-5.6 Luna route for cost-sensitive agent traffic and high-volume reasoning workloads.

OpenAI·1M context·Tiered pricing:0-272K:14 Credits / 1M in·84 Credits / 1M out·1.4 Credits / 1M cache read·17.5 Credits / 1M cache write·272K-1M:28 Credits / 1M in·126 Credits / 1M out·2.8 Credits / 1M cache read·35 Credits / 1M cache write

GPT-5.6 Sol

High-capability GPT-5.6 route for complex coding, review, and agent planning.

OpenAI·1M context·280 Credits / 1M in·1,400 Credits / 1M out·28 Credits / 1M cache read·350 Credits / 1M cache write

GPT-5.6 Terra

Default GPT-5.6 Terra route for stronger reasoning, coding, and agentic workloads.

OpenAI·1M context·Tiered pricing:0-272K:140 Credits / 1M in·840 Credits / 1M out·14 Credits / 1M cache read·175 Credits / 1M cache write·272K-1M:280 Credits / 1M in·1,260 Credits / 1M out·28 Credits / 1M cache read·350 Credits / 1M cache write

GPT-6 Astra

Long-context GPT-6 Astra route for demanding reasoning, coding, and agentic workloads.

OpenAI·1M context·Tiered pricing:0-272K:700 Credits / 1M in·3,500 Credits / 1M out·70 Credits / 1M cache read·875 Credits / 1M cache write·272K-1M:1,400 Credits / 1M in·5,250 Credits / 1M out·140 Credits / 1M cache read·1,750 Credits / 1M cache write

Grok 4.1 Fast

Default Grok route for low-latency chat and production traffic.

xAI·2M context·13 Credits / 1M in·32.5 Credits / 1M out·3.25 Credits / 1M cache read

Kimi K2.5

Moonshot AI's powerful model with strong long-context understanding and reasoning. Excels at document analysis and complex tasks.

Moonshot·256K context·42 Credits / 1M in·210 Credits / 1M out·7 Credits / 1M cache read·42 Credits / 1M cache write

Kimi K2.6

Updated Moonshot model with stronger long-context reasoning and higher output quality for research and document-heavy workflows.

Moonshot·256K context·66.5 Credits / 1M in·280 Credits / 1M out·11.2 Credits / 1M cache read·66.5 Credits / 1M cache write

Kimi K3

Moonshot's flagship long-horizon coding and knowledge-work model with a 1M-token context window and always-on reasoning.

Moonshot·1.0M context·210 Credits / 1M in·1,050 Credits / 1M out·21 Credits / 1M cache read·210 Credits / 1M cache write

Kling v2.6

Kling async video route for text-to-video and image-to-video generation through the ClawOS media API.

Kuaishou·33K context·N/A / 1M in·N/A / 1M out

Kling v3

Kling v3 async video route for higher-quality text-to-video and image-to-video generation.

Kuaishou·33K context·N/A / 1M in·N/A / 1M out

Kling v3 Omni

Kling v3 omni async video route for prompt-guided video editing with optional first-frame image input.

Kuaishou·33K context·N/A / 1M in·N/A / 1M out

Llama 3.1 70B Instruct

Larger Llama instruct model for stronger open-weight reasoning and generation.

Meta·131K context·40 Credits / 1M in·40 Credits / 1M out·40 Credits / 1M cache read·40 Credits / 1M cache write

Llama 3.1 8B Instruct

Small Llama instruct model for inexpensive chat, classification, and extraction.

Meta·16K context·2 Credits / 1M in·5 Credits / 1M out·2 Credits / 1M cache read·2 Credits / 1M cache write

Llama 3.3 70B Instruct

Updated 70B Llama instruct model for general-purpose chat and reasoning.

Meta·66K context·10 Credits / 1M in·32 Credits / 1M out·10 Credits / 1M cache read·10 Credits / 1M cache write

Llama 4 Maverick

Llama 4 model for long-context open-weight chat and multimodal-adjacent workflows.

Meta·1M context·15 Credits / 1M in·60 Credits / 1M out·15 Credits / 1M cache read·15 Credits / 1M cache write

MiMo V2 Flash

Xiaomi MiMo flash model for high-throughput, latency-sensitive production traffic.

Xiaomi·262K context·9 Credits / 1M in·29 Credits / 1M out·4.5 Credits / 1M cache read·9 Credits / 1M cache write

MiMo V2.5

General-purpose MiMo model with 1M context and tiered pricing for longer prompts.

Xiaomi·1M context·Tiered pricing:0–256K:40 Credits / 1M in·200 Credits / 1M out·8 Credits / 1M cache read·40 Credits / 1M cache write·256K–1M:80 Credits / 1M in·400 Credits / 1M out·16 Credits / 1M cache read·80 Credits / 1M cache write

MiMo V2.5 Pro

Updated MiMo pro model for broad long-context reasoning and production assistants.

Xiaomi·1M context·Tiered pricing:0–256K:100 Credits / 1M in·300 Credits / 1M out·20 Credits / 1M cache read·100 Credits / 1M cache write·256K–1M:200 Credits / 1M in·600 Credits / 1M out·40 Credits / 1M cache read·200 Credits / 1M cache write

MiniMax M2.5

MiniMax's versatile model with balanced performance across reasoning, creative writing, and multilingual tasks.

MiniMax·200K context·21 Credits / 1M in·84 Credits / 1M out·4.2 Credits / 1M cache read·26.25 Credits / 1M cache write

MiniMax M2.7

Refreshed MiniMax model with broader context and balanced performance for multilingual chat, reasoning, and creative work.

MiniMax·192K context·21 Credits / 1M in·84 Credits / 1M out·4.2 Credits / 1M cache read·26.25 Credits / 1M cache write

Mistral Nemo

Efficient Mistral model for broad multilingual chat and economical generation.

Mistral·131K context·2 Credits / 1M in·4 Credits / 1M out·2 Credits / 1M cache read·2 Credits / 1M cache write

Mistral Small 3.2 24B Instruct

Instruction-tuned Mistral Small model for compact reasoning and production assistants.

Mistral·128K context·7.5 Credits / 1M in·20 Credits / 1M out·7.5 Credits / 1M cache read·7.5 Credits / 1M cache write

Nemotron 3 Nano 30B A3B

Compact NVIDIA Nemotron model for efficient chat and instruction following.

NVIDIA·256K context·5 Credits / 1M in·20 Credits / 1M out·5 Credits / 1M cache read·5 Credits / 1M cache write

Nemotron 3 Super 120B A12B

Larger NVIDIA Nemotron model for stronger reasoning and production assistants.

NVIDIA·262K context·9 Credits / 1M in·45 Credits / 1M out·9 Credits / 1M cache read·9 Credits / 1M cache write

PixVerse C1

PixVerse C1 async video generation route using the official PixVerse API protocol.

PixVerse·33K context·Tiered pricing:360p no audio:3 Credits / 1M out·360p audio:4 Credits / 1M out·540p no audio:4 Credits / 1M out·540p audio:5 Credits / 1M out·720p no audio:5 Credits / 1M out·720p audio:6.5 Credits / 1M out·1080p no audio:9.5 Credits / 1M out·1080p audio:12 Credits / 1M out

PixVerse V6

PixVerse V6 async video generation route using the official PixVerse API protocol.

PixVerse·33K context·Tiered pricing:360p no audio:2.5 Credits / 1M out·360p audio:3.5 Credits / 1M out·540p no audio:3.5 Credits / 1M out·540p audio:4.5 Credits / 1M out·720p no audio:4.5 Credits / 1M out·720p audio:6 Credits / 1M out·1080p no audio:9 Credits / 1M out·1080p audio:11.5 Credits / 1M out

Qwen 3 235B A22B 2507

Large Qwen mixture model for economical reasoning and long-context generation.

Alibaba·262K context·10 Credits / 1M in·60 Credits / 1M out·10 Credits / 1M cache read·10 Credits / 1M cache write

Qwen 3 30B A3B Instruct 2507

Compact Qwen mixture model for fast instruction following and practical chat.

Alibaba·262K context·9 Credits / 1M in·30 Credits / 1M out·9 Credits / 1M cache read·9 Credits / 1M cache write

Qwen 3 32B

Mid-sized Qwen 3 model for affordable reasoning, chat, and content generation.

Alibaba·41K context·8 Credits / 1M in·24 Credits / 1M out·4 Credits / 1M cache read·8 Credits / 1M cache write

Qwen 3 Coder

Qwen 3 coder model for software engineering, code repair, and technical analysis.

Alibaba·262K context·22 Credits / 1M in·180 Credits / 1M out·22 Credits / 1M cache read·22 Credits / 1M cache write

Qwen 3 Coder Next

Qwen coding model for code generation, refactoring, and agentic developer workflows.

Alibaba·262K context·14 Credits / 1M in·80 Credits / 1M out·9 Credits / 1M cache read·14 Credits / 1M cache write

Qwen 3 Coder Plus

Specialized coding model from Alibaba with strong code generation, debugging, and analysis capabilities.

Alibaba·1M context·Tiered pricing:0–32K:40.18 Credits / 1M in·160.58 Credits / 1M out·4.018 Credits / 1M cache read·50.225 Credits / 1M cache write·32K–128K:60.27 Credits / 1M in·240.87 Credits / 1M out·6.027 Credits / 1M cache read·75.3375 Credits / 1M cache write·128K–256K:100.38 Credits / 1M in·401.45 Credits / 1M out·10.038 Credits / 1M cache read·125.475 Credits / 1M cache write·256K–1M:200.76 Credits / 1M in·2,006.97 Credits / 1M out·20.076 Credits / 1M cache read·250.95 Credits / 1M cache write

Qwen 3 Next 80B A3B Instruct

Qwen Next mixture model for efficient reasoning, chat, and instruction following.

Alibaba·262K context·9 Credits / 1M in·110 Credits / 1M out·9 Credits / 1M cache read·9 Credits / 1M cache write

Qwen 3 VL 235B A22B Instruct

Qwen VL model for multimodal-oriented prompts and strong text instruction following.

Alibaba·262K context·20 Credits / 1M in·88 Credits / 1M out·11 Credits / 1M cache read·20 Credits / 1M cache write

Qwen 3.5 27B

Mid-sized Qwen 3.5 model for general chat, analysis, and structured output.

Alibaba·262K context·30 Credits / 1M in·240 Credits / 1M out·30 Credits / 1M cache read·30 Credits / 1M cache write

Qwen 3.5 35B A3B

Efficient Qwen mixture model balancing quality, cost, and multilingual coverage.

Alibaba·262K context·25 Credits / 1M in·200 Credits / 1M out·25 Credits / 1M cache read·25 Credits / 1M cache write

Qwen 3.5 397B A17B

Large Qwen 3.5 model for high-quality multilingual reasoning and synthesis.

Alibaba·262K context·39 Credits / 1M in·234 Credits / 1M out·19.5 Credits / 1M cache read·39 Credits / 1M cache write

Qwen 3.5 9B

Small Qwen 3.5 model for low-latency chat, routing, and simple transformations.

Alibaba·262K context·10 Credits / 1M in·15 Credits / 1M out·10 Credits / 1M cache read·10 Credits / 1M cache write

Qwen 3.5 Flash

Fast and cost-effective Qwen model optimized for high-throughput tasks. Supports up to 1M context with tiered pricing.

Alibaba·1M context·Tiered pricing:0–128K:2.03 Credits / 1M in·20.09 Credits / 1M out·0.203 Credits / 1M cache read·2.5375 Credits / 1M cache write·128K–256K:8.05 Credits / 1M in·80.29 Credits / 1M out·0.805 Credits / 1M cache read·10.0625 Credits / 1M cache write·256K–1M:12.04 Credits / 1M in·120.4 Credits / 1M out·1.204 Credits / 1M cache read·15.05 Credits / 1M cache write

Qwen 3.5 Flash 02-23

Qwen 3.5 Flash snapshot for fast, low-cost multilingual assistant traffic.

Alibaba·1M context·10 Credits / 1M in·40 Credits / 1M out·1 Credits / 1M cache read·12.5 Credits / 1M cache write

Qwen 3.5 Plus

Alibaba's flagship language model with strong multilingual and reasoning capabilities. Supports up to 1M context with tiered pricing.

Alibaba·1M context·Tiered pricing:0–128K:8.05 Credits / 1M in·48.16 Credits / 1M out·0.805 Credits / 1M cache read·10.0625 Credits / 1M cache write·128K–256K:20.09 Credits / 1M in·120.4 Credits / 1M out·2.009 Credits / 1M cache read·25.1125 Credits / 1M cache write·256K–1M:40.11 Credits / 1M in·240.8 Credits / 1M out·4.011 Credits / 1M cache read·50.1375 Credits / 1M cache write

Qwen 3.5 Plus 02-15

Qwen 3.5 Plus snapshot with 1M context and tiered long-prompt pricing.

Alibaba·1M context·Tiered pricing:0–256K:40 Credits / 1M in·240 Credits / 1M out·4 Credits / 1M cache read·50 Credits / 1M cache write·256K–1M:50 Credits / 1M in·300 Credits / 1M out·5 Credits / 1M cache read·62.5 Credits / 1M cache write

Qwen 3.6 Plus

Latest generation Qwen model with improved reasoning and instruction following capabilities.

Alibaba·1M context·Tiered pricing:0–256K:19.32 Credits / 1M in·115.57 Credits / 1M out·1.932 Credits / 1M cache read·24.15 Credits / 1M cache write·256K–1M:77.07 Credits / 1M in·462.14 Credits / 1M out·7.707 Credits / 1M cache read·96.3375 Credits / 1M cache write

Seedance 2.0 (Official)

Official BytePlus ModelArk Seedance 2.0 async video route with direct official-provider task polling and billing.

BytePlus Official·33K context·N/A / 1M in·1,020.5413 Credits / 1M out

Seedance 2.0 (Relay)

Asynchronous text-to-video route served through the ClawOS relay at the discounted relay rate.

ByteDance Relay·33K context·N/A / 1M in·752.2 Credits / 1M out

Seedance 2.0 Asset (Official)

Official BytePlus ModelArk Seedance asset-library video route using official asset groups, asset polling, task polling, and billing.

BytePlus Official·33K context·N/A / 1M in·1,071.885 Credits / 1M out

Seedance 2.0 Asset (Relay)

Relay Seedance video generation with asset-group support for character, face, and reusable media references.

ByteDance Relay·33K context·N/A / 1M in·789.81 Credits / 1M out

Seedance 2.0 Asset Fast (Official)

Faster official BytePlus ModelArk Seedance asset-library video route with official asset ingestion and async task polling.

BytePlus Official·33K context·N/A / 1M in·777.7175 Credits / 1M out

Seedance 2.0 Asset Fast (Relay)

Faster relay Seedance asset-group video route for reference-driven generation with lower turnaround time.

ByteDance Relay·33K context·N/A / 1M in·572.985 Credits / 1M out

Seedance 2.0 Fast (Official)

Official BytePlus ModelArk Seedance fast route for basic async video generation with direct official-provider task polling and billing.

BytePlus Official·33K context·N/A / 1M in·740.5024 Credits / 1M out

Seedance 2.0 Fast (Relay)

Lower-latency Seedance text-to-video route served through the ClawOS relay at the discounted relay rate.

ByteDance Relay·33K context·N/A / 1M in·545.7 Credits / 1M out

Seedance 2.5 (Official)

Official BytePlus ModelArk Seedance 2.5 async video route for text-to-video, frame references, video references, and extended 2.5 controls.

BytePlus Official·33K context·N/A / 1M in·1,016.5 Credits / 1M out

Seedance 2.5 Asset (Official)

Official BytePlus ModelArk Seedance 2.5 asset-library video route with official asset ingestion, async task polling, and extended 2.5 controls.

BytePlus Official·33K context·N/A / 1M in·1,067.325 Credits / 1M out

Step 3.5 Flash

StepFun flash model for quick multilingual chat, extraction, and structured generation.

StepFun·262K context·10 Credits / 1M in·30 Credits / 1M out·10 Credits / 1M cache read·10 Credits / 1M cache write