Models
Browse BotConnector Cloud models, then discover local models matched to your hardware.
Most Used AI Models on BotConnector
Real usage statistics across BotConnector Cloud. The leaderboard below is derived directly from live aggregated request logs (`cloud_usage_events`), reflecting actual model adoption.
The number-one model requested by BotConnector users for rapid conversation and task execution.
Actual cumulative token volume executed in real time through BotConnector cloud inference routes.
High-performing model for complex logical reasoning, long document synthesis, and intelligent agent workflows.
Authentic BotConnector Usage Data: A total of 488 requests and 2.21 million tokens have been processed across the top 10 models on BotConnector. Synchronized from operational telemetry.
Integrated Cloud Models
Official Cloud AI models on BotConnector, verified and ready for web chat, API keys, and CLI integrations. Includes verified provider logos, clear context specs, and transparent routing indicators.
Free Starter — Chat & Reasoning
GPT-6 Luna
OpenAI model for chat, coding, reasoning, tools, and long-context work. 1M tokens/month during Launch Access. Included with Pro when subscriptions launch.
Mistral Medium 3.5
Frontier-class multimodal Mistral model for agents, coding, structured output, and document work.
Agnes 3.0 Flash
Agnes AI multimodal reasoning model for text, images, video, adjustable reasoning, and tool calling.
Space Bunny Alpha
Anonymous stealth model with fast inference, strong coding, multimodal input, adjustable reasoning, and 1M context.
Ling 3.0 Flash
Ling 3.0 Flash Fin by inclusionAI, optimized for multi-step financial workflows while retaining strong reasoning, coding, and math.
Mistral Nemo
12B multilingual model from Mistral AI and NVIDIA, strong in conversation, reasoning, and coding with long context.
NVIDIA Nemotron 3 Nano 30B A3B
Efficient NVIDIA MoE model for reasoning, coding, instruction following, and long context with few active parameters.
NVIDIA Nemotron 3 Super
NVIDIA Nemotron 3 Super for reasoning and chat with long context. The Free Cloud route was verified directly from BotConnector infrastructure.
DeepSeek V4 Flash
Efficient DeepSeek MoE model for chat, content creation, RAG, reasoning, and high-concurrency workloads.
Tencent Hy3
Tencent hybrid fast/slow-thinking MoE model for reasoning, coding, instruction following, and agent workflows.
Xiaomi MiMo V2.5
Xiaomi full-modal model for text, image, video, and audio understanding, with strong agent capabilities and 1M context.
GLM 5.3 Flash
Efficient multimodal GLM-5 model for complex tasks, visual understanding, coding, and professional agent workflows.
Ling 3.0 Flash Sante
inclusionAI Ling 3.0 Flash Sante for reasoning, coding, and analytical work.
North Mini Code
Cohere coding model for coding agents, debugging, reasoning, and tool use. Its Free Cloud route was benchmarked from BotConnector infrastructure.
Nemotron 3.5 Lightning
Long-context NVIDIA model for reasoning, coding, and agent workflows. The Free Cloud route was benchmarked from BotConnector infrastructure.
Nemotron 3 Ultra
Large NVIDIA reasoning model for planning, coding, and agent workflows with long context. The Free Cloud route was benchmarked from BotConnector infrastructure.
Gemma 4 26B A4B
Google Gemma 4 A4B multimodal model for chat, vision, video, reasoning, and tool workflows. The direct Google route was benchmarked from BotConnector infrastructure.
Gemma 4 31B
Google Gemma 4 31B multimodal model for chat, visual understanding, reasoning, and tool workflows. The direct Google route was benchmarked from BotConnector infrastructure.
Free Starter — Media & Utility
Agnes Image 2.1 Flash
Free Cloud image generation and editing for Starter accounts through the live-validated Agnes route.
PAYG — Chat & Reasoning
GLM 5.3 Flash
Efficient multimodal GLM-5 model for complex tasks, visual understanding, coding, and professional agent workflows.
DeepSeek V4.1 Flash
New-generation DeepSeek model with a large context window.
GLM 5.2
GLM model for general and developer workflows.
MiniMax M3
PAYG MiniMax M3 variant for sustained cloud workloads.
Qwen3.5 397B A17B
Large Qwen MoE model for complex workloads.
Kimi K3
Multimodal model with tools, reasoning, and structured output.
Qwen3.8 Flash
Multimodal Flash model with tools, reasoning, and a 1M context window.
GPT-5.6 Luna
OpenAI model with tools, vision, reasoning, and a large context window.
GPT-5.6 Terra
Higher-tier OpenAI model for complex workloads.
GPT-5.6 Sol
High-capability OpenAI model for reasoning and knowledge work.
Claude Sonnet 5
Anthropic model for reasoning, coding, and complex work.
Claude Opus 5
High-tier Anthropic model for complex tasks and reasoning.
Gemini 3.8 Flash
Gemini 3.8 Flash through the Google Gemini API for PAYG cloud workloads with vision, tools, and reasoning.
Gemini 3.8 Flash Flex
Gemini 3.8 Flash Flex through Google PAYG Flex routing for lower inference cost when capacity is available.
Gemini 3.7 Flash
Gemini 3.7 Flash through the Google Gemini API for PAYG cloud workloads with vision, tools, and reasoning.
Gemini 3.7 Flash Flex
Gemini 3.7 Flash Flex through Google PAYG Flex routing for lower inference cost when capacity is available.
Gemini 3.6 Flash
Gemini 3.6 Flash through the Google Gemini API for PAYG cloud workloads with vision, tools, and reasoning.
Gemini 3.6 Flash Flex
Gemini 3.6 Flash Flex through Google PAYG Flex routing for lower inference cost when capacity is available.
Gemini 3.5 Flash
Gemini 3.5 Flash through the Google Gemini API for PAYG cloud workloads with vision, tools, and reasoning.
Gemini 3.5 Flash Flex
Gemini 3.5 Flash Flex through Google PAYG Flex routing for lower inference cost when capacity is available.
Gemini 3.5 Flash Lite
Gemini 3.5 Flash Lite through the Google Gemini API for PAYG cloud workloads with vision, tools, and reasoning.
Gemini 3.5 Flash Lite Flex
Gemini 3.5 Flash Lite Flex through Google PAYG Flex routing for lower inference cost when capacity is available.
Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite through the Google Gemini API for PAYG cloud workloads with vision, tools, and reasoning.
Gemini 3.1 Flash Lite Flex
Gemini 3.1 Flash Lite Flex through Google PAYG Flex routing for lower inference cost when capacity is available.
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview through the Google Gemini API for PAYG cloud workloads with vision, tools, and reasoning.
Gemini 3.1 Pro Preview Flex
Gemini 3.1 Pro Preview Flex through Google PAYG Flex routing for lower inference cost when capacity is available.
Solar Mini 4
Fast, cost-efficient Upstage model for chat, reasoning, coding, structured output, and multilingual documents.
Solar Pro 4
Upstage Solar Pro model for richer reasoning, coding, structured output, and professional document workflows.
PAYG — Generative Media & Utility
Gemini 3.1 Flash Image
PAYG image generation and editing at 1K through Gemini.
Gemini 3.1 Flash Lite Image
Cost-efficient PAYG image model for 1K generation.
FLUX.2 [klein] 4B
Low-latency economy image generation and editing with PAYG settled from the task's actual upstream cost.
Veo 3.1 Fast
Fast PAYG video generation; video adapter is being integrated.
Veo 3.1 Lite
Lower-cost PAYG video generation; video adapter is being integrated.
Lyria 3.5
Full-song generation from prompts; music adapter is being integrated.
Qwen Image
PAYG text-to-image model from a connected provider catalog.
Qwen Image Edit
PAYG instruction-based image editing model.
Kling 3.0 Pro Text-to-Video
Text-to-video generation; video adapter is being integrated.
Qwen Audio 3.0 TTS Flash
Low-latency PAYG text-to-speech.
Qwen3 TTS Flash
PAYG text-to-speech billed by text characters.
Text Embedding V4
PAYG embeddings for retrieval and semantic search.
Ming Image 0.1 Design
Image model for UI, posters, infographics, and text-rich design. Kept under PAYG; its current rate state is Currently Free.
Ming Image 0.1 Design Layer
Image-to-image variant that decomposes designs into editable RGBA layers. Kept under PAYG; its current rate state is Currently Free.
Status Note: “Available now” indicates an active, real-time execution route. “Integration queued” means the model is cataloged while its runtime adapter is undergoing live verification.