Token Efficiency & Tips
Learn how BotConnector's caching and routing allow your plan allowance to last 2x to 3x longer, reducing costs and preventing quota waste.
1. Unlocking "Prompt Caching" (Save up to 90% – 98% on Context)
When engaging in long AI conversations, applications re-send the entire conversation history from the start on every turn so the AI retains previous context.
In BotConnector, our gateway automatically enables Prompt Caching (supported on Claude, OpenAI, Gemini, DeepSeek, and GLM). Previously processed conversation history is cached in server memory and read back at deeply discounted official developer rates:
| Model AI | Standard Input Rate / 1M | Cache Read Rate / 1M |
|---|---|---|
| Claude Sonnet 5.5 | $2.00 | $0.20 |
| Claude Opus 5.5 | $4.00 | $0.20 |
| GPT-6 Luna | $0.10 | $0.01 |
| Gemini 3.1 Flash Lite | $0.25 | $0.025 |
| Gemini 3.8 Flash | $0.75 | $0.075 |
| GLM-5.3 Flash | $0.15 | $0.03 |
| DeepSeek V4.1 Flash | $0.15 | $0.003 |
4 Golden Rules to Keep a "Hot Cache":
- Keep Turn Gaps Under 5 Minutes: Provider servers (such as Anthropic and OpenAI) maintain the cache for 5 minutes. Each new message automatically resets this 5-minute timer. As long as you reply within 5 minutes, your conversation history stays in Hot Cache and is billed at the discounted rate.
- 1,024 Token Threshold (Claude Models): On Claude (Sonnet 5.5 and Opus 5.5), Anthropic requires an initial prompt prefix of at least 1,024 tokens (~750 words) to trigger caching. Large reference files, coding files, or long guidelines uploaded in your opening message automatically lock into the cache.
- Use Append-Only Conversations: Caching relies on exact byte-for-byte prefix matching. Avoid editing older turns in the middle of a chat (e.g., modifying question #2 in a 10-turn thread), as editing earlier turns invalidates all subsequent cache. Continue forward instead.
- Static System Instructions: When configuring custom instructions, avoid injecting dynamic real-time timestamps (such as current second/minute). Token changes invalidate the prefix match (cache miss).
2. Choosing the Right Model for the Job
Matching task complexity to model capability preserves your allowance:
- Lightweight Tasks & Quick Q&A: (Definitions, grammar checks, short summaries)
→ Use Free / Flash Models: Agnes 3.0 Flash, Space Bunny, Ling 3.0 Flash, or Gemini 3.1 Flash Lite ($0.00 to $0.0003 per turn). - Daily Research & Drafting: (Brainstorming, drafting articles, strategic analysis)
→ Use Cost-Effective Frontier Models: GPT-6 Luna or GLM-5.3 Flash. Highly capable, fast, and light on allowance. - Complex Coding & Deep Analysis: (Codebases, complex debugging, scientific reviews)
→ Use Frontier Flagships: Claude Sonnet 5.5, GPT-6 Sol, or Gemini 3.1 Pro Preview. Large context files benefit heavily from prompt caching. - Extreme Reasoning:
→ Use Claude Opus 5.5 (Ultra plan exclusive).
3. Tips for Programmers & AI Coding Agents
When using the BotConnector Web App or connecting your API key to coding tools (such as Claude Code, Aider, Cline, Cursor, or OpenCode):
- Load Project Context Upfront: Provide your
README.mdand architecture rules in the initial turn so they remain in Hot Cache. - Send Targeted Requests: Request targeted changes per file or function rather than asking for full-repo rewrites in a single prompt, preserving output tokens.
4. Optimizing Your Image Quota
- Explore with Free Daily Quota: Use your 10–50 daily free images for exploratory concept prompts and rapid drafting.
- Reserve Monthly Tiers for Final Assets: Save monthly quota (Hemat, Menengah, Premium tiers) for high-resolution production assets.