Menu

I cut my LLM API bill by 70%: AirAi practical notes on routing + caching

0. Background: How did the costs explode?

I work on several side projects. In the early stages, I either had to pay the official subscription fee of $20/month or endure the rejection of my international credit card applications. When I settled the accounts at the end of the month, I discovered two frustrating issues:

  1. 80% of the requests were for simple tasks such as summarizing, categorizing, and formatting data. Using the most expensive models for these tasks was a complete waste;

  1. The subscription fee is a fixed cost – it results in losses when there are few requests, and it's not enough when there are many requests.

So I turned to AirAi: I used my own API key for pay-as-you-go usage and directed the client to its OpenAI-compatible interface.

There's also a little side story here. I never had a decent international credit card, and every time I tried to make a payment using the official direct connection, it was rejected. For a while, I used a family member's card, but that triggered risk control measures and my account was frozen. It took almost two weeks to get it unlocked after I filed an appeal. After that, I made up my mind to only use services that didn't rely on international cards. The gateway I finally chose only accepts USDT – which, for someone like me without an international card, turned out to be the option with the lowest requirements: you can deposit coins and use them immediately, there's no monthly fee, and you don't need to go through the KYC (Know Your Customer) process. Even more important is the price: in my experience, the cost per transaction is only about half of what it is with the official direct connection. That's really the reason I decided to use this service.

1. One-line switching: Point to any compatible gateway

Almost all clients that support custom endpoints (Claude Code, Cursor, Open WebUI, LibreChat, Cline, etc.) can be switched using a single environment variable without needing to modify the code. Just point the following two variables to your own gateway:

export OPENAI_BASE_URL="https://api.airai.cc/v1"
export ANTHROPIC_BASE_URL="https://api.airai.cc/v1"

Choosing which gateway to use is another matter; this article only focuses on the implementation method.

2. Code practice: Routing + Caching

Just switching lights isn't enough; what really saves money is “using the right model for the right task + caching repeated requests.” Here’s the framework I use in production (compatible with the OpenAI SDK). Both base_url and api_key are read from environment variables and are not hardcoded into the code:

import openai, hashlib, os

client = openai.OpenAI(
    base_url=os.getenv("OPENAI_BASE_URL"),
    api_key=os.getenv("OPENAI_API_KEY"),
)

# Routing: small model for light tasks, flagship only for complex reasoning
CHEAP  = "gpt-5.6-luna"   # classification / summarization / formatting
STRONG = "gpt-5.6-sol"    # only called for complex reasoning

_cache: dict[str, str] = {}   # swap for Redis in production

def ask(prompt: str) -> str:
    model = STRONG if len(prompt) > 400 else CHEAP
    key = hashlib.md5(prompt.encode()).hexdigest()
    if key in _cache:                      # cache hit, 0 cost
        return _cache[key]
    resp = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": prompt}],
    )
    answer = resp.choices[0].message.content
    _cache[key] = answer
    return answer

Adjust CHEAP and STRONG according to your budget and desired results. The model name should be based on the documentation of the gateway you are using. Remember to use the Batch interface for offline tasks, as it can usually save another 50% in costs.

3. Cost Comparison (Illustrative)

Before and After the TransformationIllustrationData (please refer to your actual bills for the real figures):

Project

Official Subscription / Direct Connection

AirAi + Routing Caching

Billing Method

Fixed Monthly Fee / Official Price

Pay-As-You-Go

Lightweight Models

Flagship Models (More Expensive)

Small models (cheaper)

Repeated requests

Full-price recalculation

Cache hits: 0 cost

Perceptual cost

Baseline: 100%

Approximately 10–30% (depending on the number of requests)

4. Honesty about the limitations (let’s be honest)

  • Not everyone will save money: A fixed subscription may be more cost-effective when the monthly number of requests is very low; pay-as-you-go pricing is only advantageous when the number of requests is high. Calculate your own usage first.

  • Latency and stability: Adding an extra layer of gateway will cause a slight increase in latency, as well as potential issues with timeouts and system degradation on critical paths.

  • Key security: Store the key in environment variables; don’t hardcode it in the front-end or in a public repository.

5. Summary

The essence of reducing the cost of using large language models (LLMs) can be summed up in one sentence: let the cheaper models do most of the work, use the more expensive models only when necessary, and eliminate any wasted resources (tokens). Also, change the requirement from “connecting with N systems” to “just modifying one line of code” (base_url). Routing and caching are two free tools that we should start using right away.


Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR