LLM Gateway Kit
Cost-Safe LLM Gateway for Swift Apps · AI Infrastructure
Semantic and vision response caching, tiered circuit breakers, and hard cost-budget enforcement for Swift LLM apps, built so a runaway prompt loop can’t blow through your API budget.
View source on GitHubHow it was built
Problem
An iOS app calling LLM APIs directly has no natural circuit breaker: a retry storm, a redundant vision call, or a provider outage can burn through a monthly budget in hours, and usage dashboards only tell you after the money is gone.
Decisions
- Added semantic and vision response caching so near-duplicate prompts and images don't re-trigger a full paid call.
- Built tiered circuit breakers that fall back tier by tier when a provider misbehaves, instead of a binary up/down switch.
- Enforced cost budgets as a hard ceiling in the gateway itself, not as a downstream alert on a usage log, so the limit holds even if nobody is watching.
Outcome
A hard budget cap stops runaway spend before it happens. The patterns were extracted from the LLM gateway in an iOS app that is in TestFlight beta, and there is no published cost benchmark yet.
Key features
- Semantic and vision response caching to cut redundant LLM calls
- Tiered circuit breakers that degrade gracefully under provider failure
- Hard cost-budget enforcement, not just usage logging after the fact
- Built for native Swift/iOS apps calling LLM APIs directly
Impact
Prevents a single misbehaving client or feedback loop from exceeding a hard cost ceiling.
Tech stack
Swift