Back to work

LLM Gateway Kit

Cost-Safe LLM Gateway for Swift Apps · AI Infrastructure

Semantic and vision response caching, tiered circuit breakers, and hard cost-budget enforcement for Swift LLM apps, built so a runaway prompt loop can’t blow through your API budget.

View source on GitHub

How it was built

Problem

An iOS app calling LLM APIs directly has no natural circuit breaker: a retry storm, a redundant vision call, or a provider outage can burn through a monthly budget in hours, and usage dashboards only tell you after the money is gone.

Decisions

  • Added semantic and vision response caching so near-duplicate prompts and images don't re-trigger a full paid call.
  • Built tiered circuit breakers that fall back tier by tier when a provider misbehaves, instead of a binary up/down switch.
  • Enforced cost budgets as a hard ceiling in the gateway itself, not as a downstream alert on a usage log, so the limit holds even if nobody is watching.

Outcome

A hard budget cap stops runaway spend before it happens. The patterns were extracted from the LLM gateway in an iOS app that is in TestFlight beta, and there is no published cost benchmark yet.

Key features

  • Semantic and vision response caching to cut redundant LLM calls
  • Tiered circuit breakers that degrade gracefully under provider failure
  • Hard cost-budget enforcement, not just usage logging after the fact
  • Built for native Swift/iOS apps calling LLM APIs directly

Impact

Prevents a single misbehaving client or feedback loop from exceeding a hard cost ceiling.

Tech stack

Swift

Interested in something like this?

Get in touch