tokenfence
Open source · self-hosted · MIT

The bill arrives later.
The loop ran all night.

TokenFence is a proxy you run in front of OpenAI or Anthropic. It prices every call, writes it to a local ledger, and returns 429 to the next request once that key's budget is spent — so a runaway retry loop stops at your ceiling instead of your provider's invoice.

The problem

A retry loop, a bad while condition, and an eval harness left running overnight look identical from the provider's side: a lot of completed requests, all billed. A dashboard tells you afterwards. Nothing in the request path says stop. That is the gap TokenFence fills — it sits between your code and the provider, and it is the thing that says no.

What it does

Per-key budgets

A daily and a monthly ceiling in dollars, per key. Whichever runs out first binds.

Refuses, doesn't warn

Over budget means 429 budget_exhausted, and the call is never forwarded.

Pre-flight on max_tokens

If the remaining budget cannot cover the requested max_tokens, it is refused before a cent is spent.

Honest about gaps

A provider that reports no usage is recorded at zero cost and flagged in a response header — never silently free.

A refusal, in full

From a local run against a stand-in provider, with a $0.001 daily budget and a call that cost $0.0025. Notice the last header: it has already gone negative, so the next call is the one that gets stopped.

HTTP/1.1 200 OK
x-tokenfence-cost-usd: 0.002500
x-tokenfence-prompt-tokens: 1000
x-tokenfence-completion-tokens: 500
x-tokenfence-remaining-usd: -0.001500

The next call on that key, before it ever reaches the provider:

{"error":{"message":"key \"ci\" is out of budget (daily budget $0.001000,
spent today $0.002500; monthly budget $20.000000, spent this month
$0.002500). The cap resets at the next UTC day or month boundary.",
"type":"rate_limit_error","code":"budget_exhausted"}}

HTTP 429

Errors use OpenAI's error shape, so a client that already renders provider errors renders these too. Budgets reset on UTC day and month boundaries.

Configuration is the whole setup

{
  "listen": { "host": "127.0.0.1", "port": 8787 },
  "upstream": {
    "flavor": "openai",
    "baseUrl": "https://api.openai.com/v1",
    "apiKeyEnv": "OPENAI_API_KEY"
  },
  "keys": [
    { "name": "ci",    "key": "tf_ci_...",    "dailyUsdBudget": 2,   "monthlyUsdBudget": 20 },
    { "name": "local", "key": "tf_local_...", "dailyUsdBudget": 0.5 }
  ],
  "models": {
    "gpt-4o-mini": { "inputPer1M": 0.15, "outputPer1M": 0.6 }
  },
  "defaultPrice": { "inputPer1M": 1, "outputPer1M": 3 }
}

apiKeyEnv holds the name of the environment variable your provider key lives in — the key itself never goes in this file. An unlisted model falls back to defaultPrice rather than being metered as free.

Endpoints

RoutePurpose
POST /v1/chat/completionsMetered, budget-enforced pass-through. Served when flavor is "openai".
POST /v1/messagesThe same guardrails over Anthropic's dialect. Served when flavor is "anthropic".
GET /v1/usageThe calling key's own spend and remaining headroom.
GET /healthzLiveness and version.

Quickstart

# Node 22.18+ — runs the TypeScript sources directly, no build step
git clone https://github.com/renisjoe/tokenfence
cd tokenfence && npm install

# 58 tests and a full CLI smoke check: no API key, no network, no spend
npm run typecheck && npm test
bash scripts/e2e-check.sh

# point it at a provider
cp examples/tokenfence.config.json tokenfence.config.json
export OPENAI_API_KEY=sk-...
npm start

What it does not do

Status

Early stage, and being upfront about it. One maintainer, one repository, no releases, and no package on npm — npx tokenfence will not work today, so clone it. There are no users to report, because this was just built. What is verified is that the parts that exist work: a clean typecheck, 58 passing tests, and an end-to-end check that boots the real CLI once per dialect against a stand-in provider and asserts the whole path — forward, meter, block, report.

Roadmap: streaming with usage reconciliation · prompt-cache multipliers in the price table · a cost estimate for a batch before it is sent · threshold warnings at 80%. A hosted version is the intended direction, but the open-source proxy is the whole product today, and it is complete enough to run.

Tech

Node and TypeScript. No runtime dependencies — the HTTP server, the ledger, the argument parsing and the cost maths are all standard library. MIT licensed.