The bill arrives later.
The loop ran all night.
TokenFence is a proxy you run in front of OpenAI or Anthropic. It prices every call, writes it to a local ledger, and returns 429 to the next request once that key's budget is spent — so a runaway retry loop stops at your ceiling instead of your provider's invoice.
The problem
A retry loop, a bad while condition, and an eval harness left running
overnight look identical from the provider's side: a lot of completed requests, all
billed. A dashboard tells you afterwards. Nothing in the request path says
stop. That is the gap TokenFence fills — it sits between your code and the
provider, and it is the thing that says no.
What it does
Per-key budgets
A daily and a monthly ceiling in dollars, per key. Whichever runs out first binds.
Refuses, doesn't warn
Over budget means 429 budget_exhausted, and the call is never forwarded.
Pre-flight on max_tokens
If the remaining budget cannot cover the requested max_tokens, it is refused before a cent is spent.
Honest about gaps
A provider that reports no usage is recorded at zero cost and flagged in a response header — never silently free.
A refusal, in full
From a local run against a stand-in provider, with a $0.001 daily budget and
a call that cost $0.0025. Notice the last header: it has already gone
negative, so the next call is the one that gets stopped.
HTTP/1.1 200 OK
x-tokenfence-cost-usd: 0.002500
x-tokenfence-prompt-tokens: 1000
x-tokenfence-completion-tokens: 500
x-tokenfence-remaining-usd: -0.001500
The next call on that key, before it ever reaches the provider:
{"error":{"message":"key \"ci\" is out of budget (daily budget $0.001000,
spent today $0.002500; monthly budget $20.000000, spent this month
$0.002500). The cap resets at the next UTC day or month boundary.",
"type":"rate_limit_error","code":"budget_exhausted"}}
HTTP 429
Errors use OpenAI's error shape, so a client that already renders provider errors renders these too. Budgets reset on UTC day and month boundaries.
Configuration is the whole setup
{
"listen": { "host": "127.0.0.1", "port": 8787 },
"upstream": {
"flavor": "openai",
"baseUrl": "https://api.openai.com/v1",
"apiKeyEnv": "OPENAI_API_KEY"
},
"keys": [
{ "name": "ci", "key": "tf_ci_...", "dailyUsdBudget": 2, "monthlyUsdBudget": 20 },
{ "name": "local", "key": "tf_local_...", "dailyUsdBudget": 0.5 }
],
"models": {
"gpt-4o-mini": { "inputPer1M": 0.15, "outputPer1M": 0.6 }
},
"defaultPrice": { "inputPer1M": 1, "outputPer1M": 3 }
}
apiKeyEnv holds the name of the environment variable your provider
key lives in — the key itself never goes in this file. An unlisted model falls back to
defaultPrice rather than being metered as free.
Endpoints
| Route | Purpose |
|---|---|
| POST /v1/chat/completions | Metered, budget-enforced pass-through. Served when flavor is "openai". |
| POST /v1/messages | The same guardrails over Anthropic's dialect. Served when flavor is "anthropic". |
| GET /v1/usage | The calling key's own spend and remaining headroom. |
| GET /healthz | Liveness and version. |
Quickstart
# Node 22.18+ — runs the TypeScript sources directly, no build step
git clone https://github.com/renisjoe/tokenfence
cd tokenfence && npm install
# 58 tests and a full CLI smoke check: no API key, no network, no spend
npm run typecheck && npm test
bash scripts/e2e-check.sh
# point it at a provider
cp examples/tokenfence.config.json tokenfence.config.json
export OPENAI_API_KEY=sk-...
npm start
What it does not do
- No streaming. A streamed response cannot be metered, so
"stream": trueis refused rather than allowed to spend invisibly. Reconciliation after the fact is the first roadmap item. - Cache multipliers are approximate. Anthropic reports cache writes and reads outside the uncached prompt. Every one of those tokens is counted and surfaced as
x-tokenfence-cache-tokens, but the 1.25× and 0.1× cache rates are not modelled — so treat the ledger as a guardrail, not an invoice. - No account, no telemetry. The ledger is a JSON file next to your config. Nothing leaves the machine except the upstream call you configure.
Status
Early stage, and being upfront about it. One maintainer, one repository,
no releases, and no package on npm — npx tokenfence will not work today, so
clone it. There are no users to report, because this was just built. What is verified is
that the parts that exist work: a clean typecheck, 58 passing tests, and an end-to-end
check that boots the real CLI once per dialect against a stand-in provider and asserts
the whole path — forward, meter, block, report.
Tech
Node and TypeScript. No runtime dependencies — the HTTP server, the ledger, the argument parsing and the cost maths are all standard library. MIT licensed.