Guardrail by NeatProxy
The runtime firewall
for AI coding agents.
See coding-agent spend. Check your budget before the next supported request reaches the provider.
Blocking requires a supported proxy integration. Codex with a ChatGPT subscription is visibility only; Cursor support is planned. Check compatibility.
Local proxy · No prompt storage · Mode-specific enforcement
The problem
Agents spend money faster than anyone can watch.
A coding agent runs unattended for minutes at a time, and a single retry loop can burn more than a month of subscription before anyone looks. The tools that exist today tell you afterwards.
Bills arrive late
Provider billing is a monthly rear-view mirror. By the time a number looks wrong, the money is already spent.
Counters only count
A token counter watches the meter run. It has no opinion about when to stop, and no way to act on one.
Alerts are not brakes
A threshold alert fires after the spend it is warning you about. A budget you cannot enforce is a preference.
Inspect a real analysis of fictional data
A total tells you what. A sample shows you why.
Try the bundled sample audit, inspect the contributing rows, and read the rule behind each finding. No upload or email required.
Spend scenario
Explore the numbers. See the assumptions.
Model a possible opportunity, or compare fictional requests with and without a cap. Neither is a savings guarantee.
Your assumptions, made explicit. Nothing here changes a real budget.
Choose assumptions, then calculate. Editing a value clears any previous result.
Same requests. Different outcome.
Eight fictional requests at $0.25 each, with and without a cap. This toy comparison is independent of your monthly spend and is not a simulation of provider pricing.
Cap checks in this sample occur before the next equal-cost request. Real usage and policies differ.
Reduced-motion preferences skip playback. Switching tools cancels and resets the demo.
Without a cap
- Request 1 · waiting
- Request 2 · waiting
- Request 3 · waiting
- Request 4 · waiting
- Request 5 · waiting
- Request 6 · waiting
- Request 7 · waiting
- Request 8 · waiting
$0.00
With sample cap
- Request 1 · waiting
- Request 2 · waiting
- Request 3 · waiting
- Request 4 · waiting
- Request 5 · waiting
- Request 6 · waiting
- Request 7 · waiting
- Request 8 · waiting
$0.00
# pricing_model
pay for the cloud, not the firewall.
Guardrail runs free on your machine, with no credit card and no time limit. Pro adds the hosted dashboard, longer history, and sync across every machine you own. Team is one flat plan for the whole group.
// free
Local firewall, budget policies, local dashboard, 7 days of history. No credit card required.
// pro
Everything in Free, plus the hosted dashboard, 90 days of history and sync across machines. $144 yearly.
// team (5 seats)
A flat bundle for up to 5 seats, with an admin console, per-member spend and shared project caps.
// TWO BILLING MODELS. ONE REAL RISK.
Your login mode determines what Guardrail can control.
Blocking requires a supported proxy integration. Codex with a ChatGPT subscription is visibility only; Cursor support is planned.
Visibility is not the same as blocking
Subscription usage and token-equivalent costs are not your provider bill. What Guardrail can see and control depends on the tool and how you sign in.
- •Claude Code requests routed through the local proxy can be checked against budget policies.
- •Codex with a ChatGPT subscription provides usage telemetry only. It cannot block requests or guarantee remaining subscription quota.
Check the policy before sending
Metered requests can add cost while an agent retries. A supported proxy integration checks the configured policy before forwarding the next request.
- •Pre-send hard dollar caps on
localhost:4000before calls hit the wire. - •A blocked request returns a policy error. How the agent retries or recovers depends on the tool.
How it works
One command in front of your agent.
For supported proxy integrations, Guardrail sits between your coding tool and the provider. It forwards requests and records usage metadata, so policies can block the next call before it is sent.
Connect your tool
Connect Claude Code with your existing login or Codex with an API key. Codex subscription mode provides telemetry, not proxy enforcement.
It reads the metadata
Supported proxied requests flow through localhost:4000 byte-for-byte. Guardrail records model, tokens, cache and estimated cost — not prompt text, responses, or code.
See it, then cap it
A live local dashboard shows spend, sessions and hidden cost. Set a budget and an over-budget call is blocked before it ever reaches the provider.
The request body is forwarded untouched. Only model, token counts and cost estimates are written locally to 127.0.0.1
What Guardrail controls
Spend control, not another dashboard to babysit.
Local visibility and control for supported integrations. Cloud sync is a separate feature. See how each one works.
Control
Budgets that block
Caps on dollars, requests, tokens or requests-per-minute. Over budget, the call is never sent, so it costs nothing.
Visibility
Hidden cost, surfaced
The cache-write tax on a first turn and the reasoning tokens you never see, attributed per request.
Privacy
Metadata stays local by default
Personal usage metadata stays on your machine unless you enable sync. Team sync defaults are explained in the privacy guide.
Coverage
Support depends on your login
Claude Code and Codex API-key requests support blocking. Codex subscription usage is visibility only; agent frameworks are in preview.
Forecast
Know the month before it lands
Spend pace and a projected month-end total from the sessions you have already run.
Flow
Pause without restarting
Flip tracking off and back on from the dashboard. Your tool keeps working throughout.
Why Guardrail is different
Prevention, not reporting.
Most tools in this space observe. Guardrail sits in the request path, which is the only place a budget can actually be enforced.
Provider dashboards
Show you the bill
Accurate and far too late. They are scoped to an account, not to the project or the session that caused the spike, and they cannot stop anything.
Token counters
Show you the meter
Useful for curiosity, limited for control. Watching a number climb does not stop the next request from being sent.
Observability platforms
Analyse after the fact
Rich analysis, but it arrives after the spend, and it usually means sending prompts and responses to a third party to store.
Guardrail
Stops the request
Runs on your machine, in the path, with your own credentials. Policies are evaluated before each call, so an over-budget request is never sent and never billed.
FAQ