slice sits in front of every AI call your team makes, routes the easy ones to cheaper models, and cuts the bill before it runs away.
measured · the same 50 prompts, sent direct and through slice
About three minutes. You sign in once with GitHub, then point your tools at slice.
(like this) means a value you fill in with your own. Everything else you can copy as-is.
The slice CLI is a small package on PyPI. Install it and the slice command lands on your path.
$ pip install slice-gateway
On Mac, use pip3 if pip is not found.
Run slice login. slice shows you a short code. Open github.com/login/device, sign in to GitHub, paste the code, and click Authorize. That is the whole login, no password typed in the terminal.
Why GitHub? You already have it, slice never stores a password, and it signs you in right from the terminal, where slice lives.
$ slice login To finish signing in, open: https://github.com/login/device and enter this code: WXYZ-1234 Waiting for you to authorize in the browser... Logged in as (your-github-username)
Run slice use claude-code. It prints three environment variables. Set them so your AI tools send their calls through slice, which routes, caches, and meters every request.
$ slice use claude-code export ANTHROPIC_BASE_URL=https://api.sliceapp.dev export ANTHROPIC_API_KEY=(your own Anthropic key) export ANTHROPIC_AUTH_TOKEN=(your slice key, from slice login) # ANTHROPIC_AUTH_TOKEN goes out as Authorization: Bearer, where slice reads its key. # ANTHROPIC_API_KEY stays your own Anthropic key in x-api-key; slice forwards it. # Claude Code notes that env auth takes precedence over your claude.ai login while # these are set, which is expected; unset the three variables to go back to normal.
ANTHROPIC_AUTH_TOKEN goes out as Authorization: Bearer, which is where slice reads its key. ANTHROPIC_API_KEY stays your own Anthropic key in x-api-key, and slice forwards it upstream.
Optional: see your cloud bill too. GitHub and an email are all slice needs. But if you want your AWS bill next to your AI spend, connect a read-only role from Settings in the dashboard. It only reads your account name and costs. It never sees your keys, your data, or anything that could change your bill.
That is it. Keep coding the way you already do. Every call now flows through slice, and you can watch spend, savings, and budgets live on your slice dashboard.
Open source. slice is yours to read line by line, or to run on your own box. It sees what your tools send to the model, the same text the provider already gets. It logs each call's model, tokens, cost, and status, keeps the first 4,000 characters of each prompt so the router learns, and caches answers for one hour. No other prompt or answer text is kept.
What it never touches. Your repo, your machine, your environment variables. Your provider key is forwarded and never stored. The AWS role is read-only: bucket settings, security groups, encryption, IAM policy attachments, and the bill. It cannot open a file, read a secret, or change anything, and you can revoke it any time.
s3://acme-user-uploads is open to the whole internet and holds files that look like IDs and passports. slice caught it before anyone else did.A fixed 50-prompt developer workload, sent twice: once straight to Anthropic, once through slice, both naming the same baseline model. Same work, smaller bill.
Routing plus cache reconcile exactly to the $0.117186 saved.
Your tool talks to slice. slice guards, caches, judges, and routes, then calls the provider. One hop in, one hop out.
The router judge is a Qwen2.5 0.5B model, trained on slice's own routed traffic, scoring 92% on 100 unseen prompts. Shrunk to 4-bit, it runs on the same small box, answering in 259 to 365 ms per decision at zero cost per call. If it is down, Haiku steps in.
A hard request does not jump to one expensive model. It runs a ladder, cheapest first, and stops the moment an answer is good enough. Gemini checks each draft, and only a fail climbs.
Take a team sending 1,000 requests a day. Straight to a strong model, a typical 1,000-in, 500-out request costs about $0.0105, so about $315 a month. Through the relay, say 700 land on the near-free NIM worker, 240 pass the Gemini check at the GPT rung, and 60 climb to Sonnet, plus the router judge running locally on all 1,000 at no cost per call. The day lands near $3 against about $10.50 going direct, roughly a 70% cut on the same work.
slice ships an MCP server with six tools, so you can ask about spend, rules, recent requests, and eval without leaving Claude Code. Type a plain question and the answer comes back inline.
Set it up in one command. See the how-to for the steps.
Set a budget. The moment a team crosses the line, slice emails you with the number, the cause, and how long you have.
Set the number, slice does the rest. Email today, Slack next.
The idea showed up while I was craving a slice of cake. The name stuck, and so did the problem.
All through 2026 I kept reading the same story: great AI tools, adopted fast, billed per token with no ceiling and nobody watching the meter. Uber torched its year's coding budget by April. Amazon burned $1.8M on a single task. That is not a tooling problem. It is a missing brake.
So I built the brake. slice meters every call, routes the easy ones to cheaper models, caches the repeats, and cuts you off before the invoice does.
It is early, and I would rather say so than pretend. One person, deploying now. Try it on your own machine and judge it for yourself.
Leaderboards pushed Claude Code adoption. The CTO said he was back to the drawing board, the budget already blown away.
Internal metrics showed a single job running roughly 860% over its budget, a catastrophically expensive blunder.
The tool caught on faster than budgets could take, so Microsoft cancelled most direct licenses and moved to Copilot CLI.
For his team, the cost of compute is “far beyond the costs of the employees.” Usage-based bills keep outrunning the plan.