the AI & cloud cost gateway

Don't let vibe coding eat your budget.

slice sits in front of every AI call your team makes, routes the easy ones to cheaper models, and cuts the bill before it runs away.

Early. One person. The gateway is live.
a slice of cake
click to cut
slice · measured run50-prompt dev workload
Easy edits, lint, one-linersrouted to Haiku$0.03
Refactors and unit testskept on Sonnet$0.11
Repeat promptsserved from cache$0.00
50 prompts, sent both ways, same models priced the same
Direct to Anthropic$0.26
Through slice$0.14
You saved
45%
$0.26 direct, $0.14 through slice, the same 50 prompts
projected · at that ratio, a team spending $8,000 a month keeps about $3,600 of it
SAVED

measured · the same 50 prompts, sent direct and through slice

get started

From install to first request.

About three minutes. You sign in once with GitHub, then point your tools at slice.

(like this)  means a value you fill in with your own. Everything else you can copy as-is.

01

Install slice

The slice CLI is a small package on PyPI. Install it and the slice command lands on your path.

Terminal
$ pip install slice-gateway

On Mac, use pip3 if pip is not found.

02

Sign in with GitHub

Run slice login. slice shows you a short code. Open github.com/login/device, sign in to GitHub, paste the code, and click Authorize. That is the whole login, no password typed in the terminal.

Why GitHub? You already have it, slice never stores a password, and it signs you in right from the terminal, where slice lives.

Terminal
$ slice login

  To finish signing in, open:
    https://github.com/login/device
  and enter this code:
    WXYZ-1234

Waiting for you to authorize in the browser...
Logged in as (your-github-username)
03

Point your tools at slice

Run slice use claude-code. It prints three environment variables. Set them so your AI tools send their calls through slice, which routes, caches, and meters every request.

Terminal
$ slice use claude-code

export ANTHROPIC_BASE_URL=https://api.sliceapp.dev
export ANTHROPIC_API_KEY=(your own Anthropic key)
export ANTHROPIC_AUTH_TOKEN=(your slice key, from slice login)
# ANTHROPIC_AUTH_TOKEN goes out as Authorization: Bearer, where slice reads its key.
# ANTHROPIC_API_KEY stays your own Anthropic key in x-api-key; slice forwards it.
# Claude Code notes that env auth takes precedence over your claude.ai login while
# these are set, which is expected; unset the three variables to go back to normal.

ANTHROPIC_AUTH_TOKEN goes out as Authorization: Bearer, which is where slice reads its key. ANTHROPIC_API_KEY stays your own Anthropic key in x-api-key, and slice forwards it upstream.

Optional: see your cloud bill too. GitHub and an email are all slice needs. But if you want your AWS bill next to your AI spend, connect a read-only role from Settings in the dashboard. It only reads your account name and costs. It never sees your keys, your data, or anything that could change your bill.

04

Use your tool as normal

That is it. Keep coding the way you already do. Every call now flows through slice, and you can watch spend, savings, and budgets live on your slice dashboard.

Open source. slice is yours to read line by line, or to run on your own box. It sees what your tools send to the model, the same text the provider already gets. It logs each call's model, tokens, cost, and status, keeps the first 4,000 characters of each prompt so the router learns, and caches answers for one hour. No other prompt or answer text is kept.

What it never touches. Your repo, your machine, your environment variables. Your provider key is forwarded and never stored. The AWS role is read-only: bucket settings, security groups, encryption, IAM policy attachments, and the bill. It cannot open a file, read a secret, or change anything, and you can revoke it any time.

on every request

What slice runs.

Routeeach request to the cheapest model that works.
Cacherepeated calls, so you never pay twice for the same answer.
Capeach team's spend. Warn as it fills, block at the ceiling.
Recommenda cheaper model, or a flat-fee tool when one fits the work better.
Watchyour AWS bill and mistakes, read-only, right next to AI spend.
Public bucket exposed
s3://acme-user-uploads is open to the whole internet and holds files that look like IDs and passports. slice caught it before anyone else did.
measured, not promised

The number.

A fixed 50-prompt developer workload, sent twice: once straight to Anthropic, once through slice, both naming the same baseline model. Same work, smaller bill.

Same 50-prompt workload, through slice
45.28%
cheaper. $0.141615 through slice vs $0.258801 direct.
Routed27 requests dropped to a cheaper model, saving $0.087675.
Cached5 repeat prompts served at $0, saving $0.029511.
Hard19 prompts stayed on the model you asked for.

Routing plus cache reconcile exactly to the $0.117186 saved.

the request path

How a request moves.

Your tool talks to slice. slice guards, caches, judges, and routes, then calls the provider. One hop in, one hop out.

client
Anthropic or OpenAI JSON.
→
slice gateway
guardcachejudgerouter
→
providers
AnthropicOpenAIGeminiNVIDIA NIM
local judge → router
Qwen2.5 0.5B on llama.cpp, local, free.
Haiku fallback
Ina request arrives as Anthropic or OpenAI JSON.
Guardauth resolves your slice key, then the budget cap blocks anything over the line.
Cachea repeat is served at once, never touching a provider.
Routethe router picks the served model: pin, then rule, then the judge on the auto path.
Outthe adapter calls the provider, then logging and the dashboard run after.
our own model

The judge is ours.

The router judge is a Qwen2.5 0.5B model, trained on slice's own routed traffic, scoring 92% on 100 unseen prompts. Shrunk to 4-bit, it runs on the same small box, answering in 259 to 365 ms per decision at zero cost per call. If it is down, Haiku steps in.

the hard tail

Hard prompts run a relay.

A hard request does not jump to one expensive model. It runs a ladder, cheapest first, and stops the moment an answer is good enough. Gemini checks each draft, and only a fail climbs.

1 WorkerNIM takes the easy, high-volume work first.
2 DrafterGPT writes the first full answer for the rest.
3 CheckerGemini reads the draft and decides good or wrong.
4 CloserSonnet steps in only for the hard part, last and least.

Take a team sending 1,000 requests a day. Straight to a strong model, a typical 1,000-in, 500-out request costs about $0.0105, so about $315 a month. Through the relay, say 700 land on the near-free NIM worker, 240 pass the Gemini check at the GPT rung, and 60 climb to Sonnet, plus the router judge running locally on all 1,000 at no cost per call. The day lands near $3 against about $10.50 going direct, roughly a 70% cut on the same work.

inside your editor

Ask slice from Claude Code.

slice ships an MCP server with six tools, so you can ask about spend, rules, recent requests, and eval without leaving Claude Code. Type a plain question and the answer comes back inline.

Claude Code: your spend this month
Claude Code showing this month's slice spend against the budget
Claude Code: your last five requests
Claude Code showing the last five requests with model, cost, and routed from

Set it up in one command. See the how-to for the steps.

before it's a problem

You'll hear it from us, not from finance.

Set a budget. The moment a team crosses the line, slice emails you with the number, the cause, and how long you have.

Gmail
slice alertsalerts@sliceapp.dev
9:42 AM
Budget warning: coding team at 82%
You've spent $4,920 of $6,000 this month. At today's pace you'll hit the cap in about 5 days. Biggest driver: Opus on lint fixes, which slice can route down to Haiku…
Open dashboard →

Set the number, slice does the rest. Email today, Slack next.

a note from me

Why I built slice.

The idea showed up while I was craving a slice of cake. The name stuck, and so did the problem.

All through 2026 I kept reading the same story: great AI tools, adopted fast, billed per token with no ceiling and nobody watching the meter. Uber torched its year's coding budget by April. Amazon burned $1.8M on a single task. That is not a tooling problem. It is a missing brake.

So I built the brake. slice meters every call, routes the easy ones to cheaper models, caches the repeats, and cuts you off before the invoice does.

It is early, and I would rather say so than pretend. One person, deploying now. Try it on your own machine and judge it for yourself.

the stories that got me obsessed