
Mostly metrics is proudly powered by Brex
Every decision a company makes is a financial decision. So finance isn't just the function that reports on the business; it's where the business gets decided.
Which makes it worth asking where your finance team's hours actually go. Reconciling last month. Chasing receipts. Policing a $30 subscription. All of it backward-looking.
That's why I use Brex: corporate cards plus AI-native workflows that absorb the record-keeping. Zero-touch receipt matching, real-time policy enforcement, continuous reconciliation. Controls fire before money leaves the account. Close takes days, not weeks.
Defense handled. Go play offense. Thousands of companies — including Anthropic, Coinbase, and DoorDash — already run on Brex. Head to brex.com/metrics to join them.
Reader Happy Hours: London + Boston
The Mostly Metrics world tour continues 🌎
Do you work in finance?
Do you want to drink cold ones with fellow metric aficionados?
RSVP below, rumor is Walter our editor in chief may make an appearance
Join us in London on October 15th and Boston on October 20th


How to Budget for Token Costs Internally
👋 Hi, it's CJ Gustafson and welcome to Mostly Metrics…
Ya ever work on a project / pptx / book report so long that when you finish you just say PLZ take this thing off my hands because it’s eaten my mind? That’s the point I’m at right now. I feel like Leo in the Revenant, emerging from the arctic after sleeping inside of a horse.

these token costs have aged and grizzled me
Too much?
Any who!
This is the second post in our series on managing LLM token costs. Both parts are recorded on video, accompanied by the source files and final outputs to download.
Part 1: How to Audit Your AI Bill for Savings (prior post)
Part 2: How to Budget for Token Costs Internally (this one… keep reading)
Last time we looked at the tokens your product burns, aka the ones you sell through to your end customers to make that dinero. Today we're inside your house…
also, watch this video of a dog who brings a toy to an intruder who breaks into the house. My boy Walter would have done the exact same thing.
In this tutorial we’re analyzing your team's Claude and OpenAI and Cursor seats, plus all the tokens getting gobbled at the company level. And then we give you frameworks to budget for them (‘I’ve got the frameworks!’ (read in an Arby’s ‘we’ve got the meats voice’)).
To set some context for what we are about to look at, our friends at Brex went through their card and bill pay data, and the median company now pays for two AI tools in a given month, with the mean a shade over three. This is up from a median of two and an average of three two years ago.
Same trend with model vendors. 61% of companies pay two or more in the same month, up from 20% two years ago, and the median company crossed over from one to two back in January (free tiers / anything not run through a Brex card are invisible to them, so read all of this as a floor).
OK, let’s budget!
How to try this yourself
First, here's the full walkthrough.
And here's everything behind it.
Open Claude, drop them in, and follow along.
The first thing I do is pull the right reports (there are two)
Last tutorial I used the Admin Usage and Cost API, which groups spend by workspace, API key, and model. That's great for your product's AI costs. But it's useless today, because it has no idea who is who within your org and what they're consuming internally.
What you want is the Spend report. Go to Settings, then Analytics, scroll down to "How much is Claude costing?" and hit Export spend report.
Then I marry it to my headcount roster and my seat list.
I grab three months, so I can make a trend out of this (in this case I used April, May, and June). You'll see net spend is what we actually paid after our negotiated rate, and gross is list price (ya boy loves a good discount).
The roster gives me department, role, level, and comp. Department is the one that is critical to this exercise. The seat list gets me what each person's seat costs, because Claude bills a lot of these as a hybrid - a subscription per employee + additional usage layered on top.
Then I hand the whole pile to Claude and have it join them using email identifiers.
In this example, every row matches to a person… no orphaned emails… Guess we didn't fire anybody at this fictional company this month.
Now every dollar of LLM spend has a name and a department attached to it. I'd encourage you to look at it per head rather than in total. Totals will def tell you engineering is your biggest department, which you already knew. Per head will tell you that finance was a lot higher than you expected.
Here's what falls out once you can see it by person.
Seats nobody is using
Most of the bill on the most expensive model
The average describes nobody (if you budget using averages you need to get your head checked)
Ten people make up a third of the bill
A ceiling on where spend flattens out
The ceiling is going to be a critical component of our budget. The move here is to budget everyone else up to the frontier. I call it frontier convergence.

I have seen the frontier. It’s a scary place.
What we're doing now is forecasting by role, rather than by a named person. We are looking at what a staff engineer costs when they're using AI well, what an SDR costs fully ramped, etc.
We have to encourage a mindset shift. AI doesn't get grandfathered in. It competes from zero, against every other dollar. If this dollar could go to a head, a tool, or a token budget, and it has to win on return, what do you actually pick?
Here are three ways to do it, and they build on each other.

First method: $ per head. Everyone starts here… Total spend, divided by headcount, times next year's headcount.
The trouble is AI usage is lumpy. A handful of heavy users spend many times what a normal person does, and they drag the average up with them.
It's fine as a top-down sanity check, or a first budget envelope pass. But it's useless for actually allocating, because it hides all the variance and nuance.
Second method: Budget AI the way companies budget healthcare. A set percentage on top of base comp, because it's becoming a cost of employing a person, same as benefits loading.

Its best use is as an output or reporting number to ground the board. We spend 20% of comp per employee on tokens.
It's clean, and for your board's higher level math they have a driver. But you as the operator responsible for the plan still don't know what's inside that percentage. It's a great communication tool, not a real lever for allocation.
Third method, and this one came from Meredith Finn, the CFO and COO of Front: Go team by team and find what your top-decile users spend per month.
The assumption is your heaviest users are what the company looks like when people are good at this. They've worked out how to get real output from the tools, and the job is getting everyone else to where they already are.
Do it at the team, then the role level. An engineering manager isn't spending like an engineering IC, and a support rep is probably spending more than their support manager.
Next, smooth the outliers. Use the top decile rather than the single max, or trim the highest and lowest in each group, so one power user doesn't set an unrealistic ceiling.
Then ramp your laggards up to that frontier over a few quarters, like someone getting their feet under them in a new job, with a margin of error on top.
Don't make it land all at once. To use a sales quota analogy, you wouldn't assume a rep is fully productive on day one.

What we've done now is shift the unit from a person to a seat. What a great engineer spends… What a great analyst spends… That's a number you can roll up and forecast with some intellectual honesty.
Interestingly enough, in the video two of mine landed within 10% of each other off completely different methods, which helps me hone in on a range. In many ways, it reminds me of working in valuations, where you run three approaches and go looking for where they agree.

Net net, the method I'd depend most on for forecasting purposes is the third, frontier convergence.
And we can convert that into a model with inputs and outputs.
In the model we built you can enter next year's headcount by role, it prices each head at the frontier, ramps adoption over time (because nobody hits frontier immediately), and adds a margin of error. I love a good buffer.
Once again, here are the files, including the model.
The prompts, in order
1. Join the spend file to your org chart.
"Attached is my Claude enterprise spend report for June and my HR roster. Join them on email. Show me total net spend, total gross spend, and any difference. Then spend by department with headcount, spend by product, and spend by model family. Also give me total requests and net spend per request."
2. Find the money you're wasting before rolling it forward.
"Attached is my seat list showing every paid seat, plus the June spend report and my HR roster. Find everyone holding a paid seat who spent less than $10 in June, including anyone with no rows in the spend report at all. Give me their department and role, and total what those seats cost us annually. Then show net spend by model family as a percent of total."
3. Test the average.
"Show me the average and median net spend per person for June. Then net spend per head by role, sorted highest to lowest, and the ratio between the top role and the bottom."
4. Find your power users.
"Show me the top 10 spenders in June by name, department, role, and net spend, plus what percent of total spend they represent."
5. Check whether the frontier has a ceiling.
"Compare the top 10 spenders against everyone else across April, May, and June. Show the month over month growth rate for each group."
6. Build the persona table.
"Build me a persona table: for each department and role, headcount, current average net spend per head, and the highest spend by anyone in that role, which I'll call the frontier. Add a column for what it would cost to bring everyone in that role up to the frontier. Sort by frontier spend."
7. Run three budgets three ways.
"Give me three budget scenarios for the next 12 months. One, flat average cost per head times headcount. Two, AI spend as a percent uplift on total base comp, including where we are today. Three, everyone converges to the frontier cost of their role. Annual number for each, and show scenario three as a percent of comp."
8. Turn it into a model.
"Build me a budget model where I enter next year's headcount by role and it calculates AI spend using the frontier cost per role. Add an adoption ramp input, since nobody hits frontier immediately, and a margin of error percentage. Show me monthly and annual totals."
What to change when you run it on your real data
Sort out permissions: Spend analytics are for owners and primary owners. So you may need to phone a friend. Finance people should definitely have access to this data.
On a seat plan you're seeing overage, not the whole picture: The report only appears once your org turns on usage credits. If your export comes back empty or looks too small, that's your reason.
You're probably budgeting one vendor out of two: The walkthrough runs on the Claude report because that's what I pulled. Run the persona table again on your other model vendor (OpenAI, Cursors, etc.) and stack them.
The comp column is optional: It's what gets you AI spend as a percent of payroll. Just be careful with the comp data if you send the finished file around.
Don't cancel the silent seats on day one: It might be nobody ever showed them what to do with it, so it's worth ten minutes of asking before you claw back a license.
Use deciles if you're big enough: Your single highest user may be kind of wasteful. So make sure it’s not totally out of whack and you’re unintentionally forecasting people to be on a run away train.
Single-person roles will show zero gap: There's only one CFO in a company, so tag them to another role and use a little art rather than science. The model isn't broken.
So go forth and budget with confidence, please.
CJ









