Was Mostly metrics is proudly powered by Brex

Most CFOs I know didn't get into finance to chase receipts.

But that's where the time goes. Reviewing expenses that should have been blocked before they happened. Closing books that take longer to close than the month took to live. Approving spend that should have been approved automatically three days ago.

That's not a people problem. That's a systems problem.

That's why I use Brex. Agentic Finance that automates the receipts, catches out-of-policy spend before it hits the books, and closes your month in minutes instead of weeks. So your finance team stops spending on maintenance and starts spending on momentum.

35,000+ companies, including Mostly Metrics, Anthropic, DoorDash, and Coinbase, have already figured this out. See why it’s time to get Brex AF.

What Will CFOs Do About Token Costs?

It's a nuanced question. And it starts with separating product usage from internal usage.

That means tracking token costs that are part of your product, which show up in your cost of goods sold, versus token costs driven by internal employees, which show up in opex.

Because the costs have gotten way too big to keep in a nebulous, centralized bucket like you’d do with rent at an early stage company.

Why?

1. It will turn into a glut that will blow up your P&L

2. You won't be able to make resource allocation decisions based on what you sell vs what you work on.

I spent time with the Data team at Brex reviewing AI spend trends for the thousands of companies on their platform.

According to their data, when we think about the mix between types of AI spend, foundation model and LLM API providers are north of 80% of all AI expenses, and that category grew 5x year on year.

Source: Brex

Everything else “AI” (design tools, coding tools, sales AI, meeting AI) is only 20% of the spend.

And when we think about where 80% of AI costs land on the P&L, there are different expectations based on geography.

COGS: There Is No "There" There

Weirdly enough, if you are seeking investment you want to prove your token costs within your COGS are significant enough to show you're an AI company (but not so significant that you are selling a dollar for thirty cents). If you still exhibit pristine 85% or 90% pure SaaS gross margins, investors will think there is no “there” there.

And it’s important to call out that the gross margin trend over time is more important than the absolute gross margin % now. 

It reminds me in many ways of Snowflake on their path to IPO. As a cloud infra company they were only recently into +50% gross margin territory in their S1. Over the next 12 quarters they crept towards 70%, which still wasn’t software margins, but made people say “wow look at that improvement and great operating leverage. Imagine what this could become as technology gets better.”

In many ways it's more impressive and investable if you're taking a long term view. So you want it to be improving over time.

Plus, you bank on Moore's law, forecasting you’ll accomplish what you're doing today at a lower cost in the future… you won't always be paying frontier LLM costs for your chatbot and your customers won't always need the Ferrari engine to reset their passwords.

OPEX: Pulling Up the Middle

Inside your company is a different story. At the employee level you start to drill into who is using AI, in what department, and for what job to be done. We’re now dissecting tokens within OPEX.

As an aside, from my experience as a people leader, the 10x employees are also going to be the 10x AI people off the jump. Employees who were amazing employees before AI are self starters and the ones who said let me get even better with AI. So you know they’ll benefit. That’s a given.

The biggest theoretical gains are where you can pull the median person up to be even 20% or 30% better.

As an analogy, I remember talking to a sales leader about quota achievements. He ran a team of ten. He said the top two people are always going to crush it. No need to give them more coaching or tools. And the eighth and ninth and tenth person on the team aren't long for this world anyway. It's the middle of the pack that would move the needle the most.

The same holds for AI, even if it's a team without a quota. Ideally productivity at the midpoint proves to be so great that you can confidently push out anyone who's bottom quartile and redeploy those costs as you see fit (like, on more tokens).

As an aside, from talking to CFOs on the podcast about AI and if it’s taking jobs, not many are getting rid of employees. They are simply choosing to not backfill people when they leave for other opportunities.

Measuring ROI

So now you know where the spending is taking place and where it should sit on the P&L. Analyzing if it’s worth it is a different conundrum.

At an aggregate level there are a few ways to measure gains.

The first is very broad: revenue per employee. That should go up over time. The median public tech company is somewhere around $450K per person, and the best are around $1M. This metric is my favorite measure of output. While it’s a horizontal measure of ROI, there’s no where to hide. It’s the GOAT of metrics.

Two goats

Then you need to come up with departmental and functional measures of success. These are vertical measures of ROI. You can't measure a finance person on code commits to GitHub. They don't code. So how do you measure value add at the functional level?

This takes a lot more work. It needs to link to the department’s value proposition in revenue generation, product creation, or helping the people in those first two buckets do their jobs better (the role of folks in G&A).

We Need to Do More, But Also Do Less

Which sounds simple until you try to build such a measure.

There's a scene in Forgetting Sarah Marshall where Peter Bretter (played by Jason Segel) is learning how to surf. Paul Rudd, his burnt out instructor, oscillates between telling him to do more, and then to do less.

It's kinda like AI. It's very easy to measure activity. But sometimes we actually need people to do less activity, but more of the right activity.

It's hard to measure ROI because it requires new frameworks that are domain specific. Queries per user is a horizontal concept. Any org can look at the number of queries their employees are leveraging on average with an AI tool. It's so much more challenging to then go a step further and say ok is this token spend aligned to, for example, a sales rep increasing deal size upon renewal from what it would have been otherwise.

It's difficult not only because it requires bespoke work; it’s hard because it requires embeddedness and connectivity with internal systems. AI tools can’t be something that lives out there, disconnected from the business.

To stay on the renewal example, you have to know how a deal connects to a CRM so you can tie the client interaction to an expected outcome, you need it to feed into whatever your ERP is to know what it means for the revenue trajectory of a business line, and it needs to tie to performance management systems. All of that is cross functional in nature and much more about enterprise transformation than just AI GTM adoption.

So it’s going from AI 101 to 301 to 501, which is the journey we are on.

The Final Boss

All of that rolls up to one number anyway. And it’s the third way to measure ROI on token spend.

The final boss is free cash flow. The value of a business is predicated on the discounted value of future free cash flows. We forget sometimes that the point of business is to make money… and the point of AI is to make more money, not to use AI.

Lurking in the background is if we ever get to a state where there's more free cash flow produced that can be realized by shareholders in the form of either dividends, share buy backs, or rising equity value per share.

If every company is doing the same things to get better by using AI, then the gains get competed away.

Brex's data showed that 50% to 70% of companies in most industries already pay for some AI, so the gap between industries is depth of spend, not whether anyone is buying.

And if we assume everyone gets to a point where they're somewhat logically deploying AI, and the spending intensity ramps across all sectors, what happens to the profits that are produced by said AI usage? Do any and all incremental free cash flows get plowed back into the business (across cogs and opex) to stay competitive, and just ultimately go back to the LLMs? Is it a zero sum game?

If so, it's not a race to win, but rather a race to not lose.

This will play out across each sector at differing paces. The Brex data shows a 100x intensity spread: software services runs at 2.29x the Brex median.

Source: Brex

If we look at legacy SaaS companies as the canaries in the coal mine, many find themselves less in a race to achieve explosive growth, and more in a race to stay in the game for the sake of staying in the game (aka, not die).

Admittedly, the future is not distributed evenly, so there will be some companies who enjoy a short term uptick from implementing AI org wide. But in the same way AI allowed them to compete to temporarily get ahead it will be accessible to others to compete away margins faster than technological waves in the past. Their alpha becomes the market’s beta.

Building a Gate

So it begs the question… if companies are hammering LLMs, what capabilities are they achieving and to what level of success?

I think we are going to need to evaluate human capital and token capital on the same plane, and ideally we find a mix between the two along some optimized frontier to produce non linear outcomes

Source: Brex

My hunch is there will be a cottage industry of model brokers to make sure firms understand the performance cost frontier and can arbitrage between open source models, frontier labs, and new entrants to get the highest utility token, not just highest volume tokens.

Much of this starts internally by building a gate for LLM costs.

What do I mean by a gate? A checkpoint of sorts. A compelling event that calls for ownership and reflection. 

Every cost line CFOs control have some sort of gate built in. It serves as a moment to pause for a human to decide. Headcount goes through approval processes built into your HRIS before someone can recruit for a new team member. Software has an annual renewal process where the department leader seeks budget and gets an assist from procurement in negotiating. Token spend doesn’t really have a gate (yet). There’s no seat count for most enterprise LLMs, calendar based renewal events aren’t really a thing because they’re usage based and you blow through your commit, and we haven’t put a lot of these on POs yet. 

The last cost that behaved this way was cloud, which took the industry about a decade to get a handle on.

This begins with tagging costs and showing it to the leaders whose teams are consuming those tokens.

Which leads me to my next point - driving to per unit numbers.

Every variable cost eventually gets driven down to a per unit number. This won’t be (or shouldn’t be) different. Token cost per employee per month is useless on it’s own because you don’t know what they’re using those tokens for. You need to get down to cost per ticket resolved, cost per document reviewed, cost per deal cycle, cost per claim completed, whatever unit your business actually solves for on the revenue end, and whatever unit your departments provide value for on the cost side. 

And finally, once you have those unit measurements, your goal is to find alpha against the market’s cost curve. Yours falling faster than the market's is a sign you have control over cost and ROI. 

When everyone has access to the same models and tools, the company that does the best work at the cheapest trackable cost wins. And that company will win by preserving the right to live and fight another day.

If you need help hiring your next FP&A, Strat Fin, or Accounting person, get in touch with my recruiting arm here.

Hoping you don’t try to say AI ROI ten times fast,

CJ

Reply

Avatar

or to participate