Discount AI: Why Good Enough Really Is Good Enough
Two weeks ago, Ramp launched a tool that decides which AI model should run a task before the task even starts. Weeks before that, Cursor shipped its own router, trained on more than 600,000 real coding requests, built to make the same call for software teams. OpenRouter, a marketplace built entirely around this problem, has grown to more than 400 models across 60-plus providers. And earlier this week, Perplexity went a step further with something called Hybrid Compute: an agent that scans a task for sensitive data, keeps that part running locally on your device, and hands the rest to whichever frontier model in the cloud is best suited to it.
These four companies are racing toward the same finish line. They're four different answers to a question that didn't exist two years ago: which model, for which job, run where?
Back then the question was simpler. Are you using AI or not. That question is basically settled now, as AI either overshadows or gets folded into nearly every tool we use to get work done. The models themselves also keep getting more capable, to the point that you no longer need the newest, priciest model just to draft an email. That rise in capability changed the question people are asking. Now that most models can do most things well enough, the real question becomes: how do we do this for less? The answer from Ramp, Cursor, and OpenRouter is essentially the same: pick the right model for the job. For a lot of early adopters, though, the ones who locked themselves into rigid workflows built around one specific model doing one specific thing on a tight timeline, the honest answer is closer to "it's complicated."
Tokenmaxxing Was Dumb
Once upon a time, in the storied past of the AI industry, a bunch of tech bros had an idea… tokenmaxxing.
Early in 2026, a strange metric took hold across the tech industry: token spend (AI usage) as a stand-in for productivity. It came to be called tokenmaxxing. The theory went that if a token measures how much computing an AI task consumes, spending more tokens must mean doing more work. Meta ran an internal leaderboard called Claudeonomics that let 85,000 employees compete for titles like "Token Legend." Company-wide consumption hit 60 trillion tokens in a single month, and its top user alone burned through 281 billion tokens. Nvidia CEO Jensen Huang said publicly he'd be alarmed if a $500,000 engineer wasn't spending a serious chunk of that salary's worth on tokens every year.
It's a little like judging a road trip by how much gas got burned instead of how far the car actually went.The employees being judged did whatever they could to burn as much usage as possible.There’s a story of some Meta engineers who built bots to purposefully run in loops, burning tokens like tickets at a street fair. Some set agents running for hours unsupervised, doing deep research (a token heavy task) on nothing in particular.
By April, Uber had blown through its entire 2026 AI budget, four months into the year, after building its own leaderboard to rank engineers on usage. Companies across the industry started calling the FinOps Foundation to brag that they were already three times over their annual token budgets. But the output wasn't matching the spend: research tracking AI-heavy engineering teams found bugs per developer up 54 percent and code churn, the amount of code rewritten or thrown out entirely, up more than 800 percent. Tokenmaxxing didn't measure productivity. It measured activity, which is only the same thing if you’re paying the bill.
By midsummer, the tone had flipped. What Associated Press technology reporter Matt O'Brien described as springtime hype had become, in his July reporting, a summertime backlash. Microsoft pulled back Claude Code licenses across major divisions. One unnamed company reportedly ran up a $500 million bill with no usage limits in place. The term that started as a badge of honor became a mark of shame almost overnight.
Where the Value Lies
Part of that backlash has been a refocus of AI adopters on efficiency. The cooler heads at many companies have realized that wasting money (as well as electricity) is a bad thing, so the best plan is to get as much productivity from these tools as they can for the least amount spent.
Part of that refocus is to use model routing. Not every task needs the latest and most expensive model. Most of what an AI system gets asked to do isn't hard. It's routine: summarize this, format that, answer a common question, draft a first pass. Routing tools like the ones from Ramp, Cursor, and OpenRouter succeed by breaking a workflow into smaller pieces that can be handled by less powerful AI instead of giving every request maximum horsepower.
Once a task gets broken down that finely, smaller and cheaper models can handle far more of the load than most people assume. The frontier model gets reserved for the tasks that actually need it: deep research, the harder reasoning steps, the messier edge cases, the kinds of things a smaller model would likely get wrong or take too long to complete. Everything else runs on something leaner and cheaper. For most of what a business actually needs done day to day, no one can tell the difference. It's just the correct match.
Then there are open-weight models. These are AI models available for anyone to download and run on their own computers or servers, instead of only gaining access through a paid app or website. Because the business owns the model outright, there's no token charges from a provider, no waiting on someone else's servers, and no data leaving the building unless the business chooses to send it somewhere. That makes this set-up ideal for companies with PII or other sensitive data concerns. For a lot of routine, repetitive work, that combination alone makes an open-weight model an attractive choice, even though it’s not the biggest and best.
The Usage Hangover
The companies that got burned by tokenmaxxing didn't retreat from AI. They got more deliberate about it. Uber's chief technology officer, Praveen Neppalli Naga, said in August that the company's AI adoption had quadrupled since January while its cost per token fell, crediting better default model choices, prompt caching, and trials with open-weight models running outside the major cloud providers. Coinbase has described doing something similar in public comments: routing the hardest problems to frontier models and pushing everything repetitive to cheaper ones.
The through-line across most of these recoveries is the same one Ramp and Cursor built their products around: stop treating every task like it deserves the most expensive tool in the shop. Some organizations are pairing that with a shift toward smaller, self-hosted models for high-volume, repetitive work, and seeing great results. None of this is about doing less with AI. Uber's own numbers show usage climbing even as cost per token drops. It's about knowing which tool actually earns its price on a given task.
So the Budget Is Saved?
Not entirely, and there’s a well-understood reason why.
In 1865, the economist William Stanley Jevons noticed something counterintuitive about coal: more efficient steam engines didn't reduce coal consumption, they increased it, because efficiency made coal power worth using for jobs that hadn't been worth it before. Apollo's chief economist, Torsten Slok, has pointed out that AI spending is following the same pattern. As tokens get cheaper, companies don't necessarily spend less. They run more agents, automate more workflows, and generate more output, which pushes total spending up even as the cost per task falls.
That sounds like a wash, but it isn't quite. What's actually changed is visibility. When token spend was tracked on leader boards, almost nobody was asking whether it produced anything. Now that a focus on efficiency has forced businesses to think about which models to use and how, the line between what gets spent and what gets produced is a lot easier to see and manage. In other words, now that they’re paying attention, it’s getting tracked. The bill might not shrink, but for the first time, most companies can actually tell you what it bought.
Is the AI Industry Growing Up?
Discount AI isn't a downgrade. It’s a sign of an industry maturing and putting emphasis on matching the right tool for the right job. The businesses deploying AI successfully right now aren't the ones still chasing the highest benchmark model for every task. They're the ones willing to take a hard look at the processes and systems that they actually use, and organize them into steps that make the most sense with the tools available. It’s a principle as old as work itself, and it’s been playing out in every warehouse, factory, bakery and auto shop for as long as those have existed.
Sam Hardin is the founder of Bright Pivot LLC, an AI automation and operations consultancy based in Covington, Louisiana. Bright Pivot helps small and medium businesses on the North Shore automate the routine and augment the rest.
Have questions about what this means for your business? I'd love to talk it through. Book a time on myCalendly, or reach me directly at sam@bright-pivot.com.
Source Links
Claudeonomics leaderboard, Meta, "Token Legend" titles, 60 trillion tokens/month
Business Insider — https://www.businessinsider.com/tokenmaxxing-ai-token-leaderboards-debate-2026-4
Jensen Huang's $500K engineer / token spend comment
Forbes — https://www.forbes.com/sites/timkeary/2026/07/10/after-tokenmaxxing-token-spend-has-become-the-new-metric-to-watch/
Jon Chu (Khosla Ventures) on looping bots burning tokens to game the leaderboard
Kingy AI — https://kingy.ai/ai/tokenmaxxing-silicon-valleys-most-controversial-new-status-game/
Uber blowing its 2026 budget by April, bugs up 54%, code churn up 800%+
Odin AI — https://getodin.ai/blog/tokenmaxxing-ai-budget/
"Springtime hype to summertime backlash," Microsoft pulling back Claude Code licenses
AP (Matt O'Brien), via ABC News — https://abcnews.com/Business/wireStory/flex-corporate-america-ai-tokenmaxxing-fades-workplaces-cut-135141078
$500 million bill with no usage limits
My Next Developer — https://mynextdeveloper.com/blogs/tokenmaxxing-crashed-and-heres-what-startups-actually-owe-their-boards-about-ai-spend
Uber CTO Naga's August comments, quadrupled adoption, falling cost per token, Jevons paradox, Torsten Slok quote
Fortune — https://fortune.com/2026/08/07/uber-ai-spending-tokenmaxxing-is-over-cto/
(Yahoo Finance ran a similar version if Fortune's paywalled for your readers: https://finance.yahoo.com/technology/ai/articles/blowing-entire-2026-ai-budget-174815792.html)

