Microsoft has 'good news' and 'bad news' for engineers using AI coding tools
Microsoft is implementing stricter controls on AI token usage for its engineers to ensure cost-efficiency. While the company remains committed to an 'AI-first' strategy, it is shifting focus from raw consumption to 'impact per token.'
Microsoft has told its engineers to keep using AI at work, and to start noticing what it costs. In an internal email to his Core AI teams, executive vice president Jay Parikh wrote that "tokenmaxxing is not what we are optimizing for," and said the company is now managing token spend with the same discipline it applies to every other critical resource. The email, first reported by 404 Media, also confirms that Microsoft has made OpenAI's GPT-5.6—a cheaper model than the frontier options engineers had been reaching for—the default for internal use.The good news for Microsoft engineers is that nobody is being told to stop. Parikh was explicit that he doesn't want to slow the company's march toward being "AI-first," and framed the change as a value exercise rather than a cost-cutting one. The bad news is everything attached to that framing. As of July 2026, every Microsoft division has an AI token budget target. Employees can now see their individual AI spending on an internal dashboard. And the updated Copilot guidelines Parikh linked to note that further restrictions may follow as the company watches where the money goes.What tokenmaxxing means, and why Microsoft wants it to stopTokenmaxxing is workplace shorthand for treating AI consumption as a proxy for productivity—the more tokens you burn, the harder you're presumably working. It produced internal leaderboards at several companies and a lot of frontier-model usage on problems that never needed a frontier model. Parikh's line lands squarely on that behaviour. "I want all of us focused on maximizing outcomes that move the needle for our customers and our business," he wrote, adding that as the company accelerates its use of GitHub Copilot, everyone needs to be aware of how they consume tokens.His closing formulation is the one Microsoft would prefer people quote: the company is not optimising for fewer tokens, but for more impact per token.The numbers behind Microsoft's AI token budgetThe guidelines don't publish a target spend figure, but they do reveal the scale of the problem. Many Microsoft engineers have been running up hundreds of dollars a month in tokens, with some reaching a few thousand. Multiply that across an engineering organisation of Microsoft's size and the line item stops looking like a rounding error.None of this is happening because Microsoft is short of money. Its most recent quarterly results beat Wall Street on revenue, operating income and net income. The awkwardness sits elsewhere, and one Microsoft employee who shared the email with 404 Media said it out loud: a company that has subsidised so much AI inference is now telling its own staff to spend less. That, the employee said, feels like the ultimate admission—and if Microsoft can't afford its own AI products, the question becomes how the customers buying them are meant to manage.Microsoft CEO Satya Nadella had warned about this weeks agoParikh's memo isn't a sudden reversal. Speaking at a live taping of the Hard Fork podcast in June, Satya Nadella was asked how much tokenmaxxing was happening inside Microsoft and answered "a lot" before the question finished. He then admitted to being part of it. "I'm a tokenmaxxer too, it's addictive," Nadella said, before offering the advice that now reads like a preview of the memo: don't use frontier models for non-frontier problems.Nadella stopped short of announcing caps at the time, pointing instead to Copilot's auto mode as the way to match tasks to appropriately sized models. Seven weeks later, Parikh has supplied the enforcement mechanism.Amazon, Adobe and Meta got there before Microsoft didAmazon, Adobe, Atlassian and Citi have all introduced throttling or spending visibility in recent months. Meta went further, imposing token budgets and then shutting down "Claudeonomics," the internal leaderboard that had employees competing to burn tokens. Microsoft had already made a quieter move in May, cancelling most Claude Code licences in its Experiences and Devices group and pushing those engineers onto GitHub Copilot CLI before the fiscal year closed. Alex Heath, writing in his newsletter Sources, reports that before the switch to GPT-5.6, Microsoft's internal Copilot setup ran on an auto-router that defaulted largely to Anthropic's models—meaning the company's own engineers were mostly coding with expensive Claude tokens on Microsoft's dime. Heath also reports that GitHub Copilot was running at steeply negative gross margins before it moved to usage-based billing earlier this year.What makes the maths uncomfortable is that none of this is happening because AI got more expensive. Per-token prices have fallen roughly 98 percent since late 2022. Bills tripled anyway, because agentic tools chew through vastly more tokens per task than the autocomplete-era interactions the original pricing was built around. Parikh's memo is Microsoft conceding that the second number matters more than the first, and that the company selling the meter has to read its own.Get the latest technology news and updates. Download the TOI App.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in