Microsoft's top AI boss has a message for engineers using GitHub Copilot
Microsoft is implementing stricter internal controls on AI token usage for its engineers to manage rising operational costs. While the company remains committed to an AI-first strategy, it is shifting focus from raw volume to optimizing the impact of each token used.
Microsoft has spent three years telling the world to use more AI. This week it told its own engineers to use it more carefully. In an internal email to his CoreAI organisation, executive vice president Jay Parikh wrote that "tokenmaxxing is not what we are optimizing for" — and backed the line with policy. OpenAI's cheaper GPT-5.6 is now the default model for internal use, every Microsoft division has an AI token budget target, and employees can pull up a dashboard showing exactly what their own AI habit costs the company.The memo was first reported by 404 Media, which obtained it from a Microsoft employee who asked not to be named. The numbers inside are the eye-catching part. Microsoft's internal Copilot guidelines say many engineers spend anywhere from a few hundred dollars to a few thousand dollars a month on tokens. No hard cap figure has been shared yet, but the guidelines warn that further restrictions may follow as spending is monitored.The company that sells AI is now rationing it internallyParikh framed the change as discipline, not distress. "As we accelerate our use of GitHub Copilot to deliver on our goals, we all need to be aware of how we consume tokens," he wrote, adding that Microsoft would manage token spend with the same discipline it applies to every other critical resource. He was careful to say this is not a retreat from being AI-first. "We are not optimizing for fewer tokens. We are optimizing for more impact per token."The financial context supports him. Microsoft's latest earnings beat Wall Street on revenue, operating income and net income. This is not a company scrambling for cash. It is a company discovering that agentic coding tools consume tokens at a rate nobody budgeted for. Per-token prices have dropped roughly 98 percent since late 2022, yet enterprise AI bills have tripled, because an autonomous agent chewing through a repository burns through vastly more than an autocomplete suggestion ever did.Nadella called himself a tokenmaxxer two months before the memo landedThe shift did not come out of nowhere. In June, at a live taping of the New York Times' Hard Fork podcast, Satya Nadella was asked how much tokenmaxxing happens inside Microsoft. "A lot," he said, cutting off the question. "I'm a tokenmaxxer too, it's addictive. But you have to step back when the novelty wears off to say, 'What is it that I'm trying to create?'" His advice then was to match the model to the task. Don't use frontier models for non-frontier problems.Microsoft had already been trimming. It cancelled most Claude Code licences in its Experiences and Devices group in May, pushing engineers to GitHub Copilot CLI before the fiscal year closed on June 30.Amazon, Adobe, Citi and Meta got here firstMicrosoft is arguably the last big name to formalise this. Amazon, Adobe, Atlassian and Citi have all introduced some form of AI throttling or spend visibility. Meta shut down "Claudeonomics," an internal leaderboard that had employees competing to burn tokens.Not everyone inside Redmond reads it as prudence. The employee who leaked the memo called it "the ultimate admission" that a company hosting AI infrastructure cannot afford its own AI products, and asked the obvious follow-up: how could the companies Microsoft sells to manage?Parikh's answer is in the memo itself. Keep using AI. Just know the bill has a name on it now, and it might be yours.Get the latest technology news and updates. Download the TOI App.
Get the full story
Sign up for Headlinne to unlock AI insights, political bias analysis, and your personalized news feed.
Create free accountAlready have an account? Sign in