AI News

Microsoft rebuilds its AI so that any model can be swapped out

Satya Nadella closed a record $331.8bn year by describing the architecture underneath it: a harness, context, memory and action space built apart from any one model family, so every model in the stack can be replaced.

Listen to this article

--:--
Editorial collage: a Microsoft four-square logo beside a rack of interchangeable model cards being lifted out of a slot, with data panels reading FY26 $331.8B +18%, AZURE $100B +41%, FOUNDRY 100,000 CUSTOMERS.

Microsoft closed its financial year on 30 June with revenue of $331.8bn, up 18%, and operating income of $155.2bn, up 21%. Azure passed $100bn for the first time, growing 41% across the year and 43% in the final quarter. On the call on 29 July, Satya Nadella spent his opening minutes on the engineering decision underneath those figures.

“We are building a new model system, where the harness, context, memory, and action space are separate from any one model family,” he told analysts, “thereby moving the frontier on the cost-to-outcome curve.” Mustafa Suleyman, who runs Microsoft AI, published the fuller version the same day: “Every model in a product or agentic system should be substitutable, and that’s only possible when you build the harness, context, memory and action space independently of a single model family.”

The four pieces Microsoft wants to own

All four sit around a model rather than inside it. The harness is the scaffolding that runs the thing: routing, retries, tool calls, evaluation, the logic that decides which model sees which task. Context is what the system puts in front of the model on any given turn. Memory is what it keeps between turns and between sessions. The action space is the set of things it is permitted to do, from reading a spreadsheet to filing a ticket.

Build those four around one vendor’s API and the product inherits that vendor’s prices, release schedule and risk profile. Build them independently and the model becomes a component with a socket behind it. Suleyman gives the resilience case directly: a business over-dependent on one family is exposed to “a security incident, a business or policy misalignment, or a geopolitical shift”.

Small models, tuned per product

Suleyman’s framing is that the industry is turning a corner on cost. “Tokenmaxxing has been the story of the last few months, but token efficiency is the next big focus,” he wrote, and Microsoft’s answer is a family of small in-house models trained for one product each rather than one large model asked to do everything.

The numbers Microsoft attaches to them, all its own:

Model Where it runs Microsoft’s claim
MAI-Cyber-1-Flash MDASH security harness No. 1 on CyberGym, beating Mythos by 12 percentage points, at 50% of the cost
MAI-Code-1-Flash GitHub, VS Code, Excel 10% higher code accept rate, 10% lower median token usage
MAI-Image-2.5-Flash PowerPoint GPU costs down up to 84%, save rates up 26%
MAI-Voice-2-Flash Dynamics 365 Contact Center GPU costs down up to 89%

On the call Nadella put the Excel case in competitive terms: MAI-Code-1-Flash is “delivering comparable quality to GPT-5.6 for the most common tasks, while operating at significantly lower cost”. The principle Suleyman states is that specialisation buys the discount: “By tuning models for a specific product, you can maintain or even exceed frontier performance, while reducing token costs dramatically.”

We have already seen the mechanism work in public. When Microsoft released MAI-Cyber-1-Flash on 27 July, the figure it led with was a score for the whole system around the model. Swapping the small model into 80% of the roles inside its multi-agent MDASH harness moved the harness itself from 88.4% to 95.95% on CyberGym, with GPT-5.4 held back for the hardest tenth of tasks. The harness was the product; the models inside it were parts, and the cheap part did most of the work.

The last discount is silicon. Microsoft says it gets “40% better performance per watt when running MAI models on Maia 200”, its inference accelerator announced in January: TSMC 3nm, 216GB of HBM3e at 7 TB/s, 272MB of on-chip SRAM, over 10 petaFLOPS at FP4, and 30% better performance per dollar than the newest hardware already in its fleet. Maia 200 went into the US Central region near Des Moines first, with Phoenix next.

The swap in numbers: Microsoft’s FY26 revenue, Azure and capital expenditure, its four in-house MAI models and the claims attached to each, the harness diagram with four model families slotting into it, and the Copilot figures.

Foundry is where the swapping happens for everyone else

The same architecture is sold as a product. Foundry now has 100,000 customers and its revenue more than doubled year on year, with the number of customers running at a one trillion token annualised rate up fourfold. The catalogue holds “over 11,000 models”, and the statistic that tells you what customers do with it is this one: since January, the number of them building with models from more than one provider has risen fivefold.

Nadella named the example. Levi Strauss & Co. “is using models from OpenAI and Anthropic on Foundry, as it brings more than 1,000 domain-specific agents into a unified enterprise AI platform”. Telefónica has adopted Foundry as the foundation of its corporate agentic platform. Both are companies buying the harness and treating the models as interchangeable stock, which is exactly the shape Microsoft is arguing for.

Copilot moves from a chat box to standing agents

“Copilot is evolving rapidly, from chat to Cowork to Autopilots,” Nadella said, and each of those three is now a shipped thing rather than a roadmap slide. Cowork went generally available last month for multi-step tasks grounded in a customer’s own work data, and usage-based billing was added to it earlier in July, with thousands of customers already paying. Autopilots arrived this quarter as autonomous, long-running agents with full enterprise compliance, “including an always-on personal agent powered by OpenClaw”. This quarter the separate surfaces converge: Microsoft will bring the Copilot experiences together, including Code, into one “super app” spanning consumer and commercial.

The adoption figures behind that: over 30 million paid Microsoft 365 Copilot seats, net seat additions more than doubling quarter on quarter, and Copilot revenue accelerating more than 60% on the previous quarter after usage-based billing arrived. Quality moved too, in the direction the model system predicts. Nadella: “Over the last three quarters, user satisfaction scores have doubled and are now at an all-time high. And this quarter alone, we cut latency by 25%.”

A 25% latency cut inside three months is the kind of gain that comes from routing and harness work rather than from a new frontier model, because no new frontier model shipped inside Copilot in that window.

What the efficiency is funding

Microsoft is spending the savings immediately. Capital expenditure was $41bn in the June quarter alone, including the effect of higher component pricing, taking the full year to $115.9bn, and Amy Hood told analysts it will grow again in FY27. Commercial remaining performance obligation, the contracted revenue not yet recognised, stands at $678bn, up 84%, which is the demand that spending is chasing.

Her guidance for the year ahead: double-digit revenue and operating income growth, operating expenses up in the mid to high single digits, full-year operating margins down by less than a point, free cash flow positive, and an effective tax rate near 20%.

One line in the release is worth separating from the operating story. Quarterly net income of $35.8bn was up 31% on a reported basis but 22% on a non-GAAP basis, and the gap is Microsoft’s stake in OpenAI. Net gains from those investments added $480m to net income and $0.07 to diluted EPS in the quarter, and $4,963m and $0.67 across the year. A year earlier the same line was a loss of $1,575m in the quarter and $3,620m for the year. The swing from loss to gain is worth roughly a point of headline growth, and none of it comes from selling software.

The claim is testable, and the next model release is the test

Substitutability is easy to assert while every model in production is your own. The check comes the first time a rival ships something Microsoft’s own family cannot match on a task Copilot depends on. If the harness is genuinely built apart from any one model family, that model appears in production in days and the cost per outcome falls again. If it takes a quarter, the sockets were never quite as standard as the architecture diagram said.

Microsoft has given everyone the yardstick to hold it to: performance per token, outcome per dollar, and a public catalogue of 11,000 alternatives sitting in its own cloud.

Sources

  1. Microsoft FY26 Q4 earnings press release, Investor Relations, 29 July 2026microsoft.com
  2. Microsoft FY26 Q4 earnings call and prepared remarks, Investor Relations, 29 July 2026microsoft.com
  3. Mustafa Suleyman, 'Optimizing the frontier performance curve', Microsoft AI, 29 July 2026microsoft.ai
  4. 'Maia 200: The AI accelerator built for inference', Official Microsoft Blog, 26 January 2026blogs.microsoft.com
  5. Microsoft AI, 'Introducing MAI-Cyber-1-Flash inside MDASH', 27 July 2026microsoft.ai