The Attach War Is Won. The Margin War Is Being Lost.
Microsoft nets out an internal transfer between application and infrastructure P&Ls that most software companies cannot — and the same quarter handed challengers a trust wedge against bundled copilots.
The margin drop matters less than the mechanism underneath it, and the mechanism is the part that transfers. Inference cost lands in the application P&L and the profit lands in the infrastructure P&L. Microsoft owns both sides, so the transfer nets out at the group level and reads as a segment-mix story. Most software companies own one side. They pay the inference and book none of the infrastructure margin.
The same asymmetry is visible a layer up. Google Cloud grew 82% to roughly $25B in the reported quarter against Azure's 43%, and The Information attributes the divergence to Anthropic outgrowing OpenAI. Google supplies one. Microsoft supplies the other, and disclosed $24B of OpenAI-related revenue, about 7% of total company revenue, in the same breath. Model selection has stopped being only an architecture decision. It is a bet on which hyperscaler's incentives stay aligned with the buyer's in 2028.
Why the renewal window is open
Both suppliers need something a buyer can hand over cheaply: a nameable proof point. Azure is guiding to one to two points of deceleration. Google is funding 82% growth through its first quarterly cash burn as a public company. A skeptic would say neither of those is a crisis, and the skeptic is correct. Neither position is durable either, which is exactly why the leverage exists while those positions hold and disappears once reservations tighten. AWS confirms the timing from the other direction: Andy Jassy disclosed customers already reserving capacity for 2028, with most AI capacity contracted on five-year terms, while OpenAI cut prices on two models weeks after release to answer bill-shock complaints.
Infrastructure is being sold five years forward while model output deflates in weeks. Almost every AI business plan built on current assumptions assumed those two prices moved together.
Read together, the sources agree on the direction and disagree usefully on magnitude. One line of reporting models 30-50% structural inference deflation. Another argues compute gets up to 10x more expensive as labs refuse to divert training capacity to serving customers. Both readings punish the same thing: a product priced as a markup on tokens. The hedge is identical under either forecast. Contracted capacity below, owned workflow above, and nothing in the middle carrying gross margin.
The wedge nobody has picked up
Two Microsoft 365 Copilot flaws, surfaced by two independent security firms, would have let a single malicious Word document pull any file or email across a tenant. One was patched in April. The second is being withheld because it is unclear whether Microsoft has patched it. Rubrik reports the exploit produced no suspicious signal in Copilot's logs, and Microsoft declined to confirm whether Purview covers sandbox activity.
For a buyer, that is a written question to put in front of a renewal rather than a headline to react to. For a seller in regulated verticals, it is a clear counter-positioning opening: provable data isolation and audit logging beats bundling in a market where the bundled incumbent cannot evidence what its assistant did. The terms agreed this renewal cycle set the ceiling on the next one.
What to do
Re-underwrite AI unit economics at the feature level before the next pricing cycle: gross margin per AI-active account with inference COGS isolated, and a contribution floor below which a feature gets repriced or retired.
Reopen cloud and model commitments this quarter while both suppliers still need named proof points, targeting a 15%+ reduction in blended cost-per-token plus contractual 2027 capacity.
Demand a written answer from your assistant vendor on whether its audit logging covers sandbox activity, and what retrospective forensic coverage exists for the pre-patch window.