Build1 publisher2 min readPublished
Microsoft publishes token and cache telemetry from 301,026 Copilot agent sessions
Microsoft released metadata from 301,026 GitHub Copilot agent sessions, covering 9.3 million LLM calls in one June week. Capacity planners now have real cache and token figures to test against, though the files measure resource use only and cannot show whether the output was any good.
The Engineer · Build desk

What happened
- The public files cover June 1 to 7, 2026, and hold a uniformly sampled set of sessions from non-enterprise Copilot users.
- The dataset card counts 631.4 billion prompt tokens, of which 540.95 billion were cached, across 1,189,581 user turns and 37 anonymized model labels.
- Timings, token counts, cache behavior and tool-call sequences are kept, while prompts, responses, source code, file paths, repository names and user or organization identifiers are excluded.
- The paper's full study is far larger, spanning 3.2 million users, 13 million sessions, 761 million LLM calls and 95 trillion tokens.
- In that full study, cache hit rates averaged about 90% within a turn, fell to 55% across turn boundaries and dropped further after model switches.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability Teams can work out calls per session, prompt tokens per call and cached share from production agent traffic themselves, without relying on a vendor's summary.
- cost A cost model that applies one blended cache rate to every call will undercount spend on multi-turn sessions if the study's drop at turn boundaries also shows up in a team's own traffic.
- constraint Anonymized model labels and a US, non-enterprise, two-IDE sample mean no figure maps cleanly onto a named model's price sheet or an enterprise fleet.
- decision Choosing between coding agents still needs a separate quality evaluation, because this release covers one product and leaves out the responses needed to judge correctness.
The dataset card is enough for a sizing baseline that does not depend on trusting the paper. The slice averages about 30.9 LLM calls per session [1] over roughly four user turns [2]. That works out to about 7.8 model calls each time a developer sends a request [8]. Each call carried around 67,900 prompt tokens on average [3]. The slice has about 0.94 tool calls per LLM call [4]. The full study reports a near one-to-one ratio, with 87% of calls initiated by the agent [11].
For a token budget, check the cache line first. About 85.7% of the slice's prompt tokens were cached [5]. The figure is a token-weighted share for the whole week, which makes it a different measure from the paper's per-turn hit rates. It still leaves about 90.45 billion uncached prompt tokens in one sampled week [6].
The paper's headline findings come from the full study. The public slice is about 2.3% of that study's sessions [7]. According to Runtimewire, which reported the release, those findings were not independently verified from the released sample [14]. Haoran Qiu, an Azure Research engineer and co-author, pointed to the files on X, and they sit in Microsoft's AzurePublicDataset repository [2]. Microsoft includes a notebook intended to reproduce the paper's figures [7]. Running it on the slice is the cheapest first check.
A figure transfers to another fleet only if that fleet's traffic resembles the source. The telemetry came from Copilot's agent in Visual Studio and VS Code, in US regions spanning no more than three time zones [9]. A week of non-enterprise traffic in June is a good sample of non-enterprise traffic in June. The model labels are anonymized, so a cache curve cannot be tied to the model a team actually pays for [10].
Retries are the largest multiplier in the paper. The researchers found that tool failures in 9% of turns could trigger retries that raised compute use by as much as four times [13]. They propose scheduling around whole sessions and reclaiming resources while a developer reads the result [15]. Their idle-time predictor captured 86% to 90% of total idle time in its evaluation, a result Runtimewire calls a systems experiment [16].
The metadata-only design is good engineering. It keeps the timing, token and cache fields a scheduler needs and drops anything that identifies a developer or their code [3]. It is published under a CC-BY license [7]. That choice also sets the limit. Without responses, the data cannot show whether Copilot's answers were correct or useful [4]. It cannot compare Copilot with Claude Code or Codex either [10].
What to watch
- Whether independent runs of Microsoft's notebook on the 301,026-session slice reproduce the paper's cache and retry figures.
- Whether Microsoft publishes enterprise sessions or a window longer than seven days, which would test how far the June figures generalize.
- Whether any trace release pairs resource telemetry with outcome signals such as accepted edits or passing tests.