Skip to content

Build1 publisher3 min readPublished

Grok's coding CLI shipped whole repos to a cloud bucket. That makes agent adoption an egress call.

An independent researcher says the CLI uploaded a repository it was told not to read, plus a .env secrets file, verbatim. That is a procurement question, not a benchmark question.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Grok's coding CLI shipped whole repos to a cloud bucket. That makes agent adoption an egress call.
Photo: blog.pragmaticengineer.com

What happened

  • xAI (Elon Musk's AI company, now part of SpaceX) released the Grok 4.5 model, built by Cursor (an acquisition) and trained on SpaceX GPUs; Cursor and Grok are now combined as part of SpaceX.
  • The newsletter says Grok 4.5 benchmarks close in coding capability to Opus 4.8 and GPT 5.5 while being 60-70% lower cost.
  • Grok 4.5 can be used via API but is easiest used via the Grok Build coding CLI, which is what many developers did.
  • Cerblab, an independent AI safety researcher, documented that xAI's official Grok Build coding CLI, on a normal consumer login, transmits the contents of files it reads - including a .env secrets file - to xAI verbatim and unredacted, with the secret appearing both in the live model turn (POST /v1/responses) and in a session_state archive uploaded and accepted (HTTP 200) via POST /v1/storage.
  • Cerblab reports the CLI uploads the whole repository - every tracked file's content plus git history - independent of what the agent reads, packaging the workspace and uploading it via POST /v1/storage.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

xAI's Grok Build coding CLI transmits the contents of files it reads, including a .env secrets file, verbatim and unredacted, and separately packages and uploads the entire repository regardless of what the agent actually read, according to an independent AI safety researcher known as Cerblab, whose findings were relayed in a free issue of the Pragmatic Engineer newsletter written by Gergely [4][5][15]. The reason this matters more than the usual agent-tooling squabble is that the destination is durable storage: a Google Cloud Storage bucket named grok-code-session-traces, named verbatim in the binary and in a captured metadata.json [9].

The adoption path is easy to reconstruct. The newsletter describes Grok 4.5 as benchmarking close to Opus 4.8 and GPT 5.5 on coding while costing 60-70 percent less, and says the model is easiest to use through the Grok Build CLI, so that is what many developers reached for [2][3].

The load-bearing evidence is a canary test. Cerblab reports prompting the agent with "reply OK, do not read any files" on a real codebase, after which Grok still uploaded the repository as a git bundle via POST /v1/storage; cloning the captured bundle recovered src/_probe/never_read_canary.txt with its unique marker intact, plus the full git history [6]. On a 12 GB repo of files the agent never read, the storage channel moved 5.10 GiB while the model-turn channel moved 192 KB, a ratio Cerblab puts at roughly 27,800x [7]. That arithmetic checks out, and it is the part that matters: the volume tracks the codebase, not the context window [16]. Even truncated mid-stream, the upload carried something on the order of 40 percent of the tree [17]. No storage upload failed; the only non-200s were a model-usage quota (402/429) on /v1/responses and one unrelated 404 [8].

Two details turn this from a bug report into a policy problem. Cerblab says the mechanism is active by default, was not found in the CLI's install or quickstart materials, and is not disabled by turning off "Improve the model" - /v1/settings still returned trace_upload_enabled: true [10]. And the researcher is careful about scope: none of this proves xAI trains on the data, which is a separate policy question; what is demonstrated is transmission, acceptance and storage [11].

The contrast the newsletter draws is the useful one for anyone writing an internal standard. Every agent ships its context window to whatever server runs the model, and other agents send the parts of the code they read [14]. Cursor, by the newsletter's account, indexes locally, creates embeddings, sends the embeddings, and does not store the user's codebase on its servers - which is the awkward part, given Cursor and Grok now sit inside the same corporate structure [12][1]. The newsletter's own verdict is that storing likely-unencrypted .env files in a GCP bucket is reason enough for a sensible company to ban the CLI [13].

Treat this as an egress decision. Any agentic CLI is a process with network access sitting in a directory that usually contains credentials, customer data and history, and the only defensible way to approve one is to watch its traffic on first run rather than read its landing page.

Worth watching: whether xAI documents the trace upload and makes trace_upload_enabled controllable; whether other researchers reproduce the 12 GB result; and whether anything in the grok-code-session-traces bucket is encrypted at rest [10][7][9].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories