Build1 distinct publisher3 min readUpdated
Glean says its customers mostly turn on automatic model selection to control spend, not to improve answers. The routing layer, not the model, is where enterprise AI budgets now get decided.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Stripe has bought OpenRouter for more than $7B [1], and at Glean the model-selection feature customers actually use is the automatic one, chosen mostly for economic reasons [6][7]. Both facts point at the same layer: the thing that decides which model handles a request, or whether a model is needed at all, is where enterprise AI spend is now controlled [8].
Arvind Jain, Glean's co-founder and CEO and a former Google Distinguished Engineer [4], gave Latent Space the arithmetic behind that. The newest frontier models cost double or quadruple the per-token rate of their predecessors, and because they can run longer tasks, users run longer tasks, so per-user spend can land at 10 to 20 times last year's level [9]. Nothing about that curve is fixed by picking a cheaper vendor. It is fixed, if at all, by not sending every request to the expensive tier.
Glean exposes three levels of control: employees can pick a model explicitly, administrators can restrict models or impose usage limits, and automatic mode selects per task [6]. "Why are people talking about model routing? Why are they excited about it? It's mostly because of cost," Jain said [5]. The sharper version of the same idea is the null route. "A big goal of Glean is to avoid using LLMs for tasks where we don't need them," Jain said, citing queries where someone is adding or multiplying two numbers and could have used a calculator [8].
The commercial claim attached to this is single-sourced and worth reading carefully. Glean engineering lead Tony Gentilcore recently said the product is "4x more cost-effective" than Claude Code, averaging $0.45 per task against $1.84 for Claude Cowork, which he attributed to Glean's harness and routing [10]. Those two figures divide out to about 4.1x [17]. No task mix, sample size, or quality comparison was given in the source, and a per-task average says nothing about how many tasks a routed system needs to finish the same work.
What is harder to copy is the observation data. Glean reports $300M in ARR this year, a three-fold increase over 15 months [2], which implies roughly $100M 15 months earlier [18]. The company was last valued at $7.2B after a $150M Series F [3], meaning Stripe paid about what the private market last put on Glean as a whole [20]. Deployment breadth is the point: Zillow reports 80% adoption across 7,000 employees [11], roughly 5,600 people [19], and Booking.com describes Glean as its first company-wide AI platform [12]. Jain says that gives Glean a view of which model people reach for first and when they escalate after an unsatisfying answer, which feeds back into routing [13].
Latent Space frames the demand as a product of frontier-model competition plus increasingly capable open-weight models such as Kimi K3 and Qwen3.8-Max [14]. Jain's own pitch is a meta-harness: "You can think of Glean today as a superset of ChatGPT, Claude, Gemini, Grok" [15]. The architecture also includes a model called Waldo [16].
Watch whether any routing vendor publishes quality-held-constant cost data rather than per-task averages, and whether Stripe's OpenRouter price shows up as pass-through pricing pressure on the enterprise routers.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Jain: "Why are people talking about model routing? Why are they excited about it? It's mostly because of cost."
Automatic mode is mostly chosen by Glean's customers for economic reasons.
Glean reached $300 million in annual recurring revenue this year, a three-fold increase over 15 months.
Zillow reports 80% adoption of Glean across 7,000 employees.
Jain says Glean observes which models users select first for different task types and when they upgrade to another model after being unsatisfied, and that this human feedback loop at scale improves its routing system.
Latent Space attributes rising demand for model routing to intense competition among frontier model companies together with the increasing power of open-weight models like Kimi K3 and Qwen3.8-Max.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Thin: one publisher, vendor-narrated
Every claim in the cluster derives from a single Latent Space interview with Glean's CEO, with one secondhand founder benchmark. Product mechanics (three-tier selection, Waldo, pre-LLM assembly) are internally consistent and specifically described, which lifts the floor. But the two most consequential quantitative claims are unverifiable from the supplied material: the >$7B OpenRouter price is a one-line aside with no filing or announcement, and the 4x cost-effectiveness figure has no methodology and names its comparator inconsistently. Cost multiples are conversational ranges. No independent, customer, or competitor voice appears.
Real enterprise deployment, self-reported
Adoption is the strongest dimension: two named large enterprises with quantified or scope-defined rollouts (Zillow at 80% of 7,000 employees; Booking.com company-wide), plus a $300M ARR figure tripling over 15 months and a shipped architectural component (Waldo). This is production usage, not pilots. It is discounted because every data point is relayed through the vendor, 'adoption' is undefined beyond headcount percentage, and no churn, seat-utilization, or routing-mode mix numbers are given.
Overstated: framing outruns verification
Positive gap. The story's framing devices — a >$7B router acquisition and a clean '4x more cost-effective' number — are the least evidenced items in the cluster, while the well-evidenced material (real deployments, configurable routing tiers, pre-LLM query assembly) is more modest than the framing implies. The dek's assertion that the routing layer is 'where enterprise AI budgets now get decided' is extrapolated from one vendor's account of its own customers' motives. The gap is moderate rather than severe because the underlying adoption is genuine and the cost-pressure problem is plainly real; the overstatement is in magnitude and generality, not in existence.
High: vendor-sourced promotional interview
The information chain is almost entirely interested. The primary source is Glean's CEO, the comparative benchmark comes from Glean's co-founder, and the customer adoption figures are relayed by Glean. Glean has direct commercial incentive to define enterprise AI value as residing in the routing/harness layer it owns, and reputational-financial incentive around ARR and valuation framing after a $150M Series F at $7.2B. The competitor named in the cost comparison was given no opportunity to respond. The publisher's format is an access interview, which structurally favors the subject's framing.
Moderate-low: direction credible, magnitudes not
Confidence is limited by structure rather than by internal contradiction. One publisher, one interview, and high incentive load mean I can be reasonably confident about what Glean says and what it ships, and about the existence of large-scale enterprise deployments, but not about the acquisition price, the efficiency multiple, or the generalizability of 'cost is the main driver' beyond Glean's book. The directional thesis — enterprises are routing to control spend, and open-weight cost pressure is newly relevant — is coherent and specific enough to hold provisionally.
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
build
A 12MB Go binary bets agent cost control is cache stickiness, not a dashboard1 distinct publisher
build
Your agent needs the API call, not the API key1 distinct publisher
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026