Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

Spanner Omni runs Google's database on customer-run infrastructure without an availability SLA

Google made Spanner Omni generally available for customer data centers, other clouds and laptops, without an availability SLA. Teams that run it take over the failure domains, upgrades and monitoring that Google handles for the managed service.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Spanner Omni runs Google's database on customer-run infrastructure without an availability SLA
Generated illustration

What happened

  • In place of Colossus, Spanner Omni writes to each node's attached local file systems and shares them across the network, splitting and rebalancing shards automatically.
  • A software time service replaces TrueTime's atomic clocks and GPS and still provides error-bounded clock synchronization across servers.
  • Google's documentation offers four reference topologies for high availability, and on the single-server layout, upgrades cause downtime.
  • Integrations that depend on Google Cloud, including BigQuery, Knowledge Catalog and Gemini Enterprise, are excluded, and Google has not said when other feature gaps close.
  • The documentation lists Google Cloud and Amazon as supported public clouds alongside on-premises servers and laptops, and Azure is not mentioned.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Running Spanner becomes a staffing line: upgrades, routine maintenance and Prometheus-and-Grafana monitoring are hours the customer's own team now spends.
  • exposure A team that moves Spanner data into a private datacenter for residency also takes on the backup, key management and audit duties for it.
  • decision Teams whose Spanner data feeds BigQuery or Gemini Enterprise have to keep the managed service or rebuild those links before moving to Omni.

Spanner can run on a software clock because of how it already spends its waiting time. Google says the database does other work while it waits out clock uncertainty, so it can get by with looser uncertainty bounds than TrueTime actually achieves [5]. On that account, much of the atomic clocks' precision was margin Spanner did not use. Google spent that margin to keep time across heterogeneous hardware without, it says, limiting availability or performance [5]. This is good engineering. Paxos consensus, automatic sharding and synchronous replication carry over unchanged [6].

Google is just as direct about storage. The new file layer is not Colossus, the company says, but a sufficient stand-in to perform comparably to the managed service for most workloads [3]. Its internal benchmarks claim millions of queries per second across petabytes in a single regional deployment [7]. That is Google's number from Google's own runs [7]. For it to transfer, the software clock's error bound on your servers has to stay inside what the overlapped waits absorb. Your workload also has to be one of the "most workloads" in Google's caveat [3].

Carlos Pérez Martín, CTO at Q2BSTUDIO, put the operating consequence plainly. "The interesting shift is operational, not architectural: once the same engine runs in your racks, the failure domains become yours," he said [8]. His first example is topology. Quorum and witness placement has to be re-derived for the local latency budget, he said, and when a whole datacenter is the unit of failure, p99 latency, not the mean, is what the application feels [9].

Google's documentation reaches the same place. It gives customer-managed infrastructure as the reason Google offers no availability SLA, and says its reference architectures help achieve comparable high availability [12]. A reference architecture does not issue service credits. The multi-zone design needs at least three zones with three servers each, so nine servers is the minimum for that layout [19].

With general availability come audit logging, backup and restore, TLS encryption, and authentication and authorization, plus worker nodes: stateless compute that handles background operations so the primary servers don't have to [18]. The query languages carry over: GoogleSQL, PostgreSQL and Spanner Graph Language [17].

I think the pilot Pérez Martín proposes is the right first test, and I would add the software clock's measured error bound under production load to what it records [11]. "A like-for-like pilot against the managed service, measured on tail latency and ops toil instead of feature parity, is the cheapest way to price that trade," he said [11].

What to watch

  • Whether Google names the remaining feature gaps against managed Spanner and puts dates on closing them.
  • Independent measurements of the software clock's error bound and p99 latency on non-Google hardware.
  • Whether Azure joins Google Cloud and Amazon on the supported-cloud list.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories