Build1 publisher3 min readPublished
Zenoh's put() returns before anyone can read it: a 3.9% stale-read rate in a tight loop
An Elixir developer measured about 78 stale reads in 2000 put-then-get iterations against the same Zenoh key. The write path acknowledges locally; the read path goes to the router.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author experimented with Zenoh via its Elixir bindings, Zenohex, using the put/get storage feature rather than the usual pub/sub use case, and found that every so often the state read back was one step behind.
- The reduced test loops 2000 iterations of Zenohex.Session.put followed immediately by Zenohex.Session.get on the same key, with a 3000 ms timeout and consolidation: :latest, comparing the returned payload to the one just written.
- Out of 2000 iterations, a small fraction printed "stale!": about 78 (3.9%) in one run.
- Querying again immediately afterwards almost always returns the correct value; the fastest the author measured was a single extra get about 1 ms later. The value does not disappear; there is a small window of lag before the write is visible.
- Zenohex.Session.put/4 is a thin Rustler wrapper around zenoh-rust's put; in the NIF, .wait() only waits for the local session to finish handing the message off (the local publish being queued), not for the remote zenohd router backing the storage to receive and apply it.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer writing up experiments with Zenoh through its Elixir bindings, Zenohex, ran 2000 iterations of put immediately followed by get on the same key, and about 78 of them, 3.9 percent in that run, read back the previous value [1][2][3]. That matters because the code under test was not doing pub/sub: it was using Zenoh's put/get as a small key-value store, which is exactly the usage that assumes a write is readable once the write call returns [1].
It is not data loss. According to the author, an extra get issued immediately afterwards almost always returns the correct value, and the fastest confirmation measured was a single additional get about 1 millisecond later [4]. The value arrives; it just is not visible yet when the caller is already asking for it.
The mechanism is visible in the binding. Zenohex.Session.put/4 is a thin Rustler wrapper over zenoh-rust's put, and the NIF's .wait() only waits for the local session to finish queueing the message, not for the zenohd router backing the storage to receive and apply it [5]. The read side is different in kind: session_get is registered as a DirtyIo NIF and genuinely blocks for a reply from the remote side within a timeout, a real request/response [6]. The author's framing is the one to keep: put behaves like a GenServer cast, get behaves like a call, and firing a cast then immediately making a call that depends on it is the classic shape of this race [7].
This is not a binding defect being blamed upstream. Eclipse Zenoh has an open design issue, #2511, titled "[Design] Acknowledged put: confirmed storage writes via query path vs protocol extension", still open as of the writeup [8]. The issue states the position plainly: the pub/sub path is fire-and-forget, and session.put() returns when the message is sent, not when it is stored [9].
Do the arithmetic on what a few percent means in an operational loop. At 3.9 percent, roughly one write in 26 is unreadable at the moment the next line of code looks for it [13]. In a service that writes state and then reads it back to make a decision, that is not an edge case you will find in a smoke test and will find in production logs, if you are comparing values at all. The failure is silent: the get succeeds, the reply is well-formed, the payload is one generation old.
The author's workaround is the honest one available today: a wrapper that puts, then polls get on the same key until the written payload reads back, retrying at a short interval until a deadline, returning :ok on confirmation and {:error, :not_confirmed} if it never lands [10]. Defaults are a 3000 ms confirm timeout, a 1 ms confirm interval, and a 3000 ms query timeout [11]. The price is at least one extra round trip per write, plus a poll loop on the unlucky percent [15].
Two things to watch. Whether #2511 lands as a protocol-level acknowledgement or as a blessed query-path confirmation decides whether every Zenoh client library needs its own version of this wrapper [8]. And measure your own rate before trusting 3.9 percent: that number is one run on one setup as reported in the post [3], and the post is an AI translation of the author's original Japanese article on Qiita [12].