Build1 publisher3 min readPublished
Your meter now runs on someone else's machine: signed receipts, fsync, and failing open
A dev.to walkthrough of metering embedded Go SDKs starts by assuming the host is hostile: Ed25519-signed snapshots, real fsync discipline, and an enforcement branch the excerpt never reaches.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Enterprise Go SDKs embedded in a customer's binary and shipped inside their infrastructure face a metering problem pure web services avoid: the process being counted runs inside a host the vendor does not control.
- Such an SDK cannot simply read a counter in Redis or call home on every request; it must count accurately, resist tampering, survive network partitions, and still enforce limits without becoming a reliability liability for the customer's production stack.
- Once a .so or a statically linked Go binary is shipped, the calling process owns the address space, the file system, the clock, and the network.
- An adversarial operator can replace the metering goroutine's ticker with a patched clock, delete or replay the local persistence file tracking accumulated usage, firewall the reporting endpoint and wait for the grace window to expire, or fork the process at a known-low counter state and restore it after heavy usage.
- cgo makes some of these attacks marginally harder to script but introduces its own attack surface: symbol interposition, LD_PRELOAD, and ABI compatibility.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A post on dev.to about usage metering in Go SDKs starts from the premise most billing code skips: when your code ships inside a customer's binary, the process you are counting runs on a host you do not control [1]. That one fact moves metering out of the request path and into local accounting, because you cannot read a counter in Redis, you cannot call home on every request, and you still have to enforce limits without becoming a reliability liability inside your customer's production stack [2].
Once a .so or a statically linked Go binary is in someone else's process, that process owns the address space, the filesystem, the clock, and the network [3]. The author's inventory of operator attacks is the useful part: patch the ticker your metering goroutine reads, delete or replay the local persistence file, firewall the reporting endpoint and wait out the grace window, or fork the process at a known-low counter and restore it after the expensive work is done [4]. Packaging does not fix this. cgo makes some of these marginally harder to script while adding symbol interposition, LD_PRELOAD and ABI compatibility to your surface [5]; pure Go is easier to audit, build reproducibly and cross-compile, and also fully introspectable with go tool objdump and patchable at the binary level [6]. The choice relocates the risk rather than removing it [7]. The stated conclusion is to assume the local process is hostile and make integrity verifiable externally rather than asserted locally [8].
Mechanically that means hot-path counters in sync/atomic, where atomic.Int64 from Go 1.19 is cache-line aligned by the compiler inside a struct and avoids false sharing [9], which matters when hundreds of goroutines in a gRPC server's handlers are calling you [10]. Then periodic durable snapshots, every N calls or every T seconds [11]. The snapshot is the product, not the counter: calls, bytes, a UTC issue time, and a HostID derived from a SHA-256 of machine-id or EKS node identity, Ed25519-signed with a key injected at construction from a license blob and never written to disk in plaintext [12][13]. Rewinding the file to an older receipt buys nothing, because the attacker cannot re-sign the current host and timestamp [14].
The unglamorous half is durability. os.WriteFile does not fsync the parent directory, so the write needs an explicit Sync() on the descriptor before the rename and a directory Sync() after it [15][17]; without that, a power loss between write and rename leaves a zero-byte receipt file and the whole snapshot period is gone [16]. On Kubernetes the writable path has to be an emptyDir or a mounted PVC, not the container overlay, which may not preserve fsync ordering semantics across node evictions [18].
Two gaps worth holding in mind. Signing per flush bounds the fork-and-restore attack to one flush interval, since only usage accumulated after the last sealed receipt can be hidden [1]. And the receipt shown carries no sequence number or previous-receipt hash, so verification detects a modified receipt but not one that simply never arrived [2].
What to watch is the enforcement branch, which the supplied text does not reach: it stops mid-sentence after naming VPC firewall rules, proxy misconfiguration and a transient AWS PrivateLink outage as things that will block your reporting endpoint [19][20]. Fail closed and a firewall rule turns your license check into your customer's incident; fail open and that same firewall rule is the cheapest attack on the list [3].