Skip to content

Build1 publisher3 min readPublished

A bot timing test that catches schedulers, and the arithmetic showing where it breaks

Log2-bucketed gap entropy on live Nostr data separates two humans from a clockwork poster by 4x. Against a burst spammer, the same metric lands almost inside human range.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying A bot timing test that catches schedulers, and the arithmetic showing where it breaks
Generated illustration

What happened

  • A developer publishing as zekebuilds on dev.to tested whether a bot can be distinguished from a human purely by when it posts, ignoring content.
  • Method: pull a pubkey's recent kind-1 notes (about 45 events each), compute the gap in seconds between each consecutive pair, bucket each gap by its log2, then compute Shannon entropy over the bucket distribution in bits.
  • The author states that without log2 bucketing you are measuring jitter noise, and with it you are measuring how many orders of magnitude of timing behaviour an account uses; a 20-second gap and a 25-second gap land in the same bucket.
  • The measurements ran through nak against the relay wss://nos.lol, using real pubkeys and real notes, with no synthetic or simulated data.
  • Two active human pubkeys measured 4.13 bits and 3.60 bits of gap-entropy.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer publishing as zekebuilds pulled about 45 recent kind-1 notes from each of four Nostr pubkeys, discarded everything the accounts actually said, and scored only when they posted: inter-event gaps in seconds, bucketed by log2, then Shannon entropy over the bucket distribution [1][2]. Two human accounts came in at 4.13 and 3.60 bits, a scheduled bot at 1.000 bits flat, and a burst spammer at 2.036 bits, which is the number that matters [5][7][9].

The log2 step is the whole trick. Without it, according to the author, you are measuring jitter noise; with it, a 20-second gap and a 25-second gap collapse into the same bucket and what you are counting is how many orders of magnitude of timing behaviour an account uses [3]. All of it ran through nak against wss://nos.lol on real pubkeys, no synthetic data [4].

Bot A is the clean case. It posts a "Random Bitcoin Podcast Spotlight" template, all 45 notes pulled shared an identical prefix, median gap around 16 hours across a 528-hour span, entropy 1.000 bits over exactly two log2 buckets [6][7]. Two things follow from that arithmetically. A flat 1.000 bits is the maximum entropy available to two buckets, so the two are equally populated: 22 of the 44 gaps in each [3]. And 528 hours across 44 intervals is a mean gap of 12 hours, well under the 16-hour median [2]. This is not a metronome. It is a timer that alternates between two adjacent scales, and the metric still buries it, because 1.000 bits against a 3.60-bit human floor is a 3.6x separation [8].

Bot B is where the author is honest. It rotates product-ad templates, dumped 45 notes in roughly half an hour at a 21-second median gap, and scored 2.036 bits [9]. Against the two humans that is 1.77x and 2.03x, straddling the 2x line the author predicted going in [8][10]. Worse than that framing suggests: a 2x rule with a human floor of 3.60 bits implies a bot ceiling of 1.80 bits, and Bot B sits about 13 percent above it [5]. And 2.036 bits cannot come from "a couple" of buckets. Four buckets cap out at 2.000 bits, so Bot B occupied at least five [4]. Twenty-second jitter, compounded over 44 intervals, genuinely spans that much.

The aggregate reads well: mean bot 1.52 bits against mean human 3.87, a 2.55x ratio [11]. The means check out against the four figures, and that is the problem, because each mean is an average of two accounts [6][1]. Four accounts on one relay is a demonstration, not a calibration set.

The author's conclusion is the useful part: both bots were trivially obvious on content, Bot A through 45 identical prefixes and Bot B through a small rotating ad pool, so template similarity catches exactly the archetype timing misses, and vice versa [13]. Timing entropy is strong against fixed-schedule posters and weak against a burst spammer that randomises cadence [12]. The cost of evading the timing half is visibly low, since Bot B was not even trying.

Watch whether the 2x threshold survives contact with a real sample, and whether it needs to be conditioned on posting volume, since a 30-minute burst and a 528-hour span are being scored on the same scale. The author says gap-entropy is one candidate input to depth-of-identity scoring at identity.powforge.dev [14]; the number to ask for next is a false-positive rate on low-volume human accounts.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories