Build1 publisher3 min readPublished
Nightly archive of auditd logs kept 234 of 696 hours on a busy database server
auditd's default 40 MB ring held 13.5 hours of events on one Postgres server, so a nightly copy built for Israel's 24-month rule saved about a third of each day. The loss dated from the archive's first week and surfaced only when its author checked coverage day by day.
The Engineer · Build desk

What happened
- A disk alert on September 6 led back to auditd's syslog plugin, which had grown auth.log to 763 MB in five days on a server where it normally adds 0.7 MB a week.
- The archive was a cron job added August 25 that forces a rotation at 02:17, gzips the rotated files and deletes archives older than 730 days.
- On the database server, 35,247 of 35,659 matched events came from a single watch on /opt/supabase, the directory holding the Postgres data volume.
- On the website server, late-evening deploys pushed the rest of the day out of the ring, and deploy days fell to about 10 percent coverage in the archive.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A size-capped ring paired with a once-a-night copy works only while the ring holds more than a day of events, and that holding time shrinks with every heavy deploy or newly watched write path.
- decision Operators watching an app directory that contains a database volume or a deploy target have to choose between narrower watch paths and a ring or copy schedule sized for that write volume.
- exposure Any fleet meeting a retention rule with auditd's default ring and a nightly copy has the same kind of hole on hosts whose ring turns over in under a day, until someone compares archived intervals against calendar days.
Israel's data security regulations require access logs to be kept for 24 months [18]. The auditd defaults in `/etc/audit/auditd.conf` are `max_log_file = 8`, `num_logs = 5` and `max_log_file_action = ROTATE` [3]. Five files of 8 MB make a 40 MB ring [4]. When it fills, auditd deletes the oldest file, and it has no concept of months [3]. Calendar retention has to come from outside. Here it came from one run a night, and anything that left the ring before that run was deleted before anyone saved it [20]. The author put the condition in one sentence: "A 40 MB ring is fine as long as 40 MB lasts longer than the gap between archive runs." [9]
How long 40 MB lasts depends on what the rules watch. The eight Ubuntu 24.04 hosts run auditd 3.1.2 with 22 to 25 rules each, mostly `-w <path> -p wa` watches on config files and app directories [2]. On the database server the app directory contains the data volume. Every file Postgres opens for writing becomes an event, and `postgres` was the executable behind 35,082 of them [6]. One write is also several lines. Each event is a SYSCALL record plus CWD, one or two PATH records and PROCTITLE, 5.4 records on average in one website deploy [8]. That deploy ran from 23:07 to 23:12 and produced 70,516 records [19]. On the voice server, the deploy rewrites about 30,000 files each run, changed or not [7].
Measured with `aureport --input-logs -t`, the database ring covered 13.5 hours [10]. Around the website deploy, three 8 MB files filled in 3 min 48 s, then 58 s, then 17 s [10]. Over 29 days the database server's archive held 234 of 696 hours [11]. That is about 34 percent, or roughly 8 hours a day [1][2]. According to the author, the syslog plugin switched on August 31 did not cause the gap [15].
Two parts of the build are good work. Each archive's filename carries a hash of its content, so a file that moves from `audit.log.1` to `audit.log.2` is not archived twice [5]. The coverage check is the right test: take each archived file's first record timestamp and its mtime, merge them into intervals, and compare the result with each calendar day [14].
Anyone scripting that check should keep `--input-logs`. When stdin is not a terminal, aureport reads stdin instead of the log files and prints `<no events of interest were found>` [16]. The author lost ten minutes to it [16].
These are one fleet's numbers. The 34 percent figure carries over only where two things hold: a watch covers a path with heavy writes, such as a database volume or a deploy target, and the copy runs less often than the ring turns over. The chatwoot host fails the first test. Its ring covers almost three days, and its only incomplete days were maintenance days: a hardening pass, a Docker update and a fleet-wide reboot [10][13].
The syslog copy behind the September 6 alert lands in `auth.log`, because Ubuntu's rsyslog routes `auth,authpriv.*` there, and logrotate keeps that file for 104 weeks [17]. 104 weeks is 728 days, two short of the archive's own 730-day expiry [3].
What to watch
- The author's figures for August 30, the day before the syslog plugin went on, which would show the size of the gap before forwarding began.
- Whether the central collector fed by the syslog plugin holds complete days, making it a usable 24-month record in place of the local archive.
- Coverage results for the two client servers the author kept out of the post.