Skip to content

Build1 publisher3 min readPublished

The 680 MB database that was really a 17 GB disk: self-hosted support platforms fail at month six

A Chatwoot operator's postmortem shows the meter you inherit watches Postgres while attachments pile up on a Docker volume that grows with tenant count, not traffic.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying The 680 MB database that was really a 17 GB disk: self-hosted support platforms fail at month six
Generated illustration

What happened

  • The author runs self-hosted Chatwoot as the WhatsApp inbox for a dozen or so small Israeli businesses, on two servers, handling a few thousand conversations a week, with a drip-sequence engine bolted on the side.
  • The author got a disk alert at 86 percent and immediately went looking at the database, which he says was the wrong place.
  • Measured sizes: Postgres database 680 MB; the chatwoot_storage_data Docker volume 17 GB.
  • Attachments live in ActiveStorage on a Docker volume, not in Postgres; every image, voice note and PDF a customer sends is a file on disk and none of it shows up when you check database size.
  • If monitoring watches the database, it will report that everything is fine right up until the container cannot write.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

An operator running self-hosted Chatwoot as the WhatsApp inbox for about a dozen small Israeli businesses got a disk alert at 86 percent and went looking at the database, which he says was the wrong place [s1 c1][2]. Postgres was 680 MB; the `chatwoot_storage_data` volume was 17 GB [3]. That is roughly twenty-five times the database, sitting in a place the default instinct never checks [1].

The reason is structural, not a bug. Attachments live in ActiveStorage on a Docker volume, so every image, voice note and PDF a customer sends is a file on disk and none of it appears in a database size query [4]. If your monitoring watches the DB, it reports healthy right up to the moment the container cannot write [5].

The accounting failure compounds the capacity failure. According to the author, the growth curve is a function of how many accounts you host, not how busy any one of them is [6]. His volume grew at roughly 0.05 GB a month until he onboarded seven new businesses over two months, after which it hit 16 GB a month [7]. That is a factor of about 320 in growth rate, triggered by sales, not by load [2]. Anyone forecasting storage from conversation volume will be wrong by two orders of magnitude on the month they close a deal.

Then there is what is actually on the volume. When he measured, almost half the outbound media was byte-identical duplicates [9]. One 14.5 MB video was stored 48 separate times; one image 325 times [10]. Chatwoot creates a new blob and a new file on every send even when the bytes match, which he calls correct behaviour for a chat app where every message owns its attachment, and expensive the moment anything fans one file out to many conversations [11]. In his case the fan-out was not campaigns but his own drip engine, re-reading and re-uploading the same file per conversation [12]. Forty-eight copies of a 14.5 MB video is 696 MB on disk to hold 14.5 MB of information [3].

He deduplicated after reading the Rails source, on the basis that `ActiveStorage::Blob#purge` is guarded by a foreign key on `active_storage_attachments.blob_id`, so removing one message's blob does not delete a file other messages still reference [13]. One pass matching on `checksum` and `byte_size` within the same account reclaimed 3.20 GB across 7,160 blobs and moved the disk from 86 percent to 71 percent [14]. That is about 19 percent of the volume recovered without deleting a single customer-visible message [4]. He recommends an index on `active_storage_blobs (checksum, byte_size)` and a size floor first, so you are not doing checksum lookups on 4 KB thumbnails [15].

The same measurement error shows up on the send path. `POST /api/v1/conversations/{id}/messages` returns 200 as soon as it inserts a row with `status = 0` and no `source_id`; actual delivery happens later in a Sidekiq job, and the WhatsApp message ID is written only when Meta acknowledges [16]. Four resend runs pushed 5,108 messages through that endpoint, the `high` queue reached 3,786 jobs with eleven minutes of latency, and inbound messages from paying customers queued behind them [17]. Nothing errored [18].

What to watch: whether your dashboard has a series for the storage volume at all, and whether your capacity plan is keyed to tenant count rather than message count [3][6]. His backpressure fix counts his own undelivered rows every 40 sends and pauses above 250 until the backlog drops under 80, which measures pressure he created rather than a global number he does not control [19]. And before any inbox deletion, note that his deleted inbox took 894 conversations and 6,160 messages with it, while contacts survived because they live at the account level [20].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories