BuildNot yet confirmed elsewhere1 publisher3 min readPublished
Object or file storage: the wrong pick stays quiet until inode exhaustion
A dev.to writeup reduces the storage choice to a property test on the calling code. The part worth arguing about is when each wrong answer shows up: on the first query, or at 10 million files.
The Engineer · Build desk
What happened
- Its worked example is user photos on an ext4 volume, where the server falls over at 10 million files and inode exhaustion is the named mechanism.
- Databases are the stated exclusion: object storage has millisecond latency and no in-place mutation, so relational engines on S3 perform badly outside Iceberg and Delta.
- The mount-based escape route is partial. The author says mount layers misreport in small ways, and RustFS documents no FUSE driver.
Why it matters
- constraint Keeping the option open does not work. Two of the four filesystem properties the post says databases need are the two object storage lacks, so "we will move the database onto the bucket later" is...
- cost The correction is billed to the filesystem side, since the object route needed no code change across four orders of magnitude.
- decision The asymmetry in detection time argues for spending the design hour on the storage question before the first upload, because only one of the two errors will announce itself while anyone is still...
- exposure Teams planning to bridge the gap with a mount take on whatever the mount layer gets wrong, and at least one S3-compatible server does not offer that layer at all.
The useful part of this comparison is not the requirement lists, it is when each mistake becomes visible. Put a database on object storage and the error reports itself early: the dev.to writeup by Ethan Carter puts object latency at millisecond level with no in-place mutation [6], against the sub-millisecond random reads, in-place byte updates, ordered fsync and lock files that PostgreSQL, MySQL, MongoDB and SQLite expect from a filesystem or block device [7]. Two of those four properties are precisely the ones object storage does not have [20]. Put user photos on an ext4 volume instead and nothing reports anything at all until inodes run out [1]. One error costs you a slow afternoon in development. The other costs an afternoon two years later, on a volume somebody else provisioned.
The two anecdotes in the post are the same workload an order of magnitude apart. The filesystem falls over around 10 million files [21]; the bucket holding the same photos reached 100 million [22], a 10x gap on identical data [23]. And the object path got there without a code change [10], which tells you where the migration work lives. It is not on the side that scaled.
So the test runs against the caller, not the data. The operating system opens, reads, writes and seeks; it does not issue HTTP PUTs, and Carter's advice is not to fight that [8]. His own shorthand is that random writes and file locks mean file storage, while HTTP PUTs and billions of objects mean object storage [9]. Uploads, data lake Parquet, backups and ML checkpoints all sit on the object side of his list [4], and none of them call seek(). Shared design files over SMB, build trees over NFS and enterprise home directories sit on the other side because they need directory browsing, file-level permissions and application transparency [16].
He also says most mature stacks run both side by side [5], which makes this a per-path decision rather than a per-company one. The lakehouse exception fits that reading: Iceberg and Delta are named as the niche where things do work over object storage [6], and that is a table format built for object semantics, not a database talked into them.
The hedge is a mount, and the writeup is careful about it. You can get filesystem semantics over S3-style storage, Carter says, but the mount layer will lie to you in small ways [13]. RustFS, whose documented feature table covers S3 core, versioning, bucket replication, event notifications and bitrot protection, ships no FUSE driver at all [14]. If your plan for the POSIX requirement you have not fully specified yet is "we will mount it", that plan depends on a component the server you picked may not offer.
This is one practitioner's checklist rather than a benchmark, and the only latency figures on offer are "sub-millisecond" and "millisecond-level" [3][6]. It is still enough, because the answer does not turn on how many milliseconds. It turns on whether anything in the path mutates bytes in place.
What to watch
- Whether the cut-off project-by-project section of the post names any S3-compatible server that documents a supported FUSE driver, rather than omitting one.
- Published latency or throughput figures for lakehouse table formats against POSIX filesystems, which would test the "niche exception" framing.
- Any measured file count at which common filesystems actually degrade, instead of the round 10M in the anecdote.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+18
- Incentives34
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The author says he has watched teams try to stretch a filesystem to 100M objects and that inode exhaustion is not a fun afternoon.
- [2]
The post says terabyte-scale ML workloads are almost always object-storage-backed in 2026, covering training data, checkpoints, immutable dataset versions and Parquet feature stores queried via the S3 API.
- [3]
The post's short answer: use file storage when you need POSIX semantics, which it defines as in-place edits, sub-millisecond random I/O and file locking (databases, OS files, NFS shares).
- [4]
The post says to use object storage when you need scale, an HTTP API and rich metadata, listing user uploads, data lakes, backups and ML datasets as the workloads.
- [5]
The post states that most mature stacks run both storage paradigms side by side.
- [6]
The post says object storage has millisecond-level latency and no in-place mutation, and that databases on S3 perform terribly with niche exceptions like Iceberg/Delta lakehouse patterns.
- [7]
The post says PostgreSQL, MySQL, MongoDB and SQLite all expect block devices or file systems with sub-millisecond random I/O, in-place mutation, strong consistency via ordered fsync, and file locking.
- [8]
The post notes the OS uses open(), read(), write() and seek() rather than HTTP PUT/GET, and advises not fighting this.
- [9]
The author's checklist: random writes and file locks mean file storage; HTTP PUTs and billions of objects mean object storage.
- [10]
The post says object storage lets you scale from 1K to 100M objects without changing code.
- [11]
The post puts Parquet/Avro/CSV data lakes accessed by Spark, Trino and DuckDB on the object side, citing a flat namespace and engines that speak S3 natively.
- [12]
For backups, the post cites versioning, lifecycle tiering, cross-region replication and Object Lock WORM retention as object storage features that save cost and operational effort at scale.
- [13]
Asked whether you can have S3 scale with filesystem semantics, the author answers that you can, but the mount layer will lie to you in small ways.
- [14]
The post states RustFS does not ship a FUSE driver, and that its README feature and status table covers S3 core, versioning, bucket replication, event notifications and bitrot protection.
- [15]
For small datasets the post calls file storage simpler: no API to learn, familiar tools such as ls, cp, rsync and grep, easy backups with tar and rsync, and local access as the fastest possible.
- [16]
The post lists SMB-mounted design files, NFS-shared build source and enterprise home directories as file-storage cases needing file-level permissions, directory browsing and application transparency.
- [17]
The post calls S3 plus CloudFront or Cloudflare the standard pattern for static web content, and says serving a global static site from an NFS mount is not something the author would want to be on-call for.
- [18]
The two wrong answers have different detection times: a database on object storage misbehaves as soon as it runs, while a filesystem holding tens of millions of objects gives no signal until inode exhaustion.
- [19]
Because the object path scales from 1K to 100M objects without code changes, the code-change cost of a late correction falls on the team that chose the filesystem.
- [20]
Two of the four filesystem properties the post says databases require, sub-millisecond random I/O and in-place mutation, are directly negated by the properties it attributes to object storage.
- [21]
The dev.to writeup describes engineers storing user-uploaded photos in /var/www/uploads on an ext4 volume and wondering why their server falls over at 10M files.
ReportedInsufficientSource: dev.to, Ethan Carter2 sources— create a free account to open themView cited source - [22]
The author says a neighbouring team put the same photos into an S3 bucket and scaled to 100M files.
ReportedInsufficientSource: dev.to, Ethan Carter2 sources— create a free account to open themView cited source - [23]
The post's two anecdotes about the same photo workload sit a factor of 10 apart: failure around 10M files on ext4 versus 100M files in a bucket.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toObject Storage vs File Storage: When to Use Which (2026)
1 article · August 23, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- RustFSFollow
- Amazon Simple Storage ServiceFollow
- s3fs-fuseFollow
- PostgreSQLFollow
- MySQLFollow
- MongoDBFollow
- SQLiteFollow
- Apache SparkFollow
- TrinoFollow
- DuckDBFollow
- Apache IcebergFollow
- Delta LakeFollow
- Amazon CloudFrontFollow
- CloudflareFollow
- ext4Follow
- NFSFollow
- SMBFollow
- Apache ParquetFollow
- POSIXFollow
- S3 Object LockFollow