Skip to content

Build1 publisher3 min readPublished

Go's csv reader allocates by design: what a 4 KB-buffer replacement actually buys

A Go pipeline over 5 million rows generated 10M-plus allocations and 540 MB of garbage. The fix is not tuning encoding/csv, because its signature guarantees the allocations.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The reader method signature in Go's encoding/csv is func (r *Reader) Read() (record []string, err error).
  • Every call to encoding/csv Read() allocates a new []string slice for the row (runtime.makeslice) and copies each parsed byte slice into a newly allocated heap string (runtime.rawstring).
  • The author was profiling a batch pipeline processing a 5-million-row CSV export; the Go process was consuming hundreds of megabytes of RAM and the garbage collector was eating a noticeable chunk of CPU time.
  • On a heavy encoding/csv run, go tool pprof shows runtime.makeslice and runtime.rawstring dominating the allocation graph.
  • Reading a 5,000,000-row CSV file with 6 columns produces 10,000,000+ allocations from slice headers and string conversions and about 540 MB of cumulative heap allocations pushed onto the garbage collector.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A developer profiling a Go batch pipeline over a 5-million-row CSV export found the process holding hundreds of megabytes of RAM with the garbage collector taking a noticeable share of CPU, and has published go-zerocsv, a zero-dependency reader and writer built on in-place typed scanning [3][6]. The part worth acting on is not the speedup; it is that the memory profile of `encoding/csv` follows from its signature, so it is not something upstream tuning can remove [1].

Read the promise: `func (r *Reader) Read() (record []string, err error)` [1]. To honour it, every call allocates a fresh `[]string` for the row and copies each parsed field into a newly allocated heap string [2]. Under pprof, `runtime.makeslice` and `runtime.rawstring` dominate the allocation graph [4]. On a 5,000,000-row, 6-column file the author counts more than 10 million allocations and about 540 MB of cumulative heap allocation pushed at the collector [5]. That count reads conservative: one slice header plus six strings per row across five million rows is 35 million allocations, so 10 million is a floor rather than an estimate [13]. In bytes it is roughly 108 of garbage per row, trivial once and expensive five million times [11].

go-zerocsv keeps a single internal 4 KB buffer, compacts it between records and reuses it, which is the mechanism behind the reported flat ~5 KB of heap whether you stream 1,000 rows or 10,000,000 [7][6]. `Read` returns a lightweight `Record`, and `Scan(&id, &name, &score, &active)` parses into your own variables in the pattern `database/sql` users already have muscle memory for [8]. There is an `IsFirst()` helper for skipping a header row and `String(i)` for direct field access [9]. The write path has the mirror-image problem: because `Writer.Write` accepts only `[]string`, emitting an int, float64, bool or time.Time through the standard library costs a `strconv` or `fmt.Sprintf` call per cell and a throwaway string with it [10].

The measurement offered is 5 million rows parsed in about 300 ms while holding 5 KB, on an AMD Ryzen 5 8400F, linux/amd64, Go 1.26 [c11b]. That is roughly 16.7 million rows per second [12]. One machine, one file, the author's own library: a direction, not a specification.

Two cautions before swapping it into a pipeline. The 540 MB and 5 KB numbers do not measure the same quantity; the source describes the first as cumulative heap allocation across the run and the second as steady-state usage [5][6]. Cumulative allocation is what the collector has to chase, not resident set, so the defensible claim is a large cut in GC work, and the 108,000x ratio between the two figures is arithmetic rather than an RSS comparison [14]. Second, zero allocations per record while scanning into a `string` variable implies that string points into the buffer that is compacted and reused for the next row [7][6]. The source's list of alternatives to `Scan` is cut off mid-sentence at copying field bytes [15], which is exactly the API whose lifetime rules decide whether you can retain a scanned string past the next `Read`.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories