Build1 publisher2 min readPublished
Claude Code traced a JNI cache leak by diffing core dumps taken 30 minutes apart
Claude Code compared gcore snapshots taken 30 minutes apart and flagged a native C cache whose entry count never dropped, a lead QA engineer wrote. The method fits servers too loaded to instrument live, since it works from frozen snapshots of the process.
The Engineer · Build desk

What happened
- Grafana showed the Java module's heap in a stable sawtooth while the C server's resident memory climbed linearly for hours, placing the leak on the native side.
- An earlier attempt by Claude at intrusive runtime monitoring accidentally crashed the high-throughput server and forced the team to change strategy.
- Subject identifiers pulled from the growing cache matched subjects for which the test client had already sent discard requests.
- Working as a log-sifting agent, Claude isolated a subject with a discard receipt in older logs that still occupied memory in the latest core dump.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure An agent with access to a live process can crash it. This crash hit a replica VM; the same permissions on a production server would have turned a diagnostic step into an outage.
- cost Agent output still costs reviewer time: the static scan's leaks were real defects, and it took a developer to establish they had nothing to do with the memory growth.
- constraint The author's baseline of days of manual pointer tracing is not set against an elapsed time for this investigation, so the post shows the method worked without showing how much time it saved.
According to the post, Claude drove GDB's gcore to write each snapshot without terminating the application [7]. It parsed the data structures in both dumps and compared the deltas [8]. The flagged structure was a cache holding the pivot results the Java module passed back for the 500,000-row grid, and each entry carried a subject name identifier [9]. Two snapshots always define a straight line, so the case for linear growth rests on the hours of RSS telemetry more than on the dumps [2].
The diff worked because this leak was easy to see in a dump. The leaked objects were instances of one structure type, and their count only went up. For the method to work on another system, the same has to hold there. The tooling has to be able to read the dump as typed structures. A leak of raw buffers with no owning structure would not show up as one type's rising count. Nor would growth from heap fragmentation. I'd expect a two-dump comparison to miss both.
The server under test was a performance-critical C process hosting a Java module over a JNI bridge [1]. A Java test application simulated hundreds of concurrent users through a proprietary SDK, while a datasource pushed random real-time updates across the 500,000-row grid [1]. Under that load, Valgrind's overhead stalled the environment almost at once [5]. The author, a lead QA engineer [17], wrote: "If we couldn't hook into the process live or use heavy runtime instrumentation, we had to look at what was accumulating by comparing frozen states in time instead." [6] The team picked the method. In my view the agent's share was the slow work that came after: reading structures out of dumps and chasing one identifier through the logs.
The logs were the heavier job. Manual grepping was out of the question, the author wrote, because the loaded environment produced hundreds of megabytes of logs every minute under aggressive rotation [11]. At a floor of 100MB a minute, one 30-minute dump interval is at least 3GB of log text [1].
The subject Claude isolated gave a timeline [13]. The client sent a DISCARD request for Subject_A. The C server processed it, deleted Subject_A from its native cache and forwarded the message to the Java module. In the same millisecond, before the Java module registered the discard, the module was doing concurrent work of its own [13]. The author wrote that the timeline revealed "a lethal concurrency anomaly" [14]. The post's opening calls the bug "a highly elusive race condition" [15].
What to watch
- The rest of the post's timeline and the fix for the discard race, to confirm what the Java module did in that same millisecond.
- Attempts to use the same gcore diff on leaks with no single countable owning structure, such as raw buffer growth, where a type-count comparison has less to find.
- Wall-clock times for dump capture and agent analysis on a comparable load, the figure needed to test the claimed saving over manual pointer tracing.