Build1 publisher3 min readPublished
Rust MCP servers: the case is resident memory at 16 per box, not throughput
An rmcp walk-through concedes its tools are I/O-bound at 100-500ms per AWS call, then makes the rewrite argument on footprint, dependency isolation and schema drift instead.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A dev.to tutorial walks through building an MCP server in Rust with rmcp, the official Model Context Protocol Rust SDK, using as its example a devops agent that manages AWS EC2 G5g instances serving Gemma 4 under vLLM; a Python version of the same server already exists.
- The server's tools are I/O bound: every one is an AWS API call (describe_instances, send_command, polling SSM), so 100-500 ms of network per call.
- The author states that the caller's language contributes nothing measurable to that latency, and that anyone selling a Rust rewrite on raw speed for this workload is selling something.
- The monorepo contains 16 rigs, each with its own MCP server.
- The author characterises the Python cost as "a gigabyte of resident Python to expose sixteen tool lists is a real cost".
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A dev.to walk-through of building an MCP server in Rust with rmcp, the official Model Context Protocol Rust SDK, opens by taking apart the usual reason for doing it [1]. Every tool in the server is an AWS API call - describe_instances, send_command, polling SSM - at 100 to 500 ms of network per call, and the author's position is that the caller's language contributes nothing measurable there, so anyone selling a Rust rewrite on raw speed for this workload is selling something [2][3].
What is left after that concession is the part worth reading. The monorepo in question is not one server: it is 16 rigs, each with its own MCP server [4]. The author puts the cost as a gigabyte of resident Python to expose sixteen tool lists [5], which is roughly 64 MB per server before any tool is called [6]. That is the load-bearing claim, and the supplied text does not close it: there is a figure for the Python side and none for the Rust binaries [7]. Direction without a delta.
The second argument is packaging, and it is more concrete. These rigs install system-wide, no virtualenvs by policy, so all sixteen share one interpreter [8]. Sixteen servers with independently drifting boto3 and mcp pins in a single Python is a standing conflict risk, according to the author [9]; a static binary has no such coupling, and each rig pins what it likes in its own Cargo.lock [10]. Note what is actually forcing this. The constraint is an install policy, not a language deficiency, and a reader with virtualenvs or uv-managed environments has already bought most of that benefit.
Third: schemars generates the tool schema from the same struct the handler destructures, so the advertised schema cannot drift from the code [11]. The author expects that one to survive longest [12], and it is the only one of the three that is a property of the type system rather than of the deployment.
Startup gets a mention because MCP transport is usually stdio, with the client spawning your binary and talking over stdin and stdout [13], which makes process startup user-visible when a session spawns a fresh process [14]. The Rust server is quoted at 2.5 ms [15]. Against one AWS call that is 0.5 to 2.5 percent of the wait [16] - real, and not the headline.
The design underneath is the reason any of this is a server rather than shell scripts, in the author's own framing [17]: nine tools covering list, start, stop, terminate, endpoint, run_remote and health [18], driving a g5g.4xlarge Graviton2 box with an NVIDIA T4G that serves Gemma 4 E2B under vLLM on port 8000 [19], with no inbound SSH, no key pair and no port 22 rule, everything going through the EC2 API and SSM Run Command under an instance profile carrying AmazonSSMManagedInstanceCore [20].
Two things to watch. First, whether anyone publishes measured resident-set numbers per server for both stacks, because the fleet argument stands or falls on that and currently rests on one round figure [5][7]. Second, the rmcp packaging trap the article flags: a bare cargo add rmcp compiles and gives you almost nothing, with server, macros and transport-io needed explicitly and client, auth, elicitation and the streamable-HTTP server transport off by default [21][22]. The author's own limit is the honest one: if you have one MCP server and it works, this is not a reason to rewrite it [23].