Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Pennyforge's scan finds steering text in tool descriptions at 8 of the 20 most-starred MCP servers

Pennyforge's dated scan of the 20 most-starred MCP servers found 41 behavior-steering phrases in tool descriptions across eight repositories. Teams wiring these servers into agents now have quoted lines and a repeatable method to hold up against the NSA's May warning on MCP security.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Pennyforge's scan finds steering text in tool descriptions at 8 of the 20 most-starred MCP servers
Generated illustration

What happened

  • Pennyforge's analyzer pattern-matched every source file on five signals: instruction phrases, permission claims, external URLs, protocol version and last-push date.
  • Ten weeks after the breaking 2026-07-28 spec revision, none of the 20 repositories showed a clean signal that it had migrated.
  • Two of the original top-20 repositories returned 404s on scan day and were replaced by alternates with 1.1k and 486 stars.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Agents read tool descriptions whether or not a human does, so a server's author can steer which tool the model picks, and a review that skips description text misses that channel.
  • exposure A team running n8n-mcp on a host whose environment holds API tokens is one diagnostic call away from putting those tokens into the model's context.
  • decision A client written only for the 2026-07-28 revision cannot confirm that any top-20 server will answer it, so client teams have to keep the old handshake path for now.
  • constraint The counts describe 20 repositories picked by stars on one day. They tell a team nothing about its own exposure until its own servers get the same scan.

The F1 signal is a phrase match. It flags strings such as "you must", "always", "prefer" and "call this before" inside tool descriptions [3]. Pennyforge files these under the "tool poisoning" family: behavior-steering text that agents read but humans usually don't [3]. The plainest examples come from large servers. Upstash's context7, at 62.8k stars, tells the model "You MUST call this function before..." [8]. ChromeDevTools' chrome-devtools-mcp, at 53.1k stars, says "ALWAYS prefer this tool over multiple individual..." [8].

A phrase count overstates steering wherever "always" is descriptive. It also misses steering written in words outside the pattern list. n8n-mcp made the list with a description that begins "Always available" and goes on to list node info, search and validation [12]. The only thing that line steers is the match count. Pennyforge says a match can be a false positive, and that none of its quoted examples are obviously malicious [4][9].

The permission findings are where I would spend review time. The reference repository's environment tool "Returns all environment variables, helpful for debugging" [10]. n8n-mcp's diagnostic tool offers to "Include detailed debug information including full environment variables and API response details" [11]. Pennyforge takes that to mean one diagnostic call can return the environment to the model. It says it flags the design choice without judging it [11].

The 2026-07-28 revision removed the initialize handshake and Mcp-Session-Id sessions [13]. Each request now declares its own protocol version and capabilities inside _meta, and a client's first call goes to server/discover [13]. The report's "ten weeks" is accurate: the scan came 71 days after the revision [19]. Eleven repos show only old-spec markers, four show a mix, and five are indeterminate because the extractor found SDK usage but no direct RPC code [14]. All 15 repos with readable RPC code still carry old-spec markers, and the 91k-star official reference repo is one of them [18][14]. For the five SDK-only servers, the spec state presumably follows whichever SDK version they pin. Migrating them would then be a dependency bump [14].

The report's own dates contradict its maintenance figure for tadata-org/fastapi_mcp. Pennyforge says the 12k-star repo has not been pushed since 2025-11-24 and calls that 14 months [1]. From that date to the 2026-10-07 scan is 317 days, about ten and a half months [17]. The repo is stale on either count. The substitute merajmehrabi/puppeteer-mcp-server was last pushed on 2025-03-14 [16].

The NSA said in its May 2026 paper that the protocol's "rapid proliferation has outpaced the development of its security model" [5]. Pennyforge says it thinks that is fair [5]. The official SDK's 242 million npm downloads last month and the 33,674 GitHub repositories tagged mcp-server measure the proliferation [6]. The scan gives the security half evidence a reviewer can check line by line. Another team can copy the method: five signals, pattern matches on source, with the analyzer and every match kept for inspection [3][4]. A shallow clone fixes a date but not a commit, so a rerun next month scans whatever is at HEAD by then. I think the format is right for this job. It dates every finding and states its limits before its results [4].

What to watch

  • A 2026-07-28 implementation landing in modelcontextprotocol/servers, which would give new-spec clients their first confirmed target among the top 20.
  • Changes to n8n-mcp's diagnostic tool or the reference environment tool that redact secrets or make the environment dump opt-in.
  • A rerun of Pennyforge's scan against pinned commits, which would let other teams reproduce the match counts exactly.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption65
Hype gap+5
Incentives
Insufficient
Confidence55
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    tadata-org/fastapi_mcp (12k stars) hasn't been pushed since 2025-11-24, which the report describes as 14 months.

  2. [2]

    On 2026-10-07 Pennyforge shallow-cloned 20 MCP server repositories: the official/reference servers plus the highest-starred third-party servers on the GitHub mcp-server topic.

    ReportedSupportedSource: Pennyforge dated evidence report, dev.toView cited source
  3. [3]

    Pennyforge ran a local static analyzer over every source file with five heuristic signals: F1 instructions in tool descriptions (patterns like "you must", "always", "prefer", "call this before"), described as the "tool poisoning" family of behavior-steering text that agents read but humans usually don't; F2 permission claims; F3 URL inventory; F4 spec state; F5 maintenance (last-push date).

    ReportedSupportedSource: PennyforgeView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 7, 2026

    The 20 most popular MCP servers, scanned: what's actually in their tool descriptions

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Entities

Loading related stories