Skip to content

Build1 publisher3 min readPublished

An OAuth login now lets Claude rewrite, or delete, your live ElevenLabs voice agent

The hosted MCP connector turns model-cost comparison into a chat query, and turns change management into a confirmation screen that reports approval rather than correctness.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • On Monday, ElevenLabs introduced a hosted MCP connector allowing Claude read and write access to chat agents built with ElevenAgents; the new server focuses on managing agents that already exist in an ElevenLabs workspace.
  • In addition to creating agents and comparing configurations, the connector can calculate expected LLM usage and cost before a change is applied.
  • The example query given is: "What would my checkout agent cost per conversation with Gemini 2.5 Flash instead of GPT-4o?"
  • The New Stack writes that this type of question is even more important now that cheaper models alone do not guarantee lower bills once agentic orchestration is involved.
  • According to ElevenLabs, Claude can update an agent's system prompt, language, voice and opening message; retrieve transcripts; explore conversation topics; check the size of a knowledge base; generate sample speech; and delete an agent.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

ElevenLabs introduced a hosted MCP connector on Monday that gives Claude read and write access to chat agents already built in an ElevenLabs workspace [1]. That moves production agent configuration out of the dashboard and into a chat session, which changes both who can make a change and what counts as evidence that the change was correct.

The read side is the easy sell. The connector can calculate expected LLM usage and cost before a change is applied [2], and the worked example is a developer asking what a checkout agent would cost per conversation on Gemini 2.5 Flash instead of GPT-4o [3]. The New Stack argues that question matters more now, because cheaper models alone do not guarantee lower bills once agentic orchestration is involved [4].

The write side is broader. According to ElevenLabs, Claude can update an agent's system prompt, language, voice and opening message, retrieve transcripts, explore conversation topics, check knowledge base size, generate sample speech, and delete an agent [5]. Access arrives through Claude's directory and an OAuth sign-in, so there is no local server to run and no API key pasted into the client, with access limited to the workspace and to the permissions approved at login [6]. That is a real improvement on the April 2025 open-source server, which developers ran locally with an ElevenLabs API key to generate speech, clone voices or transcribe audio [7].

Two layers of control sit on top: administrators can disable tools across an organisation, and users can set stricter limits for their own sessions [8]. ElevenLabs specifically flags agent deletion as destructive and recommends reviewing tool calls before approving them [9]. For comparison, The New Stack notes GoDaddy went with quote-then-execute, idempotency keys and a consent object tying every purchase to human approval, while AWS built Dogwood, a policy engine that evaluates whether a tool call is valid in context [10]; the publication places ElevenLabs' two-layer model between those approaches [11].

The interesting failure is not an unauthorised change, it is an authorised one. Trim a support agent's system prompt to save tokens, accidentally drop the line that escalates billing disputes to a human, and the shorter prompt still reads fine during review [12]. A confirmation screen records that the change was approved; it cannot show whether the prompt still works or whether a model switch altered the agent's tool calls, because a successful tool call proves the action ran, not that it produced the intended result [13]. The New Stack ties this to recent Claude containment incidents in which models completed tasks in ways that crossed operator-set boundaries [14]. Note also that one OAuth session spans both reading end-user transcripts and rewriting or deleting agent configuration [15], so untrusted text and privileged writes share a single context and the approval prompt is the only thing standing between them.

The mitigations exist and are worth wiring in. ElevenLabs ships an agent testing framework that simulates a conversation before deployment, including tool invocation, and those tests run through the CLI or API rather than by hand in the dashboard [16]. Opt-in versioning saves configuration changes on branches and routes a share of production traffic to them for gradual rollout or A/B testing, though ElevenLabs says versioning cannot be disabled once enabled [17]. The CLI can pull and push agents as code, with a documented CI/CD workflow that includes a dry run before deployment [18].

Watch whether teams default to disabling the destructive tools organisation-wide [8], whether the simulation tests become a mandatory gate in the CI/CD path instead of an optional dashboard exercise [16][18], and whether that one-way versioning switch turns out to be a cost anyone regrets [17].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories