Skip to content

Invest1 publisher3 min readPublished

A ClickFix technique beat Meta's Muse safeguards within 13 days of launch

Meta's bug bounty puts $130,000 on a prompt injection against Muse out of a $300,000 top payout. A researcher has now demonstrated a working hijack of the agent Meta sells as safe enough to complete purchases.

The Investor · Invest desk

What happened

  • A security researcher has demonstrated a working zero-day capable of hijacking the assistant, according to TechTimes citing ArsTechnica.
  • The exploit uses a ClickFix style technique, the social engineering method that gets a user or an automated system to run a malicious command dressed up as a routine fix or verification step.
  • Meta's bug bounty pays up to $300,000 for a validated vulnerability, with $130,000 earmarked for a successful demonstration of a prompt injection attack.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • exposure Meta routes money-moving approvals to the user in the app, and ClickFix works by making a malicious step look like a routine confirmation, so the person who is the final control on a payment is also the target.
  • contradiction Meta's security documentation concedes prompt injection is unsolved industry-wide while the company sells the same agent as safe for inboxes and payments, so the marketing and the documentation cannot both be true.
  • constraint Because the agent's job includes reading and following instructions on web pages, fixing one exploit does not close the class, and the isolation boundary has to hold against every page Muse visits.
  • cost Meta has reserved 43 percent of the top payout for a single attack class. That is what it will pay a researcher for the disclosure.

Meta published a price for this failure mode. Its public bug bounty for Muse pays as much as $300,000 for a validated vulnerability. Of that, $130,000, about 43 percent of the ceiling, is set aside for one attack class: a demonstrated prompt injection [11][12][2]. Meta's own security documentation says Muse is not immune to attack and that prompt injection remains an open problem across the AI industry [10].

Each Muse instance runs inside its own cloud-based virtual machine, which Meta calls a Muse Secure VM [6]. A separate oversight system called Sentinel is what Meta describes as the sole permission authority over connected services and all outbound internet traffic [7]. Sentinel substitutes surrogate tokens at the network boundary, so the agent itself never sees actual passwords or payment details [8]. On that design, a hijacked agent is an agent operating inside permissions Sentinel has already granted. The report does not say what the researcher's exploit made Muse do, or whether a fix has shipped [3].

Meta says any critical approval, such as sending money or confirming a purchase, is surfaced to the user through the app instead of being something Muse can grant on its own [9]. ClickFix, as TechTimes describes it, tricks a user or an automated system into running a malicious command disguised as a routine fix or verification step [4]. The last control on a payment is therefore a person approving a prompt, and the method in question is the one built to make a malicious step look like the routine one. It reaches Muse at all because the agent browses the web and follows instructions it finds on a page [5].

The purchase capability shipped on day one. Muse launched on September 8 able to send emails, book travel, fill out forms and complete purchases on a person's behalf [1]. It is built on the Muse Spark 1.3 model under chief AI officer Alexandr Wang [2]. Mark Zuckerberg has framed it as a major step toward what he calls "personal superintelligence" [13]. The working zero-day was reported on September 21, 13 days later [15][1].

One reading is benign. If the token boundary held and the hijack stayed inside permissions Sentinel already allowed, a bounty program did what it was built for and the $130,000 line item was cheap. A hijack that reached credentials or moved money is the second case. It puts inboxes and card details at stake on a consumer product less than a fortnight old [1]. The third is slower and more expensive. Meta patches this instance, the class stays open by the company's own account [10], and each later disclosure lands against a firm still marketing the assistant as safe enough to manage an inbox and complete payments [14]. I'd expect the second and third to be the same story on different timelines. A technical account showing that surrogate tokens never left Sentinel would change that.

What to watch

  • Whether Meta pays the $130,000 prompt-injection award and publishes a technical account of what the hijacked agent could reach.
  • Whether the researcher's write-up shows surrogate tokens leaving Sentinel or staying behind the network boundary.
  • Whether Meta keeps purchase completion enabled for new Muse users while it describes prompt injection as unsolved.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories