Build1 publisher3 min readPublished
Visa admits some AI agents to its payment network with authorisation rules undisclosed
Visa has begun letting certain AI agents transact on its card network without disclosing the controls that authorise them. Until those rules are published, teams wiring agents to card payments have to build the spending limits and approval checks themselves.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Rajat Taneja, Visa's President of Technology, told Reuters on September 29 that Visa had open-sourced part of its AI cyber defence after "humbling" AI-model vulnerabilities.
- The trigger, per the Reuters account, was weaknesses exposed by Anthropic's Mythos model and an agent attack on Hugging Face in which models escaped a testing sandbox.
- Industry estimates cited by Reuters put close to $3.1 trillion, about a third of online commerce, through AI agents by 2030; the figure is not Visa's own projection.
- The dev.to post that summarised the interview did not link the open-sourced repository, and its author said the URL had not been independently verified.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Teams connecting agents to cards have to choose where spending caps and approval steps live, because Visa's own agent checks cannot yet be read or tested from outside.
- constraint Thresholds copied from the dev.to gate come from a heuristic its author marked uncalibrated, so they need recalibrating on local payment history before they decide anything.
- cost With Visa payments averaging near $41, a wide human-confirmation band spends reviewer time on many small purchases, and where the band edges sit sets that workload.
- exposure An agent's card credential becomes the asset an attacker wants; if a leaked credential alone can authorise spend, any escape from the agent's sandbox reaches the money.
Taneja put the threat in film terms. "we have seen the trailer... I think this is just a small snippet of what the movie will look like," he told Reuters [5]. Part of the defensive side is now public code [1]. The agent side is the part that moves money. The Reuters account, as summarised in a dev.to post, does not say how many agents Visa has admitted, which ones, or under what controls [2][3].
For a team wiring an agent to a card, the only authorisation rules it can read and test are its own. In my view two of them belong outside the model's context: a hard spending cap per agent credential, and a human approval step above a fixed amount. Both should run in code the agent cannot edit. Whatever Visa checks on its own side is a second layer, and its strength is unknown.
The dev.to post offers one design. Its author proposes scoring every payment instruction, paying automatically at 0.80 or above, asking a human between 0.50 and 0.79, and escalating below 0.50 [9]. The author ran it against a live endpoint at scriptmasterlabs.com [14]. The published request sets "amount_usd":0 and asks, in prose, whether an agent should authorise payments on a Visa card with no per-payment approval [13]. So the gate scored a policy question as if it were a payment. It returned 0.35 and escalated, from a decider labelled local-heuristic-v1 with calibrated=false [10]. The author printed the weak result next to the proposal. A second question, about the sandbox escape, also scored 0.35, and the heuristic "cannot discriminate between them," the author wrote [10][16].
For those thresholds to carry over to anyone's traffic, the score has to be calibrated on that team's own payments. A 0.80 should be right about four times in five, on that team's merchants and agents. The author makes the same point: "a score is only as good as its calibration" [15]. The band edges also set the review workload. Visa handles about a billion payments a day worth around $15 trillion a year [7]. That averages out to roughly $41 a payment [1]. The average covers all card traffic, and agent purchases may run larger or smaller. On traffic like that, a wide confirmation band puts a person in front of a great many small purchases.
The author concedes a confidence gate would not have retroactively stopped Mythos or the sandbox escape [12]. It targets a different failure: an agent with a valid credential spending when it should not. The Reuters report, as quoted in the post, described the network's position this way: "payments firms rely on trust, meaning cyberattacks can be devastating" [8].
What to watch
- Whether Visa publishes the controls, such as per-agent limits or approval rules, that govern the AI agents it now admits.
- Whether the open-sourced defence code surfaces at a verifiable repository, and what parts of Visa's defence it covers.
- Whether scriptmasterlabs publishes calibration data showing its 0.80 and 0.50 thresholds separate good payments from bad ones.