Build1 publisher3 min readPublished
Replit's new black-box pen tests hand the findings to the agent that wrote the code
An August 17 release scans a sandboxed copy of a published app from the outside, then routes confirmed findings to Replit Agent to draft patches. One human click stands between audit and self-certification.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Replit added black-box penetration testing on August 17, giving its AI app builder a way to probe published applications from an attacker's perspective and pass confirmed findings to Replit Agent for proposed fixes.
- Replit introduced the feature in a thread on X and an accompanying product post by member of technical staff Alexandre Cuoci.
- A user launches a Level 3 scan from a project's Security Center; Replit then runs a white-box scan with access to the source code alongside a black-box scan that receives only the application's link.
- Replit says both scans target a full copy of the application inside a private sandbox, keeping the tests away from production users.
- The black-box scanner clicks through the application while observing its network requests; it first tests what an unauthenticated visitor can reach, then signs in as an ordinary user to check whether that account can access another user's records or restricted administrative areas.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Replit turned on black-box penetration testing on August 17, letting its AI app builder probe a published application from an attacker's perspective and pass confirmed findings to Replit Agent for proposed fixes [1]. The company announced it in a thread on X and a product post by member of technical staff Alexandre Cuoci [2]. The consequence is structural rather than technical: the system that generates the code now also assesses it and writes the remedy.
The mechanics are straightforward. A user launches a Level 3 scan from a project's Security Center, which runs a white-box scan with access to the source alongside a black-box scan that gets only the application's link [4]. Replit says both scans target a full copy of the application inside a private sandbox, keeping the tests away from production users [5]. The external agent clicks through the app while watching its network requests, first testing what an unauthenticated visitor can reach, then signing in as an ordinary user to check whether that account can read another user's records or get into restricted administrative areas [6]. Replit says the scanner also identifies the technologies behind the app and tests failure modes associated with that setup [7].
Replit's own examples make the case for the second perspective better than any framing does. The white-box pass caught a logic flaw that let a user with revoked access keep operating because the application never rechecked an old session [8]. The black-box pass found a separate admin dashboard sitting at a predictable address with no authentication [9], and in a Replit-built multiplayer game it found an endpoint that could be flooded to crash an active match [10]. Replit says the code scanner missed those runtime exposures because the underlying source did not look defective on its own [11]. That is the familiar gap between reading a repository and using a deployed system.
Then the loop closes. Findings feed directly into Replit Agent, which prepares patches for the user to inspect [12], with a human approval step before changes return to the main project and a required republish before fixes reach production [13]. Replit itself notes that an automated patch can change permissions, request handling or application behavior in ways that need product context, even when the vulnerability is real [14]. That approval click is the entire boundary between an audit and a self-certification, and it is being asked of a user who, by the premise of the product, did not write the code under review.
The tiering tells you where this sits commercially. Level 1, which Replit says is free, covers dependency checks and static analysis [15]; Level 2 adds the deeper white-box agent review [16]; Level 3 runs both agents [17]. So the attacker's-eye pass is the top of three tiers, and two of the three see only what the source reveals [25]. Auto-Protect layers on a malicious-package firewall, a web application firewall and SSL/TLS encryption [18].
This has been assembled in stages: a white-box Security Agent that maps architecture, builds a threat model and checks routes and APIs for SQL injection, cross-site scripting and request forgery launched April 21 [19], and Security Center gained multi-project vulnerability review in May [20], putting roughly four months between the first agent and this one [24]. It matches the wider consolidation, with RuntimeWire reporting in June that Agent was expanding into websites, mobile apps, pitch decks and launch videos [21] and a design suite with Figma imports arriving in July [22]. RuntimeWire's framing of the underlying obligation is fair: if someone can publish an app holding customer data in a day, security work has to move at about that speed [26].
Watch whether Replit publishes acceptance and regression rates for Agent-authored patches, whether black-box scanning drops below the top tier, and whether the human approval gate survives contact with volume.