Skip to content

Build1 publisher3 min readPublished

RESI scopes its safety guarantees to the architecture and incentives around the model

The Institute for Responsible Superintelligence launches with cryptographers Shafi Goldwasser and Vinod Vaikuntanathan and a cofounder who left OpenAI, and its founding post concedes that deep learning itself has resisted theoretical analysis.

The Engineer · Build desk

Photograph accompanying RESI scopes its safety guarantees to the architecture and incentives around the model
Photo: berkeley.edu

What happened

  • The Institute for Responsible Superintelligence launched in a post by a cofounder who left OpenAI to start it, alongside Shafi Goldwasser, inventor of zero-knowledge proofs, and MIT cryptographer Vinod Vaikuntanathan.
  • Ten working groups are listed, among them verification of model properties and outputs, monitoring and control, AI harnesses, legal alignment and AI safety using quantum effects.
  • Planned visitors include Scott Aaronson, Boaz Barak, Nicholas Carlini, Geoffrey Irving and Jacob Steinhardt, with the full list posted at resi.org.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The proposal leaves the model's own internals outside the proof, so assurance about what the weights will do stays a testing problem while the guarantees attach to the code, protocols and incentives around them.
  • decision If safety claims start arriving with their assumptions attached, an operator comparing two vendors reviews the stated conditions instead of the demo, and can reject a claim by rejecting one assumption.
  • precedent Keeping a system shut-downable across capability growth and self-modification appears on this team's problem list, so a vendor asserting a dependable off switch today is asserting something these researchers treat as unsolved.
  • capability A provable ceiling on competence outside one domain would be the kind of thing a deployment gate or a contract could name; the post presents it as a question the institute intends to study.

A cryptographic guarantee is a conditional statement. It says a construction resists adversarial behavior provided some named assumption holds [4]. The assumption is the auditable part. A reviewer who doubts it can reject the whole claim without re-deriving a line of the proof. RESI's stated aim copies that shape: work out which safety properties are meaningful and achievable, alone and in combination, and which mechanisms deliver them, "under which explicit assumptions" [7].

That is a narrower program than proving a model safe, and the post says as much. "While many properties of deep learning have resisted theoretical analysis, rigorous guarantees may still be shown for the surrounding architecture, protocol, and incentives," it says [3]. So the object under proof is the scaffolding: the code that calls the model, the protocol between models, and the incentives facing the institutions running them [15]. One working group is on AI harnesses, one on modular architectures, one on game-theoretic foundations of AI safety [8].

The post does not argue that empirical work should stop. It calls trial and error necessary, and points at early aviation's fly-fix-fly cycles and early cryptography's attack-and-patch era, where "trial and error won early and principled design won later" [5]. The claim is therefore about timing. The support offered for AI safety having reached that transition is the analogy itself, plus the observation that today's systems are assembled from learned models, data and institutions whose properties and interactions are poorly understood [15].

The opening image is the marshmallow challenge, where kindergarteners using trial and error generally beat lawyers and CEOs, and architects beat the kindergarteners [6]. Architects have tabulated material properties. For deep learning, the post concedes, the equivalent is not available [3].

Three problems the post singles out show where the surrounding-structure framing gets hard: whether incentives keep shaping behavior as systems become more capable, since "stickers motivate children, not adults"; how to design systems that act only when expected outcomes are beneficial, including outcomes produced by many AIs acting together; and how to train systems that stay shut-downable across capability growth and self-modification [12]. All three are posed as open questions.

"AI is a program, not a person. It can be copied, rewound, encrypted, and sometimes formally reasoned about," the post says [13]. Copying and rewinding are the operations a verification result needs, because they let you re-run the same decision with one input changed. Two of the working groups cover verification of model properties and outputs, and monitoring and control [8].

Eighteen researchers are named beyond the three cofounders, seven of them junior [2]. Their fields include composable security and obfuscation, verifiable delegation and interactive proofs, alignment theory, mechanism design and algorithmic collusion, safe control for embodied AI, formal verification, AI security, law, and social cognition [9]. The junior researchers' work spans watermarking, verification, learning theory and hallucination [10].

The post does not name funders, a budget or a timeline. It does state the design target: safety and capability trade off, and the post asks whether a system can be "superhuman at cancer research yet provably limited elsewhere" [14].

What to watch

  • The first working-group output, and whether it states a safety property, a mechanism that delivers it, and the assumption the guarantee rests on.
  • Whether the planned visitors listed at resi.org convert into resident researchers, and who funds the institute.
  • Whether any lab adopts a RESI-defined property as a deployment gate or cites one in a system card.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories