Skip to content

Product1 publisher3 min readPublished

Record, don't prompt: two labs converge on demonstration as the agent interface

Record & Replay and Record a Skill shipped weeks apart. Both bet the context agents need lives in what people do, not in what they can be bothered to write down.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Record, don't prompt: two labs converge on demonstration as the agent interface
Photo: pcmag.com

What happened

  • In June, OpenAI launched Record & Replay, a feature that lets ChatGPT and Codex users demonstrate a workflow and turn it into a reusable skill.
  • Weeks after OpenAI's launch, Anthropic unveiled Record a Skill inside Claude Cowork: users record their screen doing a task, narrate their reasoning, and Claude turns it into a skill it can run again.
  • The Fast Company essay's author argues that the two companies converging on the same solution within weeks is an admission that prompting alone was never going to get AI where it needs to go.
  • The author says they worked at Apple for 12 years and, as a founding engineer, spent their formative years building the Chinese version of Siri.
  • The essay's example: telling a voice assistant to set an alarm for 6 every day leaves it unclear whether the time is morning or night, because people know their own schedule and expect whoever is listening to know it too.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

In June, OpenAI launched Record & Replay, which lets ChatGPT and Codex users demonstrate a workflow and turn it into a reusable skill [1]. Weeks later, Anthropic unveiled Record a Skill inside Claude Cowork: record your screen doing a task, narrate your reasoning, and Claude turns it into a skill it can run again [2]. Two competitors landing on the same input method within weeks is worth reading as a product decision, not a launch.

The argument for reading it that way comes from a Fast Company essay whose author treats the convergence as an admission that prompting alone was never going to get AI where it needs to go [3]. The author spent 12 years at Apple as a founding engineer, building the Chinese version of Siri [4], and frames the problem as continuous with that work: tell a voice assistant to set an alarm for 6 every day and the system cannot tell morning from night, because the user already knows their own schedule and assumes the listener does too [5].

The concept underneath is tacit knowledge, a term coined in 1966 by the philosopher Michael Polanyi for what people know but cannot quite articulate [6]. The essay cites one study estimating that 40% of a company's valuable knowledge sits inside individual employees' heads and is never written down [7]; it does not name the study, so treat the figure as illustrative rather than measured. The concrete version is better. Ask someone how they file an expense report and they will say they upload a receipt, categorize it, and submit; they will leave out that meals over $75 go to their manager for review, and that client dinners are classified differently from team lunches [8]. A prompt cannot correct its way to context that was never stated [9]. A demonstration captures the sequence, the decision points, and the small judgment calls along with the actions [10].

The operator-relevant part is what a captured skill becomes. Paired with a scheduled task, a recorded skill runs autonomously in the background without someone reopening it [11]. That is the difference between a macro and a coworker, and it is also where the cost of a bad recording compounds. The essay's stated prize is recovering the time already spent supervising these tools: a study from Glean found employees spend roughly 6.4 hours a week botsitting, meaning correcting output and reexplaining things the tool should already know [12]. On a 40-hour week that is about 16% of working time [13], or roughly 333 hours a year per employee [14]. Any capture feature has to beat that bar, not just exist.

The longer bet is workflow mining: technology that observes how work is done, extracts reusable knowledge, and builds a library of workflows people can select and personalize, with the argument that the library sharpens as versions accumulate, the way open-source code improves as more developers build on it [15].

That analogy is doing a lot of work. Open-source improves because code is readable, diffable, and testable by people who did not write it. Nothing in this account establishes that a recorded workflow is legible enough to review, or how often one reruns correctly after the underlying app moves a button. Watch for three things: whether either vendor publishes a success rate for replayed skills, whether recordings can be edited rather than only re-recorded, and whether anyone measures botsitting hours after adoption. If that number does not fall, demonstration has changed the authoring step and nothing else.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories