Product2 publishers3 min readPublished
Siri Recap turns a colleague's words into a summary only the wearer can reach
Apple's new audio intelligence pipeline deletes the raw audio twice and keeps the text. Every switch described in the announcement sits on the wrist of the person listening, not with the person being summarised.
The Product Desk · Product desk

What happened
- Apple said on Wednesday that the Apple Watch Series 12 and Ultra 4 ship with four opt-in audio intelligence tools driven by the watch microphones, including sound recognition, Live Rewind and a conversation recap.
- Siri Recap can run all the time or on a schedule, with a dedicated model deciding when speech is happening before any audio enters the protected buffer on the S11 chip.
- Live Rewind and Siri Recap both depend on Siri AI, which Apple says arrives in beta later this year in English, with more languages after that.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- constraint Consent in this design is collected once, from the wearer, at setup. The other half of a summarised conversation has no toggle and no indicator, and cannot exclude a single exchange, which pushes the whole problem onto social negotiation.
- exposure What survives the pipeline is a titled summary of what other people said, sitting on the wearer's phone and watch under the wearer's encryption. Those people cannot see it, correct it or ask for it to go.
- precedent Apple has now set the shippable template for ambient capture on the body: a hardware enclave paired with an opt-in screen, but no bystander mechanism at all. Cheaper wearables will copy the language long before they copy the exclave.
- decision Anyone who writes meeting or clinical-space policy has one beta cycle to decide whether a wrist counts as a recording device, given that Apple's own account says no recording is created.
A wearer double-taps the crown while a waitress runs through the lunch specials, and fifteen seconds of her speech becomes text he can ask Siri about or save to the Siri app [14][18]. Apple's privacy account of that moment is entirely about him. No audio recording is created or stored, and the transcript is end-to-end encrypted [3][16]. She is in the pipeline without any point in it she can touch [22].
That pipeline runs through three compute environments before a summary exists: the Secure Exclave on the watch, the on-device speech and language models on the iPhone, and Apple's foundation models in Private Cloud Compute [21]. Raw audio is deleted twice on that path, once on the watch after transfer and once on the phone after transcription [20]. What persists is the condensed version with a generated title, encrypted and sent back to phone and watch [10]. Which makes the load-bearing sentence in Apple's framing, that the features do not create or store audio recordings [3], both accurate and slightly beside the point. The durable artifact is text about what somebody else said.
That text does not travel alone. Apple says it also sends Now Playing data, calendar data, high-level location labels such as home, work and school, and point-of-interest categories like grocery store or park, while withholding precise location [11]. A safety model on the phone screens the text to omit potentially harmful terms before it leaves [9]. Both are reasonable choices, and both are made about a second person's words with no input from that person.
Sound Recognition is the clean case in this set. It notifies the wearer about doorbells, sirens, alarms or a baby crying without sending anything off the watch [4], works when the phone is elsewhere, and Apple points it at people with hearing loss [19]. It reacts to sound, not to anyone's speech.
Product teams tell themselves users do this with something like Siri Recap: open the setup screen, weigh it, and set a schedule that covers only the hours where nobody would mind [6][15]. What users actually do with a setting that has an always-on option [6] is take the default and stop thinking about it. Live Rewind and Siri Recap need Siri AI, which arrives in beta later this year in English [17], and there isn't yet any retention or usage data to check that assumption against.
The forcing function for anyone shipping ambient capture, or writing a policy about someone else's, is a two-axis grid: whose speech gets processed, and who holds the control. Wearer's speech under wearer control is dictation, settled. Others' speech under their own control is a meeting recorder that announces itself, also settled. Others' speech under wearer-only control is the cell these two features occupy, and the sources describe no outward indicator and no per-conversation exclusion for the other party [22].
So the policy line that bites is a device-class line rather than a recording rule, because on Apple's account nothing is recorded [3]. If a colleague wants out of a summary, the only mechanism available is the wearer's mouth.
What to watch
- Whether the Siri AI beta ships any outward signal that other people in the room can see or hear, rather than another wearer-side toggle.
- Whether Siri Recap gains a per-conversation exclusion, so a wearer can drop one meeting without switching the feature off entirely.
- Whether Apple publishes how wearers actually configure Siri Recap, in particular the split between always-on and scheduled listening.