All notes
On-device MLBy Valentino HartantoUpdated 19 Sep 20266 min read1,095 characters

Keeping presentation coaching on device

A practical look at privacy, latency, and synchronizing local speech analysis with a live Keynote deck.

Privacy changes the architecture

Presentation coaching handles material people may not want to upload: unreleased work, research, client information, or a nervous first rehearsal. Waverd treats local processing as a product requirement rather than a technical preference.

WhisperKit and Core ML keep transcription on the Mac. That removes a cloud round trip and makes the privacy promise easier to explain: the recording and analysis stay close to the person presenting.

A transcript is only half the context

Feedback becomes more useful when it can answer which slide was visible when a phrase, pause, or filler occurred. Waverd polls Keynote through AppleScript and aligns slide changes with the speech timeline.

The resulting model is temporal rather than document-based. Speech events, slide transitions, and recording progress share a common clock, which makes per-slide feedback possible after the rehearsal.

Useful intelligence stays quiet

The goal is not to show every signal the models can produce. It is to surface a small set of observations that someone can act on before the next rehearsal. Local intelligence earns trust when the interface remains calm, specific, and easy to dismiss.

Next notePutting PSO inside the experiment loop