Wiretap · voices on the tape

2-min multi-speaker excerpt → pyannote speaker separation → audio cleanup of a degraded intercept → local voice-print verification. Spectrograms rendered at a matched 16 kHz scale.

speakers separated
4
cross-talk regions
4
cleaned re-ID (Speaker A)
100 / 100
other panelist (Speaker B)
58 / 100

1  —  Speaker separation  (4 clean tracks, one per voice)

Speaker A REFERENCE
Speaker A spectrogram
64.0s speech5 turns
dominant voice — enrolled as the reference print
Speaker B
Speaker B spectrogram
30.3s speech6 turns
a second panelist — the negative control
Speaker C
Speaker C spectrogram
18.0s speech4 turns
third voice, zero cross-talk
Speaker D
Speaker D spectrogram
5.9s speech3 turns
brief fourth voice

2  —  Audio cleanup  (degraded intercept → recovered)

Degraded intercept SYNTHETIC
degraded spectrogram
Speaker A's track, band-limited to 300–3400 Hz with added hiss + room tone — a labeled demonstration of a noisy intercept.
After voice isolation ELEVENLABS
cleaned spectrogram
Noise floor drops from a full-spectrum haze to black; the cleaned track re-identifies as Speaker A at 100/100.

The noise floor visibly collapses — the degraded track is a uniform haze, the enhanced track a black background with isolated speech (voice isolation also restores high-frequency bandwidth beyond the 16 kHz view). The cleanup preserved voice identity: enrolling the cleaned reference still re-finds Speaker A on the tape.

Voice-print scores are 0–100 rank scores, NOT liveness — a cloned or synthetic voice can score high, and the same speaker scores lower across languages/compression. A match is a corroborated lead, never an identification.

Source: a ~2-min analysis excerpt of the Lightcone podcast (channel: Y Combinator, youtube.com/watch?v=e1Yhs9BEOSw). Speakers labelled A–D by the diarizer, not named. The "degraded intercept" noise is synthetic and labeled as such.