2-min multi-speaker excerpt → pyannote speaker separation → audio cleanup of a degraded intercept → local voice-print verification. Spectrograms rendered at a matched 16 kHz scale.
The noise floor visibly collapses — the degraded track is a uniform haze, the enhanced track a black background with isolated speech (voice isolation also restores high-frequency bandwidth beyond the 16 kHz view). The cleanup preserved voice identity: enrolling the cleaned reference still re-finds Speaker A on the tape.
Voice-print scores are 0–100 rank scores, NOT liveness — a cloned or synthetic voice can score high, and the same speaker scores lower across languages/compression. A match is a corroborated lead, never an identification.
Source: a ~2-min analysis excerpt of the Lightcone podcast (channel: Y Combinator, youtube.com/watch?v=e1Yhs9BEOSw). Speakers labelled A–D by the diarizer, not named. The "degraded intercept" noise is synthetic and labeled as such.