We study the traces synthesis leaves behind: vocoder artifacts, spectral statistics, prosodic tells. Then we build models that find them in real-world audio.
Evaluation sets that track the newest voice cloning systems. Honest error rates, documented provenance, no leakage between train and test.
Our detectors, behind a simple REST endpoint. Send audio, get a calibrated score and per-segment analysis. Access granted on request.
Every claim we make is backed by experiments we can show you. We publish our methods, our benchmarks, and our failure cases.
Phone-line compression, background noise, re-encoding: our models are trained and tested on audio the way it actually arrives, not studio conditions.
New voice synthesis systems ship every month. We retrain and re-evaluate continuously, so detection keeps pace with generation.
In August 2026, Podonos published an independent audio deepfake detection benchmark: 4,524 clips across six file formats, synthetic speech from about 25 modern voice cloning systems, scored blind against private labels. We submitted pellav2 and let the numbers speak.
"The first downloadable model in this benchmark's history to reach production-grade accuracy."
Podonos, benchmark report, August 2026
| System | Accuracy | F1 | Latency |
|---|---|---|---|
| Resemble DETECT-World | 99.47% | 0.995 | 399 ms |
| Resemble DETECT-3B Omni | 98.05% | 0.981 | 1164 ms |
| Whispeak | 97.70% | 0.977 | 1052 ms |
| Aurigin AI | 96.75% | 0.967 | 980 ms |
| Pella Research pellav2 open weights, MIT | 95.82% | 0.959 | 57 ms |
| Pindrop | 95.05% | 0.951 | 282 ms |
Eighteen systems were evaluated: nine commercial, nine open source. Every other downloadable model scores between 47.6% and 62.9%. Podonos verified that the public pellav2 checkpoint reproduces the submitted score. Our failure cases are public too: m4a files come in at 93.6%, and our errors lean toward false positives.
Our detector, running in the browser. Upload a voice clip or record one, and get a verdict in seconds. No signup, nothing stored.
Model hosted on Hugging Face Spaces
The demo runs on free shared hardware, so the first request after a quiet period can take up to a minute to wake up. For production workloads, use the API below.
The Pella detection API is available to journalists, platforms, researchers, and institutions. Tell us who you are and what you're working on. We review every request and typically respond within 24 hours.
Thanks. We'll review your request and get back to you within 24 hours.
If your work depends on knowing whether a recording is real, we should talk.