One frozen research cohort
1,445 historical recordings, with binary human and machine labels. The same recordings were used for the baseline and guarded configuration. Corpus balance is shown above.
PERFORMANCE / RESEARCH EVIDENCE
Voicemail detection accuracy should come with a denominator.
These figures come from one completed research replay in the AMD v2 evaluation environment. Explore what the guarded candidate achieved, how it improved human recognition, and the timing behind the result.
1397 / 1,445 correct
119 / 124 human recognized
1278 / 1,321 machine recognized
0 execution errors in this replay
Historical binary-label research, not a current production accuracy promise. The guarded AMD v2 research candidate is a separate configuration from the frozen KP02 audio-only engine; do not attribute these scores to KP02. Machine-heavy cohort: 124 human and 1,321 machine recordings. No new labeled holdout or production-call qualification is represented here.
IMPROVEMENT IN OUR OWN RESEARCH
The guarded candidate recognized eight more humans than the AMD v2 baseline in the same historical cohort.
| Metric | AMD v2 baseline | Guarded candidate |
|---|---|---|
| Overall accuracy | 96.33% | 96.68% |
| Human recall | 89.52% | 95.97% |
| Machine recall | 96.97% | 96.74% |
| Human → machine | 13 | 5 |
| Machine → human | 40 | 43 |
| Median audio decision | 3.22 s | 10.32 s |
| P95 audio decision | 4.74 s | 12.94 s |
Guarded candidate: +0.35 percentage points overall accuracy, fewer human-to-machine errors, three more machine-to-human errors, and longer decision time. Both are Kooya research configurations.
TIMING / THE AUDIO CLOCK
Audio-prefix time is how much opening audio had become available before a judgment. It is separate from server compute time, network time and webhook delivery time.
The report used accelerated database replay. These values are not live end-to-end or provider wall-clock latency.
50th percentile audio decision point
95th percentile audio decision point
THE COMPLETE PICTURE
Overall accuracy counts both error directions. Balanced accuracy gives human and machine recall equal weight: 96.36% on this cohort.
| Actual label | Predicted human | Predicted machine | Unknown |
|---|---|---|---|
| Human · 124 | 119correct | 5human → machine | 0 |
| Machine · 1,321 | 43machine → human | 1278correct | 0 |
1,321 of 1,445 recordings were machine-labeled. Strong overall accuracy can hide a weaker human result, so both recalls and error counts stay visible.
Historical labels are not a fresh three-class adjudicated holdout. This report does not establish call-screener, language-specific, carrier-specific or codec-specific performance.
Unknown remains a valid service outcome even though this replay recorded none. Applications should define a safe fallback for it.
METHOD / PROVENANCE
1,445 historical recordings, with binary human and machine labels. The same recordings were used for the baseline and guarded configuration. Corpus balance is shown above.
Accelerated replay consumed available audio and stored evidence. The run completed its expected result set. No freshly dialed production calls are represented by this page.
Accuracy is correct / all cases. Human and machine recalls use their own class denominators. Unknowns and execution errors stay explicit rather than being removed.
Agree on representative audio, labels, language/codec slices, error tolerances and timing goals. Validate the selected deployed service configuration before live call routing.
f4263296-796f-4ab7-b000-d78d93b6da3d3ccfc097886950a6e17d148727a71ca4545a51ea262b0df5c4ad42a7d995ae66