PERFORMANCE / RESEARCH EVIDENCE

The evidence
behind the answer.

Voicemail detection accuracy should come with a denominator.

These figures come from one completed research replay in the AMD v2 evaluation environment. Explore what the guarded candidate achieved, how it improved human recognition, and the timing behind the result.

96.68%

Overall accuracy

1397 / 1,445 correct

95.97%

Human recall

119 / 124 human recognized

96.74%

Machine recall

1278 / 1,321 machine recognized

0

Unknown results

0 execution errors in this replay

Historical binary-label research, not a current production accuracy promise. The guarded AMD v2 research candidate is a separate configuration from the frozen KP02 audio-only engine; do not attribute these scores to KP02. Machine-heavy cohort: 124 human and 1,321 machine recordings. No new labeled holdout or production-call qualification is represented here.

IMPROVEMENT IN OUR OWN RESEARCH

More people
recognized as people.

The guarded candidate recognized eight more humans than the AMD v2 baseline in the same historical cohort.

HUMAN RECOGNITION / SAME 124 RECORDINGS
AMD v2 baseline89.52%
111 / 124 humans
Guarded candidate95.97%
119 / 124 humans
+6.45percentage points
in human recall
Same-cohort outcomes
MetricAMD v2 baselineGuarded candidate
Overall accuracy96.33%96.68%
Human recall89.52%95.97%
Machine recall96.97%96.74%
Human → machine135
Machine → human4043
Median audio decision3.22 s10.32 s
P95 audio decision4.74 s12.94 s

Guarded candidate: +0.35 percentage points overall accuracy, fewer human-to-machine errors, three more machine-to-human errors, and longer decision time. Both are Kooya research configurations.

TIMING / THE AUDIO CLOCK

A decision needs
evidence.

Audio-prefix time is how much opening audio had become available before a judgment. It is separate from server compute time, network time and webhook delivery time.

The report used accelerated database replay. These values are not live end-to-end or provider wall-clock latency.

GUARDED CANDIDATE / MEDIAN10.32s

50th percentile audio decision point

GUARDED CANDIDATE / P9512.94s

95th percentile audio decision point

See timing fields in the result →

THE COMPLETE PICTURE

Every outcome
has a place.

Overall accuracy counts both error directions. Balanced accuracy gives human and machine recall equal weight: 96.36% on this cohort.

Guarded candidate · historical confusion matrix
Actual labelPredicted humanPredicted machineUnknown
Human · 124119correct5human → machine0
Machine · 1,32143machine → human1278correct0

Read accuracy in context.

1,321 of 1,445 recordings were machine-labeled. Strong overall accuracy can hide a weaker human result, so both recalls and error counts stay visible.

Historical labels are not a fresh three-class adjudicated holdout. This report does not establish call-screener, language-specific, carrier-specific or codec-specific performance.

Unknown remains a valid service outcome even though this replay recorded none. Applications should define a safe fallback for it.

METHOD / PROVENANCE

A result you can trace.

01 / CORPUS

One frozen research cohort

1,445 historical recordings, with binary human and machine labels. The same recordings were used for the baseline and guarded configuration. Corpus balance is shown above.

02 / REPLAY

Completed causal replay

Accelerated replay consumed available audio and stored evidence. The run completed its expected result set. No freshly dialed production calls are represented by this page.

03 / SCORING

All terminal outcomes count

Accuracy is correct / all cases. Human and machine recalls use their own class denominators. Unknowns and execution errors stay explicit rather than being removed.

04 / NEXT VALIDATION

Prove it on your workflow

Agree on representative audio, labels, language/codec slices, error tolerances and timing goals. Validate the selected deployed service configuration before live call routing.

Report date
2026-09-16
Aggregate extraction checked
2026-10-08
Run ID
f4263296-796f-4ab7-b000-d78d93b6da3d
Source report SHA-256
3ccfc097886950a6e17d148727a71ca4545a51ea262b0df5c4ad42a7d995ae66
Download own aggregate evidence →

THE NEXT STEP

Make the next benchmark yours.

Define your own audio cohort and success criteria with the Kooya Answer team.

Discuss an evaluation