Send isolated far-end audio, receive a judgment, and keep control of your call in your own application.
KOYAA ANSWER / API V1GUIDED SERVICE INTEGRATION
01 / START HERE
Your application owns the call.
Kooya Answer classifies the far-end opening audio as human, machine or unknown. Your dialer decides what to do next. The service does not originate calls or hang them up.
Register a client and source.Kooya provisions a scoped API key and a source configured for WAV, WSS, optional direct SRTP, WebRTC or SIP.
Send audio with call identity.Declare the source and provide a stable client reference. Isolate the callee or far-end track.
Receive the judgment.Use the HTTP response, WSS judgment message, LiveKit data message, configured HTTPS webhook or authenticated session polling.
Apply your call policy.Map human, machine and unknown to your own routing, disposition and agent workflow.
Use an API key only from a trusted backend.
Keep it out of browser code, URLs and application logs. WebRTC clients receive a short-lived room token; WSS browser clients should connect through your backend so it can add the Authorization header.
Direct SRTP is an optional UDP media path. The client creates an authenticated session over HTTPS, supplies one SDES key line, then sends encrypted media to the returned UDP address. This implementation accepts AES_CM_128_HMAC_SHA1_80 and rejects plaintext, tampered and replayed packets before audio is counted. The public service keeps direct SRTP disabled until its deployment and network path are explicitly qualified. SIP signaling and its SRTP media remain a separate LiveKit integration. SDES keying specification ↗
03 / AUDIO FORMATS AND FRAMING
Send one isolated callee channel.
HTTP WAVPCM16
Uncompressed signed 16-bit PCM in a WAV container. Up to 10 MiB. The WAV format must match the source rate and channel selection.
WSSPCM16LE · PCMU · PCMA
Authenticated WebSocket accepts little-endian PCM16 at qualified source rates, plus G.711 μ-law or G.711 A-law at 8 kHz mono. L16 is not supported over WSS.
DIRECT SRTP · OPT-INPCMU · PCMA · L16 · G.722
Direct UDP media uses the negotiated peer tuple and SSRC with AES_CM_128_HMAC_SHA1_80. G.722 uses static payload type 9, 16 kHz decoded mono audio and an 8 kHz RTP timestamp clock. The hosted endpoint stays disabled until deployment and network qualification are complete.
PCMU and PCMA use 8 kHz mono. Direct G.722 decodes to 16 kHz mono PCM while its RTP clock runs at 8 kHz. RTP payload mappings are checked.
For stereo source audio, select the left or right callee channel. Mixed agent and callee audio is not accepted as an isolated caller source.
WebRTC sends negotiated Opus media through LiveKit; the media adapter decodes it before evaluation.
Adapters normalize selected audio to mono PCM16 at 16 kHz and feed the detector in 20 ms frames: 320 samples / 640 bytes.
Send raw audio chunks, not a WAV header, over WSS. Each binary message is bounded at 65,536 bytes; use 20 ms chunks where possible.
The detector commits a judgment from at most the first 15 seconds of audio. Recording can continue after that decision until the transport ends.
04 / HTTP WAV EVALUATION
Upload audio and receive a session result.
POSThttps://api.amd.kooyaai.com/v1/evaluationsBearer API key
Send source_id, callee_isolated=true, the optional client_reference, and the audio file as multipart form data. The source must be configured for WAV upload.
A completed evaluation returns session_id and result. If processing is still underway, the response is 202; use the returned session identifier with GET /v1/sessions/{session_id}. Reuse the same Idempotency-Key when retrying the same request.
05 / AUTHENTICATED WSS
Open a session, then stream audio.
POSThttps://api.amd.kooyaai.com/v1/sessionsJSON + Bearer API key
Create a session with the configured WSS source. Include source_id, callee_isolated:true and a call reference. Every session admission requires an Idempotency-Key header: use one unique value per call and reuse that value only when retrying the same request. The response supplies session_id, the configured audio contract, and audio_websocket_path.
Connect to wss://api.amd.kooyaai.com plus the returned path and send the same Bearer key in the WebSocket Authorization header. Send binary audio only, then send the exact control message {"type":"stop"} when the far-end stream ends.
CREATEAuthenticated session
→
CONNECTWSS audio path
→
STREAMBinary chunks
→
JUDGMENTJSON message
The service sends {"type":"judgment","result":{...}} when the detector commits. The result is also pollable by session ID. Send stop or call the session end endpoint so full-session recording can finish; the 15-second detector decision does not stop recording.
06 / DIRECT SRTP · OPTIONAL
Encrypted UDP media with an HTTPS admission.
The service returns a UDP address only after it admits the source, validates the RTP contract and matches a ready media worker. Direct SRTP is disabled by default. Do not send production traffic until the deployment operator enables the dedicated ingress profile and confirms the public UDP route and carrier compatibility.
Send the key-bearing attribute only in the authenticated HTTPS JSON request. The API never returns or stores the key; it retains a one-way fingerprint and passes the key through its private non-persistent control queue to the media worker. The worker verifies SRTP authentication before RTP parsing, decoding or usage metering. UDP media is direct network traffic; HTTP, WebSocket and Nginx do not proxy it. Configure the caller's negotiated remote IP and source port in the request.
Require an explicit network qualification.
Confirm UDP reachability, one-way media direction, codec/payload/clock, SSRC, key rotation, packet loss and the end-of-session signal before enabling a carrier integration. The SDES attribute must be sent over authenticated HTTPS; SRTP media itself uses UDP.
07 / WEBRTC AND SIP
Browser media and telephony have separate paths.
WEBRTC / LIVEKIT
A room, one callee track, one judgment.
Create a session through the HTTPS API with the intended callee identity and track.
Join the returned room at wss://rtc.amd.kooyaai.com using its short-lived token.
Publish the isolated far-end microphone track; negotiated Opus is decoded by the media adapter.
Receive a reliable room data message on amd.judgment, or recover by session polling.
Use the returned TURN hostname when a peer connection needs relay. The service selects only the configured callee identity and track.
SIP / LIVEKIT SIP
TLS signaling, SRTP media.
Kooya provisions an allowed inbound trunk with media encryption set to SIP_MEDIA_ENCRYPT_REQUIRE and maps it to the source.
Your telephony edge routes the designated call leg to sip.amd.kooyaai.com:5061;transport=tls.
SIP/SDP negotiates SRTP on its separate UDP media path; the LiveKit SIP bridge attaches only the admitted session and callee attributes.
Receive a configured HTTPS webhook or poll the session for its judgment.
LiveKit SIP can negotiate G.722 on the wire and decodes it before the analyzer; the application source remains configured for decoded PCM. SIP admission remains disabled until an operator provisions the trusted trunk and dispatch mapping and explicitly enables SIP. The SIP service and its TCP 5061 / UDP 10000–10999 ports are published only when AMD_SIP_ENABLED=true; the default release leaves them closed. When enabled, service readiness checks that every admitted inbound trunk requires SRTP. The consumer publishes a quiet return audio track, not an audible prompt. Do not route live calls through this path until a real trunk-to-room test proves answer, callee audio, recording, hangup and judgment delivery. The service does not originate outbound calls. CommPeak or another customer SBC is not changed by Kooya's deployment.
WebRTC signaling uses WSS on rtc.amd.kooyaai.com; WebRTC media uses its encrypted ICE/DTLS-SRTP path and TURN when required. When SIP is explicitly enabled, secure signaling uses TLS on sip.amd.kooyaai.com:5061 and SRTP media uses its own UDP RTP range. Plain SIP 5060 is never published. These media ports are direct network traffic, separate from the HTTPS API.
08 / JUDGMENT AND POLLING
Use the decision that was actually made.
human
Use your human-answer workflow.
machine
Use your voicemail or machine workflow.
unknown
Apply an explicit fallback or review policy.
“Machine” does not identify a precise voicemail-beep time or guarantee that a mailbox is ready to record. The result does not authorize hangup or other call control.
Poll GET /v1/sessions/{session_id} or its returned status_url after a timeout or lost connection. End a live session with POST /v1/sessions/{session_id}/end; its response may remain 202 while recording finishes.
09 / WEBHOOKS AND CUSTOM BODY MAPPING
Your endpoint. Your JSON shape.
Configure an HTTPS webhook at the client or source level; a source-specific configuration overrides its client's configuration. Delivery is stored durably and sent after the judgment transaction commits. The stable event identifier is sent as X-AMD-Event-ID. Use custom body mapping to fit the result into your existing event contract.
Map supported event fields into JSON values with {{variable}} placeholders; whole-value placeholders retain their JSON type.
Configure a decision map such as human → HUMAN, machine → VOICEMAIL and unknown → REVIEW.
Add custom headers; secret references are encrypted and secret values are not returned by normal read APIs.
Return a 2xx response after durably accepting the event. Retries may repeat a delivery, so deduplicate by event_id or X-AMD-Event-ID.
Only public HTTPS destinations are accepted. Destination IPs are checked and DNS is revalidated for each delivery; redirects are not followed.
Webhook delivery is at least once. Configure your preferred body fields and names during onboarding, or use the session status endpoint as the recovery path.
10 / USAGE, RECORDINGS AND ONBOARDING
Meter audio, not tokens.
ANALYZED AUDIO VOLUMEAnalyzed minutes = total audio_analyzed_ms ÷ 60,000
Evaluation count and analyzed audio milliseconds are recorded separately. Received audio, decision audio, full recording duration, bytes and processing time provide additional operational and research measures.
Usage tracking is enabled for measurement; automated customer billing is not enabled. Commercial options are pay-as-you-go usage and capacity planned around concurrency. Rates and qualified capacity are agreed during onboarding; no rate or 250+ concurrent-session result is promised on this page.
When recording is enabled, full session audio continues until the transport stops, even after the detector has committed its judgment. Private recordings are encrypted and retained without automatic expiration. See data handling.
What to bring to a pilot
Your dialer or telephony platform and the source of callee audio
Expected evaluation volume and peak concurrent calls
Codec, sample rate, channel layout and network source addresses
Preferred judgment path and webhook fields/body mapping
A representative labeled evaluation set and acceptable error/timing goals
Retry with the same idempotency key and poll the session.
A disconnected stream, missing RTP packets or unavailable worker does not mean machine. Treat the result as unavailable until a judgment is committed; use the returned session ID and webhook event ID for recovery and deduplication.