Architecture of a voice AI deployment on AWS
Every service in the call and the control plane, and the path a call takes through them. Pick a flow to trace it.
Swipe sideways to see the whole map.
Call plane services
Every service a call touches, in the call plane account in Sydney.
| Service | What it does | How it scales |
|---|---|---|
| Admission | Takes the carrier's call events, applies answering hours and caller blocks, claims a worker and answers the call. | Two tasks in two zones |
| Voice services A and B | Call workers. A supervisor keeps idle workers ready, and each worker takes one call and exits. | On the worker slots calls need, with a daytime minimum |
| Post-call service | Writes call records, runs audio metrics and quality checks, renders greetings and serves chat. | On queue age |
| Jobs service | Usage meters, call analysis, alerts, follow-up and customer webhooks, read from queues. | On the oldest message's age |
| Test service | Simulated test calls on their own workers, never on the workers that answer customers. | From zero when tests are queued |
| Route controller | Gives each task a route on the media load balancer when it starts, and removes it only after the task stops. | One function per task event |
| Database | Aurora PostgreSQL Serverless, encrypted under a customer-managed KMS key. | Serverless capacity, with a reader in single-tenant deployments |
| Object storage | S3 for recordings, call records in transit and cached greetings, with Block Public Access on. | Not applicable |
| Queues and registry | SQS queues for post-call work, each with a dead-letter queue, and a DynamoDB table of workers and call ownership. | On demand |
| Failover handler and canary | Answers calls in Melbourne when Sydney cannot, and places test calls around the clock. | One function per event |
Control plane services
The dashboard, sign-in and releases, in our account or in yours.
| Service | What it does | How it scales |
|---|---|---|
| Web app | Dashboard and agent editor, behind CloudFront and a web application firewall. | Two tasks in two zones |
| Sign-in | GoTrue, with SAML 2.0 and OIDC single sign-on. | Two tasks in two zones |
| Data API | PostgREST, with row-level security on every table. | Two tasks in two zones |
| Release pipeline | Builds each image once and promotes it to each environment by digest. | Runs per release |
Design rules
One call per process
Each call runs in its own operating system process, which exits when the call ends. In our measurements a second call on the same event loop added 37% to median turn latency.
Answer first, then place
Admission acknowledges the carrier at once, then places the call on a named worker before any audio arrives.
One owner per call
Only the handler that answered a call sends it later commands, so a failover never splits control of a call.
Calls do only calls
Analysis, metrics, webhooks and tests run on other services from queues, so post-call work never stalls a live call.
Protected while busy
A task holding a call is protected from scale-in, and a release waits for its calls to end.
A second failure domain
Failover answering, canary calls, paging, backups and the audit trail sit in Melbourne.
Security questionnaires
We answer security questionnaires in writing and take your team through the architecture on a call. Send yours to hello@verticalai.com.au.
