Architecture Reference
How CX-AI works
under the hood
CX-AI is a voice-powered AI customer service platform built on OpenAI's Realtime API and Vercel's edge infrastructure. This page explains the full architecture โ from WebRTC audio to tool execution to RAG-powered policy lookup.
Overview
What CX-AI Does
A unified platform for AI-augmented customer service, covering both full automation and human assistance.
Mode 01 ยท Self-Service
Customer Talks to AI
A customer speaks directly with the AI agent "Alex" via voice in real time. Alex can look up orders, process returns, answer policy questions, and manage account details โ no human required.
Mode 02 ยท Agent Assist
AI Assists the Human Agent
The AI listens silently to a live customer call and proactively surfaces relevant data โ orders, products, policies โ on the agent's screen, along with a suggested reply to read aloud.
12
AI Tools Available
2
Guardrail Layers
pgv
Vector Search Engine
WebRTC
Audio Transport
Technical
Tech Stack
A modern, minimal stack optimized for low latency and developer velocity.
Voice Session Flow
How a customer's voice travels from browser to OpenAI and back in near-real time.
Core Flow โ WebRTC Session Setup
Browser (Customer)
01
User clicks Connect
โ
02
POST /api/session
mode: self-service
โ
03
Creates RTCPeerConnection
getUserMedia() โ mic capture
createDataChannel('oai-events')
createOffer() โ SDP offer
createDataChannel('oai-events')
createOffer() โ SDP offer
โ
05
setRemoteDescription(answer)
WebRTC established
Audio streams bidirectionally
Audio streams bidirectionally
Vercel Backend
โ
Receives session request
Validates mode param
โ
โ
Calls OpenAI to mint token
Returns client_secret
10-minute TTL, single-use
10-minute TTL, single-use
OpenAI Realtime
04
Direct browser โ OpenAI WebRTC
POST /v1/realtime/calls
Bearer <ephemeral_secret>
Body: SDP offer โ SDP answer
Bearer <ephemeral_secret>
Body: SDP offer โ SDP answer
โ
โ
Audio + events flow
Mic โ OpenAI voice
DataChannel carries JSON events
DataChannel carries JSON events
Key design decision: The browser talks directly to OpenAI for audio. Vercel only issues the short-lived secret and executes tool calls โ it never proxies audio, keeping latency minimal.
Push-to-Talk (PTT)
User holds Talk button
mic track.enabled = true โ audio flows to OpenAI
User releases button
mic track.enabled = false โ server VAD detects silence (threshold 0.8, 1500ms) โ AI responds
Tool Execution Flow
How the AI agent fetches real data when it needs to answer a question.
Example โ Customer asks about order #100001
OpenAI โ Browser DataChannel
01
AI emits function_call event
response.output_item.done
{ name: 'lookup_order', call_id: '...',
arguments: '{"order_id":"100001"}' }
{ name: 'lookup_order', call_id: '...',
arguments: '{"order_id":"100001"}' }
โ
04
AI receives tool result
conversation.item.create โ function_call_output
response.create โ AI speaks result
response.create โ AI speaks result
Browser โ Vercel โ Supabase
02
useRealtimeSession โ executeTool()
POST /api/tool/lookup_order
{ order_id: '100001' }
{ order_id: '100001' }
โ
03
Vercel queries Supabase
lib/tools/orders.ts โ orders table
Returns: id, customer, status, items, totalโฆ
Returns: id, customer, status, items, totalโฆ
Security: The tool router /api/tool/[name] enforces an allowlist of exactly 12 tool names. Unknown tool names return 404 immediately.
Session Insights
A live plain-English summary of the call, updated automatically after every agent response.
How Session Insights are generated
Agent finishes speaking
Transcript updated with a new "assistant" entry
2-second debounce fires
useEffect reads latest transcript + toolHistory via refs (stale-closure-safe)
POST /api/summary
{ transcript: [...], toolHistory: [...] } sent to Vercel
OpenAI gpt-5-nano generates summary
Direct API call โ no gateway. max_tokens: 180. Returns plain-English paragraph.
Session Insights pane updates
"Sources accessed" chips reflect tools used so far in the session
Design note: An AbortController cancels any in-flight summary request when a new one starts, so rapid multi-turn conversations never produce out-of-order results.
Available Tools (12)
Functions the AI agent can invoke to retrieve or mutate data in real time.
lookup_order
Fetches full order details by order ID
โ orders table
search_customer_orders
Returns all orders for a given customer
โ orders + customers tables
search_order_items
Finds orders containing a specific product across history
โ orders + order_items tables
initiate_return
Starts a return process for an order item
โ orders table (status update)
cancel_order
Cancels an undelivered order
โ orders table (status update)
change_shipping_address
Updates delivery address for an active order
โ orders table
add_customer_note
Appends a note to an order record
โ orders table
lookup_product
Retrieves product info by name or SKU
โ products table
rag_query
Answers policy questions via semantic vector search
โ policy_chunks + pgvector
search_knowledge_base
Finds relevant help center articles
โ knowledge_base + pgvector
get_support_history
Retrieves past support tickets for a customer
โ support_tickets table
calculate
Safe arithmetic evaluator for totals, prices, comparisons
โ server-side (no DB)
RAG Pipeline
How the AI answers policy and knowledge-base questions with grounded, accurate responses.
Example โ "What is your return window?"
Customer voice query received
"What is your return window?"
AI invokes rag_query tool
lib/tools/knowledge.ts โ ragQuery()
OpenAI embeds the query
text-embedding-3-small โ 1536-dimension vector
Supabase pgvector nearest-neighbor search
match_policy_chunks(embedding, top_k=4) โ cosine similarity
Top 4 chunks returned โ AI speaks grounded answer
Seed: ~15 customers ยท ~25 orders ยท ~20 products ยท 4 policy docs ยท ~30 knowledge articles
Content Guardrails
Two independent layers prevent the agent from going off-topic or being misused.
Layer 1 โ System Prompt
Voice Agent Instructions
- Allowed scope defined: orders, products, shipping, returns, policies
- Named refusals: medical/legal advice, politics, roleplay, prompt injection
- Scripted redirect phrase for out-of-scope requests
- Abusive language handling (calm redirect, no escalation)
Layer 2 โ Suggest Endpoint Guard
/api/suggest Protection
- Fast regex pattern matching blocks obvious jailbreak strings
- OpenAI Moderation API checks for hate, harassment, violence, sexual content
- On API outage: fails open โ no false positives, users not blocked
- Every block logged to Vercel structured logging
Reference
Key Design Decisions
The reasoning behind architectural choices that aren't obvious from the code alone.
Why WebRTC direct to OpenAI instead of proxied?
Proxying audio through Vercel would add 100โ200ms per hop. WebRTC direct cuts the browserโAI path to a single network leg, giving near-real-time response feel that makes the demo convincing.
Why ephemeral client secrets instead of sending the API key to the browser?
The OpenAI key never leaves the server. The browser receives a 10-minute-TTL token that can only start one Realtime session โ eliminating any risk of key exfiltration or abuse.
Why push-to-talk (PTT) as primary instead of always-on VAD?
PTT gives the presenter full control during a demo โ no accidental triggers from background noise, audience questions, or side conversations. Server VAD (threshold 0.8, 1500ms silence) is enabled as a backup but PTT takes priority in the UI.
Why Supabase for the database?
pgvector support is built-in, enabling RAG without a separate vector store. The same database handles both structured data (orders, customers) and semantic search (policies, knowledge base), keeping the stack simple and the latency low.
Why build two modes (Self-Service vs Agent Assist)?
Together they demonstrate the full spectrum of AI roles in customer experience: full automation and human augmentation. Sharing the same WebRTC session infrastructure shows how one platform supports both patterns with minimal added complexity.
Why a dedicated calculate tool instead of letting the AI do the math?
Language models can hallucinate arithmetic. Routing all financial calculations through a server-side evaluator guarantees exact results โ important when discussing order totals, fees, or price comparisons with real customers.
Why gpt-5-nano (direct API) for session summaries instead of a gateway?
The summary is a single short text generation call (โค180 tokens). An AI gateway adds auth complexity and billing overhead without adding value here. A plain fetch to the OpenAI API using the existing key is simpler, faster, and cheaper.
Database Schema
PostgreSQL via Supabase with pgvector for semantic search.
customers (id, name, email, phone, address, created_at)
orders (id, customer_idโcustomers, status, items jsonb,
total_amount, shipping_address, notes, created_at)
order_items (id, order_idโorders, sku, name, qty, price_cents)
products (id, name, sku, description, price, inventory, specs jsonb)
support_tickets (id, customer_idโcustomers, order_idโorders,
subject, status, messages jsonb, created_at)
policy_chunks (id, source_file, content, embedding vector(1536))
knowledge_base (id, title, content, category, embedding vector(1536))
pgvector indexes on policy_chunks.embedding and knowledge_base.embedding enable sub-millisecond nearest-neighbor search at demo scale.
Environment Variables
CX-AI Architecture Reference ยท April 2026
โถ Launch Demo