3CX Call Control API with OpenAI Realtime Agent, built by Hoop.
A UAE insurance platform needed an AI voice agent sitting inside their enterprise 3CX PBX, not alongside it. Our team built the Python and FastAPI system that bridges 3CX's call control API with OpenAI's GPT-4o Realtime API via persistent WebSockets, handling bidirectional audio, OAuth token management, dynamic prompt compilation, and tool calls, all within a Dockerised deployment.

- Client
- Insurance Platform (UAE)
- Industry
- InsurTech / Voice AI
- Location
- United Arab Emirates
- Timeline
- 2 Months, Aug 2025
In short
The 3CX Agentic AI Integration is a sophisticated Python-based system that bridges modern conversational AI with enterprise telephony infrastructure. It integrates OpenAI's GPT-4o Realtime API with a 3CX PBX system via WebSockets, enabling intelligent, real-time voice interactions, specifically configured for vehicle insurance sales callbacks in the UAE market.
- Python and FastAPI backend bridging 3CX PBX with OpenAI GPT-4o Realtime API
- Persistent WebSocket connections for low-latency bidirectional audio
- Automatic audio resampling: 8kHz (3CX) ↔ 24kHz (OpenAI)
- OAuth 2.0 token management with auto-refresh for uninterrupted 3CX API access
- Dynamic Prompt Compilation: per-call system prompts from caller metadata
- Docker containerisation for consistent deployment
Getting GPT-4o inside a 3CX enterprise PBX
Enterprise telephony platforms like 3CX are not built to hand audio to a cloud AI model. They use SIP trunks and their own call control APIs, output audio at telephony-standard 8kHz, and expect low-latency responses. OpenAI’s Realtime API runs at 24kHz over a WebSocket. Every part of the integration needed to reconcile those two worlds.
Beyond audio plumbing, the system needed to authenticate against 3CX’s OAuth 2.0 layer without disrupting active calls, track caller metadata across a conversation to inform the AI’s context, and execute tool calls and database actions mid-call without adding perceptible latency.
Bridging two incompatible audio protocols
3CX outputs telephony-standard 8kHz; OpenAI’s Realtime API operates at 24kHz. Real-time resampling in both directions was required.
Persistent low-latency connection to OpenAI
A phone conversation can’t tolerate connection setup overhead per utterance; a persistent WebSocket had to stay open for the call’s duration.
OAuth 2.0 token lifecycle for 3CX API
3CX access tokens expire. Auto-refresh logic had to run without interrupting in-flight API calls.
Per-call dynamic context for the AI agent
A static system prompt gives every caller the same generic start. The agent needed caller-specific context at session open.
Tool calls and database actions during live calls
The agent needed to fetch data, log outcomes, and trigger follow-up actions mid-conversation without stalling the call.
Consistent audio quality across callers
Volume varies dramatically between callers and devices; RMS normalization kept the agent’s responses consistent regardless of input level.
A five-component FastAPI backend as the integration layer
We built a Python and FastAPI backend that acts as the integration layer between 3CX and OpenAI’s GPT-4o Realtime API. When a SIP-trunked call arrives at 3CX, the Event Listener triggers the agent pipeline; audio is downsampled from 8kHz to the format OpenAI expects, sent via WebSocket, and the AI’s response is upsampled to 24kHz and streamed back to the caller. NumPy handles the resampling and RMS normalisation throughout.
The system is containerised with Docker for consistent deployment and carries a full OAuth 2.0 token management layer so 3CX API access never drops mid-call.
How the call flows through the system
Incoming call routes through the SIP trunk to the 3CX PBX, call info and audio pass to the FastAPI backend, audio is downsampled to 8kHz and routed through the five backend components, then upsampled with instructions and sent at 24kHz to OpenAI's Realtime API. The output audio streams back to the caller.
Five modules inside the FastAPI backend
01 Event Listener
Call Detection
Monitors incoming calls from the 3CX system and triggers agent engagement on each new call event.
02 Audio Handler
Bidirectional Audio
Processes call audio in both directions: downsampling from 3CX (8kHz) for OpenAI and upsampling the AI’s response (24kHz) back to the line.
03 Realtime Agent Client
WebSocket to OpenAI
Manages the persistent WebSocket connection to OpenAI’s Realtime API, handling session lifecycle, audio streams, and conversation events.
04 Token Management
OAuth 2.0 Auth
Handles OAuth 2.0 authentication with 3CX, with automatic token refresh to ensure uninterrupted API access during active calls.
05 Participant Info System
Caller Context
Tracks caller metadata and integrates it into the AI’s context window, enabling personalised rather than generic interactions.
06 Dynamic Prompt Compilation
Per-Call Prompts
Generates a contextual system prompt for GPT-4o at the start of each call, using caller information and conversation requirements rather than a fixed static prompt.
What the integration delivers across four areas
Call Handling
- 3CX Call Control API Integration
- SIP Trunked Call Routing
- Event Listener System
- Bidirectional Audio Streaming
Audio Processing
- 8kHz ↔ 24kHz Auto Resampling
- NumPy Audio Manipulation
- RMS Normalisation
- HTTP/2 Audio Transport (httpx)
AI Conversation
- GPT-4o Realtime WebSocket Session
- Prompt & History Management
- Dynamic Prompt Compilation
- Tool Calls & Actions
Infrastructure
- Docker Containerisation
- OAuth 2.0 with Auto Token Refresh
- Database Handling
- WhatsApp Follow-Up Capabilities
Technical depth
Functional Components In The Backend
6
Audio Sample Rate Conversion Each Call
2×
Enterprise Platforms Bridged
3CX + OpenAI
Build & Deployment
2 Months
The results
3CX PBX Integration
Enterprise
GPT-4o Realtime Voice
Real-Time
Docker Deployment
Containerised
Full Build Delivered
2 Months
The services behind this integration
Need an AI agent inside your enterprise phone system?
Tell us where you want to go. We'll show you how to get there — and then build it with you.