Skip to main content
DevelopmentVoice AIEnterprise TelephonyAI / ML

3CX Call Control API with OpenAI Realtime Agent, built by Hoop.

A UAE insurance platform needed an AI voice agent sitting inside their enterprise 3CX PBX, not alongside it. Our team built the Python and FastAPI system that bridges 3CX's call control API with OpenAI's GPT-4o Realtime API via persistent WebSockets, handling bidirectional audio, OAuth token management, dynamic prompt compilation, and tool calls, all within a Dockerised deployment.

3CX Call Control API with OpenAI Realtime Agent
Client
Insurance Platform (UAE)
Industry
InsurTech / Voice AI
Location
United Arab Emirates
Timeline
2 Months, Aug 2025

In short

The 3CX Agentic AI Integration is a sophisticated Python-based system that bridges modern conversational AI with enterprise telephony infrastructure. It integrates OpenAI's GPT-4o Realtime API with a 3CX PBX system via WebSockets, enabling intelligent, real-time voice interactions, specifically configured for vehicle insurance sales callbacks in the UAE market.

  • Python and FastAPI backend bridging 3CX PBX with OpenAI GPT-4o Realtime API
  • Persistent WebSocket connections for low-latency bidirectional audio
  • Automatic audio resampling: 8kHz (3CX) ↔ 24kHz (OpenAI)
  • OAuth 2.0 token management with auto-refresh for uninterrupted 3CX API access
  • Dynamic Prompt Compilation: per-call system prompts from caller metadata
  • Docker containerisation for consistent deployment
The Challenge

Getting GPT-4o inside a 3CX enterprise PBX

Enterprise telephony platforms like 3CX are not built to hand audio to a cloud AI model. They use SIP trunks and their own call control APIs, output audio at telephony-standard 8kHz, and expect low-latency responses. OpenAI’s Realtime API runs at 24kHz over a WebSocket. Every part of the integration needed to reconcile those two worlds.

Beyond audio plumbing, the system needed to authenticate against 3CX’s OAuth 2.0 layer without disrupting active calls, track caller metadata across a conversation to inform the AI’s context, and execute tool calls and database actions mid-call without adding perceptible latency.

01

Bridging two incompatible audio protocols

3CX outputs telephony-standard 8kHz; OpenAI’s Realtime API operates at 24kHz. Real-time resampling in both directions was required.

02

Persistent low-latency connection to OpenAI

A phone conversation can’t tolerate connection setup overhead per utterance; a persistent WebSocket had to stay open for the call’s duration.

03

OAuth 2.0 token lifecycle for 3CX API

3CX access tokens expire. Auto-refresh logic had to run without interrupting in-flight API calls.

04

Per-call dynamic context for the AI agent

A static system prompt gives every caller the same generic start. The agent needed caller-specific context at session open.

05

Tool calls and database actions during live calls

The agent needed to fetch data, log outcomes, and trigger follow-up actions mid-conversation without stalling the call.

06

Consistent audio quality across callers

Volume varies dramatically between callers and devices; RMS normalization kept the agent’s responses consistent regardless of input level.

Our Solution

A five-component FastAPI backend as the integration layer

We built a Python and FastAPI backend that acts as the integration layer between 3CX and OpenAI’s GPT-4o Realtime API. When a SIP-trunked call arrives at 3CX, the Event Listener triggers the agent pipeline; audio is downsampled from 8kHz to the format OpenAI expects, sent via WebSocket, and the AI’s response is upsampled to 24kHz and streamed back to the caller. NumPy handles the resampling and RMS normalisation throughout.

The system is containerised with Docker for consistent deployment and carries a full OAuth 2.0 token management layer so 3CX API access never drops mid-call.

PythonFastAPIOpenAI GPT-4o Realtime API3CX Call Control APIWebSocketsNumPyDockerhttpx (HTTP/2)
The Architecture

How the call flows through the system

Incoming call routes through the SIP trunk to the 3CX PBX, call info and audio pass to the FastAPI backend, audio is downsampled to 8kHz and routed through the five backend components, then upsampled with instructions and sent at 24kHz to OpenAI's Realtime API. The output audio streams back to the caller.

Incoming call
SIP trunk → 3CX PBX
FastAPI backend (5 components)
OpenAI GPT-4o Realtime API
Audio streamed back to caller
Key Functional Components

Five modules inside the FastAPI backend

01 Event Listener

Call Detection

Monitors incoming calls from the 3CX system and triggers agent engagement on each new call event.

02 Audio Handler

Bidirectional Audio

Processes call audio in both directions: downsampling from 3CX (8kHz) for OpenAI and upsampling the AI’s response (24kHz) back to the line.

03 Realtime Agent Client

WebSocket to OpenAI

Manages the persistent WebSocket connection to OpenAI’s Realtime API, handling session lifecycle, audio streams, and conversation events.

04 Token Management

OAuth 2.0 Auth

Handles OAuth 2.0 authentication with 3CX, with automatic token refresh to ensure uninterrupted API access during active calls.

05 Participant Info System

Caller Context

Tracks caller metadata and integrates it into the AI’s context window, enabling personalised rather than generic interactions.

06 Dynamic Prompt Compilation

Per-Call Prompts

Generates a contextual system prompt for GPT-4o at the start of each call, using caller information and conversation requirements rather than a fixed static prompt.

The Full Feature Set

What the integration delivers across four areas

Call Handling

  • 3CX Call Control API Integration
  • SIP Trunked Call Routing
  • Event Listener System
  • Bidirectional Audio Streaming

Audio Processing

  • 8kHz ↔ 24kHz Auto Resampling
  • NumPy Audio Manipulation
  • RMS Normalisation
  • HTTP/2 Audio Transport (httpx)

AI Conversation

  • GPT-4o Realtime WebSocket Session
  • Prompt & History Management
  • Dynamic Prompt Compilation
  • Tool Calls & Actions

Infrastructure

  • Docker Containerisation
  • OAuth 2.0 with Auto Token Refresh
  • Database Handling
  • WhatsApp Follow-Up Capabilities

Technical depth

Functional Components In The Backend

6

Audio Sample Rate Conversion Each Call

2×

Enterprise Platforms Bridged

3CX + OpenAI

Build & Deployment

2 Months

The results

3CX PBX Integration

Enterprise

GPT-4o Realtime Voice

Real-Time

Docker Deployment

Containerised

Full Build Delivered

2 Months

Need an AI agent inside your enterprise phone system?

Tell us where you want to go. We'll show you how to get there — and then build it with you.

More case studies