Explainers

PBX high availability: what it covers and how to verify it

What a supplier means by high availability, what the SLA and recovery metrics actually commit to, and the architecture and connectivity questions to settle before you sign.

Cover graphic for a guide on cloud PBX continuity, from the datacentre to each company site

On a Monday morning, the phone system of a multi-site company stops distributing calls. Staff watch their handsets register and then drop off again. Queues go quiet and internal transfers fail. At a clinic, patients calling reception get nowhere. At a mid-sized company, sales, support and the coordination between sites all slow down.

PBX high availability therefore means more than keeping servers powered on. The phone system has to stay reachable, route calls, keep its configuration and accept changes from an administrator during an outage. That continuity depends on three things: the supplier's infrastructure, what the contract commits to, and the connectivity that actually links your users to the supplier.

These fail independently of each other. A redundant cloud PBX is out of reach from an office whose fibre has been cut. In the other direction, a backup internet line at the office won't help if the supplier runs its whole platform from one site. IT directors, integrators and telecom resellers need to assess the full chain, from the power supply in the datacentre to the last carrier connection at the building.

Why high availability matters for cloud telephony

Flow of the five functions a cloud PBX must keep running: registration, signalling, routing, configuration and admin access

The phone sits underneath almost every operational process, and most organisations only learn how much they depend on it when it stops. With reception down, inbound calls, voice menus, queues and transfers all stop at once. Staff can fall back on email or chat for a while, but customers, patients and suppliers still call the main number.

In a cloud PBX, availability covers several separate functions:

  • registration of desk phones, softphones and mobile apps

  • signalling between the platform and those endpoints

  • routing each call to the right person, queue or menu

  • the configuration, including extensions, opening hours and on-call rules

  • administrative access, so someone can change the service during the incident

Duplicated servers don't guarantee that all five fail over together, and a buyer usually finds out which one lags during the first real incident.

Rule of thumb: judge availability by the path a call takes, never by the status a server displays.

Server redundancy is one layer of several

A platform can run on several machines and keep a single point of failure in its power, cooling, storage, network or carrier connection. Maintenance creates the same risk. If several instances on separate servers depend on one shared component, taking that component offline takes all of them down.

The Uptime Institute's Tier standard describes how a facility is built. A Tier III facility is concurrently maintainable: each capacity component can be taken out of service on a planned basis without affecting IT processes. A Tier IV facility is fault tolerant, so each capacity component and distribution path can sustain a failure without affecting them. Some vendor material quotes 99.982% availability for Tier III and 99.995% for Tier IV. The Uptime Institute itself explains that it removed references to expected downtime per year from the Tier Standard in 2009, and that those references "were never a part of the Tier definitions". Today a tier tells you about power and cooling paths. It says nothing about the supplier's PBX software, its carriers, or the internet access at your own sites.

Three conditions to meet together

Resilient cloud telephony needs three conditions at once. The first is redundant infrastructure: sites, networks, components and power paths with separate dependencies. The second is a set of measurable objectives, meaning a service level agreement (SLA), a recovery time objective (RTO) and a recovery point objective (RPO), written into the business continuity plan and the disaster recovery plan. The third covers data and access. The data is hosted and processed in the European Economic Area, and a backup route exists for the day the main line goes down.

Many organisations also have a legal duty to plan for this. Article 21(2)(c) of the NIS2 directive, Directive (EU) 2022/2555, requires essential and important entities to include "business continuity, such as backup management and disaster recovery, and crisis management" in their cybersecurity risk-management measures. Telephony belongs in those plans whenever reception, on-call duty, medical coordination or a public service depends on it.

The metrics that describe availability

Comparison of 99.9%, 99.95% and 99.99% SLA levels, with yearly downtime converted from each percentage and typical use

You can't compare two offers on an availability percentage alone. A supplier can advertise a high SLA and leave the buyer with no clear answer on how long a restore takes, how much data is lost, or how the switchover works. Read the figures below as one set.

SLA, RTO and RPO

The SLA is the supplier's contractual commitment. It defines the availability covered, what gets measured, the exclusions and any compensation when the target is missed. Check the measurement method, the notification deadlines and the conditions for compensation. Without those three, a percentage in a contract is hard to enforce.

The RTO is the target time to bring the service back after an incident. The RPO is the maximum amount of data the organisation accepts losing, counted back from the failure to the last usable copy.

The French cloud provider DoliCloud shows why you need both, scenario by scenario. Its published SLA, in French, targets an RPO of 2 minutes and an RTO of 50 minutes for a limited server failure. For a major disaster, such as a fire in the datacentre, the same page gives an RPO of 24 hours and an RTO of 48 hours. Both scenarios sit under one headline availability of 99.9%, excluding planned maintenance, so a buyer comparing headline figures would never see the 48-hour case.

SLA levelDowntime per yearTypical use
99.9%About 8 hours 46 minutesStandard service, moderate criticality
99.95%About 4 hours 23 minutesImportant business service
99.99%About 52 minutesCritical telephony, multi-site operations

The durations in this table are arithmetic conversions of the percentages, for orientation. They aren't contract terms. A tender should ask the supplier to confirm the applicable figure, the measurement window and the exclusions.

MTBF and MTTR

Mean time between failures (MTBF) is the average time between two faults. It indicates how reliable a component or service has been, and it can't promise that no incident will happen. Mean time to repair (MTTR) is the average time needed to restore the function after a fault.

For a PBX, ask for MTTR per type of failure. Restoring an application node, losing a whole site, losing a carrier and losing the customer's own line each follow a different procedure. Also check whether calls in progress survive a switchover, whether queues and menus are replicated, and whether administration stays accessible during the incident.

Audio quality is part of availability too. A service that connects calls with broken audio is unusable for a conversation. Our guide to VoIP call quality covers the factors to put in the specification next to uptime. In practice, an available phone system is one that people can reach, that routes calls correctly, and that carries audio you can follow.

Technical architectures for resilient telephony

A sound architecture removes shared dependencies first and adds sophisticated mechanisms after that. Geo-redundancy, load sharing, replication and failover have to work as one system. A secondary server in the same technical zone doesn't create geographic resilience.

Geo-redundancy and replication

Geo-redundancy spreads services and data across separate sites. It protects against a local incident that affects power, cooling, the network or access to the building. The supplier should list exactly what it replicates, item by item: accounts, extensions, routing rules, queues, interactive voice response (IVR) menus, opening hours, recordings and administration data.

Replication runs active-passive or active-active. In the first model, a primary site handles the calls while a secondary site waits for the switchover. In the second, several nodes take part in the service and each carries a share of the traffic. Active-active uses the resources more fully. It also requires stricter control of data consistency and of what happens during a network partition.

Failover and load balancing

Automatic failover relies on health checks that can tell three cases apart: a slow server, an unreachable service and a complete failure. When a node fails, the system removes it from the traffic and sends new sessions elsewhere. It also has to stop sending calls to a node that still responds at network level after its signalling has stopped working.

Load balancing distributes connections across several servers and removes bottlenecks, though it won't repair a badly designed data layer. The balancer itself has to be redundant, and so do name resolution, monitoring and network distribution.

Backup and restore

Replication isn't a backup. A corrupted configuration or an accidental deletion spreads to every node, and then every copy holds the same error. Backups have to be isolated, encrypted and checked. Someone also has to restore them into a test environment at regular intervals.

Until someone has tested the failover, it's an assumption about how the system should behave. The integrator documents the expected behaviour for a node failure, a site failure, a storage failure and a connectivity failure. The procedure names who triggers the switchover, who checks inbound calls afterwards and who authorises the return to normal operation.

Data sovereignty and connectivity resilience

Residency and sovereignty

A datacentre in the European Union settles where the data is stored. That is residency. Sovereignty also depends on which law applies to the data and which authority can demand its disclosure, and the two answers can differ for the same server.

United States law gives a concrete example. Under 18 U.S.C. § 2713, added in March 2018 by the CLOUD Act, a provider of electronic communication or remote computing services that is served with a warrant or other legal process under that chapter must preserve, back up or disclose the customer content and records in its possession, custody or control "regardless of whether such communication, record, or other information is located within or outside of the United States". A datacentre in Europe run by a provider that falls under that law is in scope.

For a cloud PBX, the analysis covers calls, recordings where you use them, logs, metadata, backups and administration tools. Metadata deserves attention because call detail records show who called whom, when and for how long. The CISPE Data Protection Code of Conduct for cloud infrastructure service providers, for which the French data protection authority (CNIL) is the designated supervisory authority, requires an adhering provider to let customers store and process their data entirely within the European Economic Area. With that option, Chapter V of the GDPR, which governs transfers to third countries, doesn't come into play, and it's an option you can ask a supplier to confirm in writing.

A nearby datacentre can still go down

A datacentre close to your sites can lower latency. It doesn't lower the risk of an event that takes out a whole region.

On 25 April 2023, a cooling water pipe leaked in a non-Google part of a building used by Google Cloud's europe-west9 region in Paris. The water reached an uninterruptible power supply (UPS) room and caused a fire. The fire brigade attended, the building was evacuated and its power was shut down. According to Google's incident report, europe-west9 has three buildings with independent cooling, power and networking, and only one of them was hit. The whole region still went down. Regional Spanner, the database behind the region's control plane, kept two of the replicas it needed for quorum in the building that lost power, which Google calls a misconfiguration. The report records the incident from 25 April to 26 April, with pumping of the standing water starting on 29 April. On 10 May, The Register reported that the facility "remains offline" two weeks after the leak, with no estimate for the full recovery of instances in the europe-west9-a zone.

Separate buildings protect you only if every critical dependency is spread across them. That includes the databases and control systems the supplier runs behind the scenes, so ask where its second site is, how far it is from the first, and what the two sites still share.

The local fibre blind spot

The PBX can be fully available at the supplier and unusable at a site whose fibre is cut. Calls can no longer take their usual path, handsets lose registration and users can't reach queues or menus.

Plan a fallback that matches how critical each site is. The options are a 4G or 5G backup connection, a second carrier on physically separate infrastructure, or a controlled diversion of calls to mobiles. Before you sign with a second carrier, check that its line doesn't run through the same duct as the first one. For a multi-site estate, a software-defined wide area network (SD-WAN) can steer traffic across several paths, provided the backup links are genuinely separate.

Implementation checklist for IT directors and integrators

Six-step checklist for PBX continuity, from mapping the current setup to testing the switchover and recording the results

Start with the dependencies, before anyone chooses an interface. The IT department and the integrator build a map the operations team can work from: numbers, sites, carriers, equipment, routing rules and escalation contacts.

From audit to objectives

  1. Map what exists. List the internet access at each site, the current telephony, the SIP trunks, the critical sites, the equipment on an uninterruptible power supply and the network dependencies. The result separates single points of failure from components that are genuinely redundant.

  2. Set the business objectives. The business continuity plan states which calls must still work in degraded mode, for how long and in what order of priority. An emergency service, a patient reception desk and an administrative switchboard don't necessarily share the same RTO.

  3. Choose the fallback model. Compare geo-redundancy, active-passive, active-active, a second carrier connection, 4G or 5G, and diversion to mobiles. Pick the one that fits the risk, the budget and the skills available.

Configure, test, document

  1. Replicate the configuration. Check extensions, queues, IVR menus, calendars, exception schedules and on-call rules. A backup that holds only user accounts can't rebuild the full service.

  2. Test the switchover. Simulate a server failure, a site failure and the loss of the main line. Test inbound calls, outbound calls, transfers, queues, announcements and administrative access. Pinging the servers doesn't test any of this.

  3. Record the outcome. Write down the results, the gaps and who owns each one. The procedure states when to use the backup line, how to inform users and how to return to the normal path without creating a routing loop.

Validate the firewall rules before go-live. The SIP signalling ports, the Real-time Transport Protocol (RTP) media range and Transport Layer Security (TLS) need to pass on the backup path as well as the main one. Voxbi's firewall settings give the integrator a starting point to adapt to the equipment actually deployed.

Questions to ask your cloud PBX supplier

An effective tender asks for answers you can verify. The supplier should be able to describe its architecture, its contractual scope and how it handles an incident.

AreaKey questionExpected answerWarning sign
InfrastructureWhere are the sites and what dependencies do they share?Locations, zone separation and a description of the critical pathsA single site presented as redundant, with no detail
ResilienceHow does failover work?Detection, trigger, what switches over and the failback procedureFailover described as automatic, with no documented test
DataWhere are data and backups hosted and processed?European Economic Area, data categories, sub-processors and GDPR safeguardsAn answer that gives only the physical location of the server
SecurityWhat protects signalling and sessions?SIP TLS and WebRTC encryption, access management and loggingPartial encryption or missing documentation
ConnectivityWhat happens if the local fibre goes down?Diversion, secondary access, 4G or 5G, with clear responsibilitiesThe supplier refers you to the carrier with no scenario
ContractWhat does the SLA cover and what compensation applies?Scope, measurement, exclusions, notification and penaltiesA percentage with no RTO, RPO or calculation method
OperationsWhen are recovery tests run?Scheduled tests, written reports and corrective actionsNo evidence of testing, or monitoring offered in place of tests

The data question has to cover recordings and metadata as well as account information. Choosing a cloud service also means assessing transfers, sub-processor access and the applicable law, which is part of choosing a cloud phone system in general.

For integrators, the supplier should hand over a file the IT department can use: the logical architecture, responsibilities, escalation contacts, the limits of the SLA, the conditions for diversion and the test results. I would take an open account of past incidents over a theoretical availability figure that nobody outside the supplier can check.

Putting PBX high availability into your recovery plan

High availability for a cloud PBX rests on three conditions that go together. The architecture removes single points of failure, and its switchover mechanisms are tested. The commitments combine an SLA, an RTO and an RPO with concrete recovery procedures. The data and its processing meet European sovereignty requirements and the GDPR.

A cloud supplier can't deliver a service to a site that can no longer reach it, so carrier connectivity belongs in the same plan. Every organisation should document its main access, its fallback and how calls should behave in degraded mode.

You can use the implementation checklist and the question grid straight away, during an audit, a contract renewal or a PBX replacement. The telephony part of your disaster recovery plan then belongs in the wider business continuity plan, with exercises and an update after every change to the architecture.


Voxbi provides a cloud PBX hosted in European Union datacentres, with web traffic over HTTPS and SIP TLS covering SIP devices and WebRTC. IT directors, integrators and telecom resellers can write to hello@voxbi.com to assess a cloud PBX architecture, its continuity options and how it fits into a recovery plan.

See Voxbi in your business.

Talk to us or to a certified Voxbi partner.