"99.99% availability" in a specification: is the cable cut in the calculation, or not?
6 min read
“99.99% availability” turns up in tenders and specifications as if it were a fact. It isn’t: it is the result of a calculation, and a calculation has inputs that can — and should — be stated. ITU-T G.911 (04/97), Parameters and calculation methodologies for reliability and availability of fibre optic systems, is the Recommendation that fixes those parameters and that method: it is nearly thirty years old and, we checked on the official ITU page, it is still in force, while the previous edition, from March 1993, is superseded. No later text has replaced it. Whoever writes “99.99%” without saying how they got there is not stating a commitment: they are writing a number.
A figure is a calculation, not a promise
G.911 defines availability performance like this: “The ability of an item to be in a state to perform a required function at a given instant of time or at any instant of time within a given time interval, assuming that the external resources, if required, are provided” (clause 3.3). The closing condition is not decorative: a note to the same definition spells it out without ambiguity — “In the definition of the item, the external resources required must be delineated” (clause 3.3, note 2). Availability is not a property of the cable or of the equipment: it is a property of the whole system, maintenance included — and the method brings that into focus at the very first step of the calculation. Among the inputs the Recommendation lists for Step 1 (clause 6.2.1) are “repair rates based on MTTRs of the various elements and sub-systems”. Three items to demand, not one figure to accept: which MTTR, whether the cable is in the model, which protection architecture. One at a time.
MTTR is not the time it takes to swap a card
G.911’s definition wrongfoots anyone who pictures MTTR as the time of a technical intervention and nothing more: “mean time to repair (MTTR): The total corrective maintenance time divided by the total number of corrective maintenance actions during a given period of time” (clause 3.18). And “maintenance time”, in the same Recommendation, closes off every loophole: “The time interval during which a maintenance action is performed on an item either manually or automatically, including technical delays and logistic delays” — it includes technical delays and logistic delays (clause 3.17). Standby call-out, travel to the access pit, digging permits, night work if it comes to that: under the Recommendation, they sit inside MTTR, not outside it. It should not be confused with MTBF — “mean time between failures (MTBF): The expectation of the time between failures” (clause 3.6) — which measures how often an item fails, not how long it takes to repair it: two different figures, and availability depends on both.
How much the gap between an equipment MTTR and a cable MTTR actually weighs is shown by the Recommendation’s own access-network case study (Appendix III, Table III.1): the repair rate of the OLTM, the line terminal, is 1-2 per day; the cable’s is 0.5-1 per day. The text defines repair rate as the inverse of the average time taken to repair a failure — “repair rate (inverse value of average time taken to repair a failure)” — so the implied MTTR in the example is 12-24 hours for the equipment and 24-48 hours for the cable — the conversion into hours is ours, the recommendation gives the rates, not the hours: double, and that is still just the technical intervention, before counting permits and call-out on a real urban site.
Is the cable in the model, or has it been left out
The failure-prediction method in G.911 (clause 4) is a parts-count method applied to hardware: the Recommendation states this without qualification — “The scope is limited to hardware-related failures under steady-state conditions” (clause 4.1). It is a clean method, built for equipment and electronics. The cable follows a different logic, and G.911 says so at the very start of its own reliability parameters for fibre and cable: “Extrinsic failure modes, such as cable dig-ups or installation/maintenance errors, dominate the overall reliability” — extrinsic failure modes, meaning digging accidents that sever the cable or installation and maintenance errors, dominate overall reliability; intrinsic modes, such as fibre fatigue, remain a small fraction (clause 2.4). A calculation that applies parts-count to the equipment alone describes a secondary failure with great care and ignores the one that, on the Recommendation’s own admission, weighs more. This is not a detail specific to any one urban installation: G.911 says it of the access network in general, where — “users do not have alternate path connections via another central office” — the customer has no alternative route to another exchange, unlike connections between exchanges (Appendix III). That is why the same appendix opens by citing the “need for line terminals and cable redundancy to reduce service interruptions in the event of cable cut or line equipment failure”: redundancy is needed on line terminals and on the cable, for a cable cut just as much as for equipment failure.
Which protection, and whether the standby route shares the run
The reason redundancy gets designed in at all, G.911 puts in one line: “the net availability performance obtained by cascading large numbers of network elements is incompatible with the expectations of many customers” (clause 6.1). But the calculation for a protected network only holds if one assumption is true: that failures on the two routes are independent events. The same Appendix III states it explicitly for cable redundancy: “Cable redundancy assumes an alternate path is provided via a physically diverse route” — and, describing cable protection, adds: “This alternative assumes that working and protection fibres follow physically separate paths to the same CO”. “Assumes”: it is a modelling assumption, not an automatic description of what is in the ground. If the standby route shares the same duct, the same bridge or the same access pit as the primary route — as we wrote about subsea cables — a single dig can cut both at once, independence collapses, and the figure calculated on that assumption no longer describes the network that was actually built.
Applied to the specification
A declared availability figure enters the specification with its three inputs, exactly as the calculation behind an optical budget is demanded instead of a declared class on its own: the assumed MTTR and what it is built from — equipment or cable, with or without call-out and permits —, the presence of the cable cut alongside equipment failure, the type of protection and the declared physical independence of the two routes. At acceptance, that independence gets checked for real: the standby route’s path on the ground, not just on the design drawing. And every failure that follows — date, cause, actual time to restore, the run affected — feeds a measured MTTR that, over time, replaces the one assumed at tender, inside the same map the network is designed on: an AI flags where the promised availability does not hold and which run brings it down, and the team — ours, together with CSIDIA — steps in. Always within the client’s perimeter: on autonomous on-premise machines that need no deep integration into the existing network, or in a dedicated cloud with a data centre in Italy; always with shared management.
Do you need to write an availability clause into a specification, or check one you have been handed without its inputs? Talk to an engineer: the first session comes at no cost, and the assumed MTTR fits in one line, written before anyone signs.