FluxBilling

SLA Credits You Can Actually Automate

Most hosting SLAs are marketing documents never designed to be executed. They work until a real outage, when you find out the promise is either unenforceable or ruinous.

Mario MarinMario Marin6 min read

Most hosting SLAs are marketing documents that were never designed to be executed. They promise a number, define the remedy vaguely, and rely on almost no customer ever claiming. That works until a real outage, at which point you discover your SLA is either unenforceable or ruinous, and nobody is sure which.

Here is how to write one you can actually automate, and what your billing system needs to do with it.

Not legal advice. SLA terms interact with consumer protection law, contract law and your liability position. Have the document reviewed by a qualified adviser before you publish it.

Define the thing you are measuring

"99.9% uptime" is meaningless until four questions are answered, and every dispute lives in the gaps.

Uptime of what? The network, the power, the hypervisor host, or the customer's service? A provider measuring network availability and a customer measuring whether their site responded are measuring different things and will disagree honestly.

Measured how, and by whom? Your monitoring is the practical answer and the customer will not fully trust it. Naming the measurement method and making the data visible removes most of the argument before it starts.

Over what period? 99.9% monthly allows about 43 minutes; annually it allows about 8.7 hours and lets you have a bad month with no consequence. Monthly is the customer-friendly and the more common choice.

What is excluded? Scheduled maintenance with proper notice, customer-caused outages, force majeure, and issues in the customer's own software. Exclusions are legitimate and necessary — but an SLA whose exclusions swallow the promise is worse than no SLA, because it reads as bad faith the first time it is invoked.

Make the remedy specific

The remedy is where vague SLAs become unenforceable. Define it as a table: a band of availability, and a percentage of that service's monthly fee credited. Something like a 10% credit below 99.9%, 25% below 99.5%, 50% below 99%, and 100% below 95%.

The exact numbers are yours. What matters is that a person — or a script — can read the availability figure and produce the credit without judgement.

Then settle the four boundary questions in writing: is the credit capped at the monthly fee for that service (it should be), is it a credit against future invoices rather than a cash refund (usually yes, and say so plainly), must the customer claim it within a window (common, and reasonable if the window is not absurdly short), and does it apply to the affected service only or the whole account (the affected service).

Claim-based or automatic?

The honest split in this industry.

Claim-based is nearly universal, and the reason is that almost nobody claims. It is cheap, and it quietly converts your SLA into a document rather than a commitment. It is also the model most likely to annoy the customer who does claim, because they now have to prove your outage to you.

Automatic means you calculate the credit and apply it without being asked. It costs more in credits and it is a genuine differentiator, because it says the number in the SLA is real. For business and colocation customers it is a meaningful trust signal, and it removes an entire category of support interaction.

If you go automatic, model the cost first against your actual historical availability rather than your target. If your worst month last year would have cost you 25% of a segment's revenue, you need either better infrastructure or a more conservative table before you publish it.

What automation requires

Four things have to connect, and the chain breaks at the same place in most stacks.

  1. Monitoring that records availability per service, not per host or per rack. A credit is owed against a specific customer's specific service, so the availability figure has to attach to that service.
  2. A mapping from monitored object to billable service. This is the break point. Monitoring usually knows IP addresses and hostnames; billing knows customers and services. If nothing authoritatively joins them, the credit calculation is manual forever.
  3. A calculation that runs on the billing period boundary and produces a credit amount per affected service.
  4. A credit mechanism that applies it correctly as a document rather than an ad-hoc adjustment — see credit notes, refunds and clawbacks.

Step two is the one worth checking before you promise anything automatic. In a stack where monitoring and billing are separate products, that join is an integration you own; where they share records, it is a relationship that already exists.

Maintenance windows are part of the SLA

Excluded maintenance is only excluded if you actually gave the notice your SLA requires. That means the notice period is an operational commitment, not a clause — and emergency maintenance needs its own defined treatment, because "emergency" is otherwise an exclusion with no boundary.

Publish scheduled maintenance somewhere customers can see it, and keep the record. When a customer disputes an outage that was a maintenance window, the published notice is the entire argument.

How FluxBilling fits

FluxBilling holds services, customers, invoices, hardware inventory and network devices in one schema, which is what makes step two above a relationship rather than an integration — a monitored device is connected to the rack position and the service it belongs to. Network device tracking with SNMP monitoring is part of the DCIM layer on every tier, and credit notes are first-class documents that reference the invoice they adjust rather than being an untracked balance tweak.

Being straight about scope: FluxBilling does not ship an automated SLA credit engine. There is no built-in availability-to-credit calculation that runs on the period boundary and issues credits unprompted. The records the calculation needs are in one place, and the credit mechanism exists, but the rule itself is something you would build against them — realistically a plugin-builder project. If automated SLA credits are a requirement for you, ask us for specifics rather than assuming.

See DCIM and network monitoring for hosting providers.

Closing thoughts

Take your current SLA and try to compute what you owed for your worst month last year. If you cannot do it — because availability was never recorded per service, or the remedy is too vague to apply — then your SLA is a marketing document, and the moment a serious customer's lawyer reads it properly is the wrong moment to find out.

Tagged
hosting SLA creditsuptime SLA hostingservice level agreement hostingautomated SLA creditSLA remedy tablemaintenance window SLA
Written by
Mario Marin
Mario Marin
View all posts →