By the time an SLA Report shows a breach, the period is already over and there's nothing left to do but explain it. Predictive SLA Breach exists to catch the problem while there's still time to act.
How it works
Every SLA target has an implicit downtime "budget" for the period — a 99.9% monthly target on a 30-day month allows roughly 43 minutes of downtime before it's breached. As the period progresses, iPM tracks how much of that budget has already been used and projects, based on the current pace, whether the host or service is on track to stay within it or is likely to breach before the period ends.
Reading the projection
A service that's used a disproportionate share of its downtime budget early in the period — for example, 80% of its allowed downtime used in the first ten days of a thirty-day month — gets flagged as at risk, well before the SLA report would otherwise confirm a breach. This is a projection based on downtime consumed so far, not a guarantee of what will happen, but it's a useful early-warning signal precisely because it arrives while there's still budget left to protect.
What to do with an at-risk flag
- Investigate whether the downtime so far has one identifiable, fixable cause (see Alert History for the specific events behind it)
- Treat any further planned maintenance for that host extremely carefully for the rest of the period — a change that would normally be low-risk carries more weight when there's little downtime budget left
- Consider whether the target itself was realistic in the first place, if a host is chronically at risk period after period, since that's a different problem than any single incident
Predictive SLA Breach is most useful reviewed regularly (weekly, at minimum) rather than only checked once — a projection made on day 2 of a period tells you far less than one made on day 20.