How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message, or ask a question in your own words.
Escalation: when nobody picks up an alarm
A notification rule tells the team that should deal with a fault. Escalation is the backup for when nobody does: if the alarm is not acknowledged in time, it goes on to somebody else, and on again if need be, until somebody says they have it. It is set up per rule, so only the alarms that matter escalate.
Two kinds of rule · Setting it up · Acknowledging · How an alarm climbs · What stops it · The messages · Testing
Two kinds of rule
Every rule under Setup, Actions has a Kind.
| Kind | What it is |
|---|---|
| Notification | The everyday rule. It is bound to part of the estate, is about the states you tick, may have a schedule, and tells the people who should be dealing with a fault. It may escalate. |
| Escalation | A backup. It is bound to nothing and has no states or schedule of its own: it is reached only when another rule's alarm has not been acknowledged in time, and sends whatever that rule escalates. It has its own channel, recipients and repeats, and may escalate again. It has no schedule because an escalation that can be silenced is no escalation. |
In the rules list an escalation rule shows escalation in place of its states and bindings, and a rule that escalates carries a small climbing icon beside its name. An escalation rule's page lists, under Reached from, the rules that escalate to it.
Setting it up
A rule can only escalate to an escalation rule that already exists, so build the chain from the top down. For example, an on-call engineer behind the NOC, and the duty manager behind them:
- Create
Duty manager, Kind Escalation, on SMS to the duty manager's mobile, sending once. - Create
On-call, Kind Escalation, on SMS to the on-call mobile, sending up to 3 times every 5 minutes, and set it to escalate toDuty managerafter 15 minutes. - On the notification rule,
NOC, set it to escalate toOn-callafter 10 minutes, for CRIT.
| Field | What it does |
|---|---|
| If not acknowledged, escalate to | Nobody, or one of your escalation rules. A rule cannot escalate to itself, and a chain that would come back round on itself is refused. |
| After | From 5 minutes to 4 hours: how long the alarm waits at this level for somebody to acknowledge it before it climbs to the next. |
| Escalate | Notification rules only. Which of the rule's states are serious enough to escalate, CRIT unless you choose otherwise. A rule about WARN and CRIT that escalates CRIT only tells the team about both, and escalates only the critical. |
Acknowledging
Acknowledging says I have it and I am on it. It is done on the dashboard, where each rule's alarm awaiting acknowledgement sits above Current alarms with an Acknowledge button, or on the phone view, which has the same button. Each card shows the rule, how long ago it sent, how many objects are in alarm, and what happens next, such as Escalates to On-call in 4m.
- It stops everything at once. Whoever acknowledges, at whatever level the alarm has reached, the whole chain stops: the escalation rules above and below alike, and the notification rule's own repeats of what it has already sent.
- It does not change the estate. The sensors stay in alarm on the wheel, and when they recover the clearance is sent as usual.
- Something new afterwards is a new alarm, with its own clock. Acknowledging the first fault does not cover a second one an hour later.
- A reply to the message does nothing. Answering an email, a Matrix message or an SMS does not acknowledge the alarm. Use the dashboard or the phone view.
This is not the same as acknowledging a sensor under Current alarms. That acknowledges one object, for a period you choose, and it also keeps that object out of the escalation. In a large outage, acknowledging the rule's alarm is one press however many sensors are failing. Pausing the affected part of the estate, for planned work say, stops the alarms and their escalation as well.
How an alarm climbs
The clock starts when the notification rule sends its alarm. If nobody acknowledges it within that rule's After, the alarm climbs to the escalation rule. That rule sends, repeats on its own settings, and after its own After climbs to the next, and so on. The figure follows the example above, with the NOC repeating every 15 minutes and the alarm acknowledged at 40 minutes.
| Minute | NOC | On-call | Duty manager |
|---|---|---|---|
| 0 | Alarm | ||
| 10 | Escalated, first send | ||
| 15 | Repeat | Repeat, second send | |
| 20 | Repeat, third and last send | ||
| 25 | Silent from here on | Escalated, first send | |
| 30 | Repeat | Repeat, second send | |
| 35 | Repeat, third and last send | ||
| 40 | Acknowledged on the dashboard: nothing further is sent by any of them | ||
| 45 | Repeat not sent | ||
The rules of the climb
- It only goes up. Once the alarm has climbed past an escalation rule, that rule never sends for it again, whatever its repeat settings, and the alarm never comes back down to it.
- Only the level the alarm is at repeats. Two escalation rules never repeat side by side.
- Repeats have to fit inside After. On an escalation rule that climbs further, a repeat only happens if it falls before the climb, so the page refuses a repeat interval as long as After or longer. Sends beyond what fits are simply never made: On-call above could send three times in its 15 minutes, but not five.
- A climb beats a repeat. When both fall due at the same moment, the alarm climbs and the repeat is not sent, so nobody gets two messages for one step.
- The top of the chain stops on its own. The last escalation rule repeats up to its send count and then goes quiet. The alarm stays on the dashboards until somebody acknowledges it or it clears.
- The notification rule keeps going. Its team is the one that should be dealing with the fault, so it goes on repeating at its own interval, up to its own send count, while the alarm climbs. The escalation rules are the backups, not a replacement.
Recipients do not carry up
Because each escalation rule falls silent once the alarm climbs past it, the people on it hear nothing more about
that alarm from it. If the on-call engineer should still be chased once the duty manager has been called, put the
on-call number on Duty manager as well. List on each escalation rule everyone who should be hearing
about it at that level.
What stops it
- An acknowledgement, on either dashboard, at any point.
- Everything clearing. When nothing in the alarm is still in the rule's states, the alarm closes within a minute and nothing further climbs. The clearance goes to the notification rule's recipients as usual; the escalation rules do not send one.
- The objects being dealt with in the estate. Acknowledging the failing sensors under Current alarms, or pausing, suspending or scheduling off that part of the estate, takes them out of the alarm. When nothing is left in it, it closes.
- Nothing serious enough. If what is still wrong is in none of the states the rule escalates, a WARN on a rule escalating CRIT only, the alarm waits where it is. Should something then go critical, the climb starts again from the first escalation rule, with a fresh clock.
- A rule switched off. Switching off or deleting the notification rule ends its alarms. An escalation rule that has been switched off stops the climb at the level below it, and the dashboard card says so by name, as Not escalated: On-call is switched off.
Schedules
A notification rule outside its schedule sends nothing, and an alarm it has not sent cannot escalate. Escalation rules have no schedule, so once an alarm is climbing it climbs whatever the hour.
The messages
An escalation goes out through the escalation rule's channel, template and recipients, listing what is still wrong
in the alarm as new, with charts and the poll log where the channel carries them. Its kind is
ESCALATION, and it goes out under the name people know, the notification rule's, so the default
templates show Oversight ESCALATION: NOC. A ticket channel works as an escalation too, and raises one
ticket when the alarm reaches it.
| Placeholder | Gives, on an escalation |
|---|---|
{{kind}} | ESCALATION |
{{rule}} | The notification rule nobody acknowledged, such as NOC |
{{tier}} | The escalation rule sending this message, such as On-call |
{{raised}} | When the notification rule first sent the alarm, UTC |
On a rule with Use AI, the written overview says plainly that the alarm has not been picked up, and that acknowledging it on the dashboard stops the escalation.
Testing
Send a test on an escalation rule sends a sample message to its recipients straight away, which proves the channel and the addresses. To see a whole chain climb, make a notification rule bound to a single sensor that nothing else covers, with the shortest After, and leave the alarm unacknowledged. Do not test it on a sensor your real rules cover: every rule covering it will send, to everyone on them.