How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
Every object in the tree has a state, and each one is worked out from what sits directly beneath it. Rollup is where you say how. The same three settings appear at every level; only what is being counted changes:
| On a | Rollup counts |
|---|---|
| Sensor | Its probe groups, each judging the sensor's readings on its own |
| Device | Its sensors |
| Group | The devices and groups in it |
| Site | Its groups |
How the state is worked out
- Only those with an opinion are counted. Anything OK, WARN or CRIT is counted. Anything UNKNOWN, STALE or HELD, switched off, suspended, still a draft or outside its schedule is left out altogether: it is neither healthy nor failing, so it cannot drag the state either way. Where nothing is left to count and something beneath is HELD, cut off by something it depends on, the object is HELD too; see Dependencies. A sensor with Counts towards its parent's state cleared is left out of its device's count too.
- WARN and CRIT both count as failing.
- How many are failing puts the object in one of four buckets:
- none failing: always OK
- ANY: exactly one failing, and at least one other not
- SOME: more than one failing, but not all
- ALL: every one failing. With only one counted, one failing is all of them
- The bucket gives the state you chose for it, OK, WARN or CRIT.
- The state is never worse than the children. If the bucket says CRIT but none of the failing children is CRIT, the object is WARN. A device whose only problem is a warning never goes critical.
The settings
| Setting | What it does |
|---|---|
| ANY, SOME, ALL | The state for each bucket. |
| Minimum reporting | Below this many with an opinion, the state is UNKNOWN rather than a verdict reached from one straggler. On a sensor in three probe groups, 2 means one group on its own cannot decide. |
| Counts towards its parent's state | Clear it to watch something without letting it condemn what it sits on. A certificate expiry warning is worth seeing, and is not the device being down. |
| Counts towards availability reporting | Clear it to keep the object out of availability figures, for something whose downtime is expected or not your concern. |
The defaults
| Object | ANY | SOME | ALL |
|---|---|---|---|
| Sensor, group, site | OK | WARN | CRIT |
| Device | WARN | CRIT | CRIT |
A device with one failing sensor needs a look, and with more than one is in trouble, so a device defaults to WARN on ANY and CRIT on SOME and ALL. Keep each setting at least as bad as the one before it: a device set to CRIT on ANY but WARN on SOME would look better when a second sensor fails.
Sensors and probe groups
A sensor in one probe group is simple: when that group's readings fail, one of one is failing, which is ALL, so CRIT.
With more than one group, the defaults ask for agreement. In two groups, one failing is ANY, which is OK: the device is still reachable from the other, and the fault is most likely on the path from the first. Both failing is ALL and CRIT. Change the sensor's settings to suit what you are watching:
| You want to know | ANY | SOME | ALL |
|---|---|---|---|
| Is it down for everyone (the default) | OK | WARN | CRIT |
| As soon as anyone loses it | WARN | CRIT | CRIT |
| Can every location reach it, such as a public web site | CRIT | CRIT | CRIT |
SOME can only happen with three or more groups; with two, it is never reached.