How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
Rules turn readings into a state. Each rule tests one slot and says what the sensor should be when the test is true: WARN or CRIT. A sensor matching no rule is OK. Rules only test slots, so anything you want to judge must be in a slot first; see Extraction.
How rules are judged
- Every rule is tested on every reading, and the worst state reached wins. CRIT beats WARN beats OK. The order the rules are listed in makes no difference, so a WARN rule can never hide a CRIT rule below it, and you can add rules in whatever order reads best.
- Every rule that reached the winning state is named in the alarm, with the value it saw: Guest 133 is 0, not equal to (1,2,3); Guest 105 is 0, not equal to (1,2,3). The name comes from the slot's Reading column, so name your slots.
- Rules act on a single reading. The failure threshold on the Polling panel counts failed readings (no answer, timeouts, refusals) and does not apply to rules. One reading over the line is enough to change the state, and the next reading under it clears it again.
- Rules only judge readings that arrived. A failed poll has no slots to test, and is dealt with by the failure threshold instead.
- A slot that could not be read is not tested at all, so its rules are skipped rather than failed. See the guard rule below.
- With several probe groups, each group's readings are judged on their own, and the sensor's Rollup settings decide what the sensor is when the groups disagree.
The operators
| Is | Use it for |
|---|---|
| equal to, not equal to | A number or a piece of text. Compared as numbers when both sides are numbers, so 200 and 200.0 are equal; otherwise as text, exactly, including capital letters. Both also take a bracketed list: see below. |
| greater than, at least, less than, at most | Numbers only. Never true for text. |
| between, outside | Numbers only, with the lower bound in Value and the upper in And. Both bounds count as between. |
| containing, not containing | Text that includes the value anywhere. Capital letters count: Error does not contain error. |
| matching, not matching | A regular expression, delimiters and all: /^5\d\d/. Add i after the closing delimiter to ignore case: /timeout/i. |
A rule with no value is dropped when you save, as is between or outside without an upper bound. Values are plain text: accented letters and symbols such as £ are removed, so test those with a pattern instead.
Lists: one of these, none of these
equal to and not equal to take a bracketed, comma separated list as well as a single value. equal to (2,4) means "is one of these"; not equal to (1,3) means "is none of these". Items can be numbers or text: (running,paused).
Lists are how you judge a state code, where several values are fine and the rest are not. A network interface is healthy when its status is up (1) or testing (3), and faulty in any of the other five. not equal to (1,3) gives CRIT says that in one rule. Two separate rules, not equal to 1 and not equal to 3, would each go off on the other's healthy value. Name the healthy values or the faulty ones, whichever is the shorter list.
To test text that genuinely starts and ends with brackets, use matching.
Stacking rules on one slot
Several rules can test the same slot, which is how a reading gets both a warning and a critical level:
| Slot | Is | Value | Then |
|---|---|---|---|
| SV03, Disk used | greater than | 80 | WARN |
| SV03, Disk used | greater than | 90 | CRIT |
At 85 the first matches and the sensor is WARN. At 95 both match and CRIT wins. There is no need to write the WARN rule as between 80 and 90.
The same goes for a range with a safe middle. outside 10 and 35 gives WARN with outside 5 and 40 gives CRIT on a temperature warns early and goes critical at the extremes, in either direction.
The guard rule
Most sensor types fill some slots on every reading whatever happens, and a rule on a slot that could not be read is skipped. So an HTTPS sensor whose API token has expired gets a 401, fills its status code, finds nothing at your JSON paths, skips every rule on them and reads OK. It is blind, and says it is fine.
Every sensor that reads values out of a response needs a rule that fails when the response itself is wrong:
- HTTP(S): SV01, Status code, not equal to 200, gives CRIT. Use a list,
(200,204), where more than one code is normal. - An API that answers errors with a web page: SD02, Content type, not containing application/json, gives CRIT.
- A response with a field that says all is well: a rule that goes CRIT when that field is anything else.
Examples
| To catch | Rule |
|---|---|
| A web page that is not OK | SV01 not equal to 200, CRIT |
| A page that has come back empty | SV02 less than 500, CRIT |
| A slow web page | A DIRECT row into SV03, then SV03 greater than 2000, WARN |
| A maintenance page served with a 200 | A REGEX row /maintenance/i into SD03, then SD03 matching /maintenance/i, WARN. The pattern searches the whole page; a WHOLE row would keep only its first 255 characters |
| An SMTP server answering with the wrong name | SD01 not containing mail.example.co.uk, WARN |
| A Proxmox guest not running | The guest's slot not equal to (1,2,3), CRIT: running, in a backup or migrating are all fine |
| Any of a group of guests down | With uf_pveguests and its groups argument, the group's count greater than 0, CRIT. One rule for the whole group, with a text row reading the group's _down value to say who |
The last two show there is more than one way to cover the same thing. One rule per guest gives each its own line on a chart and its own name in the alarm; one rule per group keeps the rule list short. Both are right: pick the one that reads best for whoever answers the alarm.