How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
How often this sensor is polled, how long each poll may take, when a failure counts, where it is polled from, and what that costs.
| Setting | What it does |
|---|---|
| Interval | Seconds between polls, 10 at the least. 60 suits most things; go shorter only where a minute matters, such as a core switch. Every probe group polls on its own schedule, so halving the interval or adding a group doubles the cost. |
| Connect timeout | How long to wait to reach the service at all: the connection and, where there is one, the TLS handshake. 5 seconds is generous on a local network. |
| Response timeout | How long to wait for the answer once connected. Raise it for a slow query or a heavy page, not for a device that is simply unreachable. |
| Failures before a probe calls it down | How many failed polls in a row a probe group needs before it calls the sensor down. 3 is the default and absorbs a single lost packet or a dropped connection. One good reading resets the count. |
| Retention, days | How long this sensor's readings are kept. Blank means the service default of 90 days. Raise it where a long history is wanted for reports, lower it for a chatty sensor whose detail is worth nothing after a fortnight. |
| Target override | Leave empty and the sensor uses the device's own address, so moving the device moves the sensor with it. Set it only to point somewhere else. |
| Probe groups | Which of GEN's locations poll this sensor. At least one is needed. Several give a second opinion, which is what tells a device being down apart from one path to it being down. |
Connect plus response timeout must be less than the interval, and the form refuses a save that breaks this. Otherwise a slow poll is still running when the next is due and polls pile up. On a 10 second interval, 1 and 3 leaves room; 5 and 10 does not.
Two types do not use both timeouts. A ping makes no connection, so only the response timeout applies. A TCP sensor is finished once it has connected, so the connect timeout is the one that matters.
Failures, and what they are not
The failure threshold counts failed polls: no answer, a timeout, a refused connection. It does not apply to rules. A reading that arrives and breaks a rule changes the state at once, on that one reading. Raising the threshold makes a flapping service quieter; it does nothing about a value bouncing over a threshold, which belongs in the rule.
Several probe groups
Each group judges the sensor on its own readings and keeps its own failure count. What the sensor becomes when the groups disagree is set on its Rollup panel, where ANY, SOME and ALL each map to a state. The default asks for agreement: with two groups, both must lose the device before the sensor goes critical.
An email delivery sensor is run centrally by GEN and has no probe group to choose.
Store raw response
Keeps the whole response with every reading, or the converted one where a user function is in use, so you can see exactly what the extraction rows and rules were tested against. Invaluable while setting a sensor up, and it should be turned off afterwards:
- It adds 2 credits to every reading, on top of the sensor type's own cost and the 4 credits a user function adds.
- The response is kept for as long as the reading is, and anyone who can see the sensor can read it. Leave it off where responses carry personal data or anything else that should not be sitting in a monitoring history.
Charts in notifications
On by default. When this sensor raises an alarm, its charts are drawn into the message on channels that carry pictures: email, tickets and Matrix. Clear it where a chart says nothing, such as a list of guests that are each either up or down.
What it costs
The foot of the panel shows the credits for one reading, and an estimate for a month at the settings as last saved. The sum is simply the cost of a reading times the number of readings, times the number of probe groups. An interval of 10 seconds is six times the cost of one minute, from every group. The charge is for readings actually accepted.
Chart
Any numeric reading can be put on one of five charts, numbered 1 to 5, or left on None. The setting appears in two places: the Chart column of each numeric slot in Extraction, and Chart, response time on the Polling panel for the sensor's own timing (connect time on a TCP sensor, and the value read on an SNMP sensor). Text slots cannot be charted.
The number belongs to the device, not the sensor. Every reading on a device given Chart 2 is drawn on the same chart, whichever sensor it comes from. Give each of a server's volumes Chart 2 and they plot side by side; give CPU and memory Chart 1 and they plot together on another.
- On the dashboard, the chart picker for a device offers Chart 1, Chart 2 and so on, each drawing exactly the readings given that number. Series picked by hand lets you tick your own instead. Your browser remembers which you last chose for each object.
- In alarms, where Charts in notifications is ticked on the Polling panel, the message carries every numbered chart the alarming sensor's readings are on, drawn across the whole device. A volume filling up arrives with the other volumes beside it. A sensor with nothing numbered sends up to four of its own readings instead. A recovery carries no charts.
Two units to a chart. A chart has a scale on each side, one for each of the first two units it holds. A reading in a third unit is left off and the legend says so. The unit is matched exactly as typed, so write it the same way every time: ms and Ms count as two units.