How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
Every reading a sensor takes has thirty slots to fill: twenty numeric ones, SV01 to SV20, and ten text ones, SD01 to SD10. Rules test slots, charts draw them and alarms quote them. Extraction is where you say what goes in each one: which part of the response to read, what to call it, and what unit it is in.
The order things happen in
- The user function, if one is chosen, converts the response into something easier to read. What it returns replaces the response body for everything after it.
- The sensor type's own slots are filled. Most types fill a few by default (see below).
- Your rows are read. A row on a slot the type already fills replaces the type's value.
- The multiplier is applied to each numeric slot, and the reading is stored.
Slots the type fills for you
These are filled on every reading without you setting anything. On the Extraction panel they show their name under the slot and From the type in the Method list. Choosing a method on one of these rows replaces what the type puts there, so to add your own readings start from the first free slot instead.
| Type | Filled by default | First free |
|---|---|---|
| HTTP(S) | SV01 status code, SV02 response size, SD01 status text, SD02 content type | SV03, SD03 |
| SNMP | SV01 the value, SD01 the raw value, SD02 its declared type | SV02, SD03 |
| DNS | SV01 response code, SV02 answer count, SD01 response, SD02 answer | SV03, SD03 |
| SIP | SV01 status code, SD01 status line, SD02 user agent, SD03 allowed methods | SV02, SD04 |
| MySQL | SV01 the value, SV02 rows, SD01 the value as text, SD02 the column | SV03, SD03 |
| FTP | SV01 login time, SV02 transfer time, SV03 bytes read, SV04 file size, SV05 file age, SD01 modified, SD02 content | SV06, SD03 |
| MongoDB | SV01 ok, SD01 the reply | SV02, SD02 |
| SV01 return time, SD01 to SD04 outcome, detail, mail server and TLS | SV02, SD05 | |
| TCP, SMTP, IMAP4 | SD01 the banner | SV01, SD02 |
| RDP | SD01 the protocol | SV01, SD02 |
| Ping | Nothing | SV01, SD01 |
Numeric and text slots
- Numeric slots hold a number to six decimal places, and are what rules compare and charts draw. Anything that is not a number is refused and the slot is left empty. That includes JSON true and false: read those into a text slot, where they arrive as 1 and nothing, or have a user function turn them into 1 and 0.
- Text slots hold up to 255 characters and are shown on the reading. Anything longer is cut short. Rules can still test them with equals, contains and matches.
The columns
| Column | What it does |
|---|---|
| Reading | The name shown on charts, in alarms and in the rule list. Say what it is: Queue length, not SV03. |
| Unit | Numeric slots only. Shown beside the value and on the chart's scale. |
| Chart | Numeric slots only. See Chart, below. |
| Multiplier | Numeric slots only. Applied before the reading is stored, so a value that arrives in bytes can be stored in megabytes at 0.000001, or one in tenths of a degree at 0.1. Rule thresholds are then written in the stored units. Changing it does not alter readings already taken. On an SNMP sensor, a row for SV01 replaces the multiplier on the SNMP panel, even at 1, so set it again here. |
| Method | How the value is found. See the methods below. |
| Labels | What each raw value means, as JSON: {"1":"up","2":"down"}. The reading is still stored as the number and rules still test the number, so this changes only what is shown. SNMP needs none, since the MIB carries the meanings. |
| Expression | What the method looks for: a path, a pattern or a header name. Methods that read the response as a whole ignore it. |
The methods
JSON: one value from a JSON response
The expression is a path of names separated by dots, with a number to pick an item from a list, counting from 0. Given this response:
{"data": {"cpu": 0.12, "node": "pve1"},
"PowerControl": [{"PowerConsumedWatts": 212}],
"healthy": true}
| Expression | Gives |
|---|---|
data.cpu | 0.12 |
data.node | pve1, so into a text slot |
PowerControl.0.PowerConsumedWatts | 212 |
healthy | true, which only a text slot will take |
data | The whole object as JSON text, cut to 255 characters in a text slot |
This is not JSONPath. There is no $, no [*], no wildcards and no filters, and a name that itself contains a dot cannot be reached. A path always means one fixed place in the response. The usual problem is a list where the item you want can be anywhere, such as a guest or an interface found by name, not by position, since positions change. There are two ways round it:
- Ask the API for the one item. Many take a filter in the URL, for example RouterOS's
/rest/interface?name=ether1, and the item you want is then always at position 0. - Use a user function (below) that finds it and hands back a flat answer.
JSONCOUNT: how many items
The number of items in the list, or keys in the object, at the path. data on a Proxmox resource list is the number of guests; an empty expression counts the top level. It is the one thing a path cannot say, because a path to a list gives you the list.
XPATH and XPATHCOUNT: XML and SOAP
For XML, not JSON. The expression is an XPath 1.0 expression, and the text of the first node it finds is the value, stored as a number where it is one. XPATHCOUNT gives the number of nodes found instead.
Namespaced documents, which SOAP always is, need the prefixes declared. Give the expression as JSON with a prefix map:
{"xpath": "//t:Temperature", "ns": {"t": "urn:example:thermal"}}
Or match on the local name and skip namespaces altogether: //*[local-name()='Temperature'].
REGEX: a pattern in text
For HTML pages, banners and anything else that is plain text. The expression is a full PCRE pattern, delimiters and all, like /Uptime:\s*(\d+)/. The value is the first bracketed group, or the whole match if there is none. Only the first match counts.
- Flags go after the closing delimiter:
iignores case,slets.cross line breaks,mmakes^and$work per line. - If the pattern needs a
/, use another delimiter:#href="(/status[^"]*)"#. - A number captured into a numeric slot is stored as a number, so
/(\d+) users online/into SV03 can be charted and ruled on.
HEADER: a header or named value
A response header by name, such as Content-Type or X-Queue-Depth, matched without regard to case. Non-HTTP types report some of what they find this way too, which is how DNS fills its response code and FTP its file age.
STATUS, STATUSTEXT, SIZE
The status code (200), the status line (200 OK) and the number of bytes of body read. No expression.
WHOLE: the entire body
The body as it is: a number if it is nothing but a number, such as the result of a SQL query that returns one value, otherwise text, cut to 255 characters.
DIRECT: the response time
How long the response took, in milliseconds. Putting it in a slot is how you write rules on it, such as warning when a web page slows down. On a TCP sensor the connect is the whole check, so its response time is always zero.
User functions
Some responses cannot be read with a path: a list you need to search, values you need to add up, a guest you need to find by its ID wherever it sits. A user function turns a response like that into a small, flat answer that an ordinary JSON row can read.
- Choose it at the top of the Extraction panel. Each one shows its own description once chosen, which says what it returns and what arguments it takes.
- Arguments are JSON and tell the function what to pick. They are how one function serves many sensors.
- What it returns replaces the body for JSON, JSONCOUNT, XPATH, REGEX and WHOLE rows. STATUS, STATUSTEXT, SIZE and HEADER still read the real response, so the type's own status code slot still means what it did.
- It costs 4 extra credits a reading, which is why a path is always the first thing to try.
- User functions are written by GEN, not in the portal. If you need one that is not in the list, raise a ticket saying what the endpoint returns and what you want from it.
For example, uf_pveguests on a Proxmox /api2/json/cluster/resources?type=vm sensor, with the arguments {"expect":[114,133]}, returns {"114":1,"133":0,"notrunning":"133 proxy stopped on pve2"}. A JSON row with the expression 114 then reads guest 114's state wherever it sits in the cluster, and a text row reading notrunning gives the alarm something useful to say.
When a slot cannot be read
If a path does not resolve, a pattern does not match or a value is not a number, the slot is left empty rather than set to zero, and any rule on that slot is skipped. The sensor only goes UNKNOWN when nothing at all could be read.
This matters because most types always read something. An HTTPS sensor whose token has expired still gets a status code, 401, so the type's slots are filled. Your JSON rows find nothing to read, their rules are skipped, and the sensor reads OK while seeing nothing. Always add a guard rule: on an HTTPS sensor, SV01 not equal to 200 gives CRIT. The same happens when a user function fails, since the body is then left as it arrived.
Checking your work
Tick Store raw response on the Polling panel and each reading keeps the body your rows were tested against: the converted one where a user function is in use. Compare your paths with what is actually there, then untick it, as it costs 2 extra credits a reading.
Chart
Any numeric reading can be put on one of five charts, numbered 1 to 5, or left on None. The setting appears in two places: the Chart column of each numeric slot in Extraction, and Chart, response time on the Polling panel for the sensor's own timing (connect time on a TCP sensor, and the value read on an SNMP sensor). Text slots cannot be charted.
The number belongs to the device, not the sensor. Every reading on a device given Chart 2 is drawn on the same chart, whichever sensor it comes from. Give each of a server's volumes Chart 2 and they plot side by side; give CPU and memory Chart 1 and they plot together on another.
- On the dashboard, the chart picker for a device offers Chart 1, Chart 2 and so on, each drawing exactly the readings given that number. Series picked by hand lets you tick your own instead. Your browser remembers which you last chose for each object.
- In alarms, where Charts in notifications is ticked on the Polling panel, the message carries every numbered chart the alarming sensor's readings are on, drawn across the whole device. A volume filling up arrives with the other volumes beside it. A sensor with nothing numbered sends up to four of its own readings instead. A recovery carries no charts.
Two units to a chart. A chart has a scale on each side, one for each of the first two units it holds. A reading in a third unit is left off and the legend says so. The unit is matched exactly as typed, so write it the same way every time: ms and Ms count as two units.