How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message, or ask a question in your own words.
MikroTik: RouterOS routers and switches
MikroTik routers and CRS switches can be watched two ways, and most estates use both. SNMP works on RouterOS 6 and 7 and is the cheapest route for hardware health, interfaces, PoE, optics and LTE signal. The REST API, RouterOS 7 only, is the only way to see BGP and OSPF sessions and WireGuard handshakes, and gives power supply state as plain words. The MikroTik MIB is already loaded into Oversight, so its objects can be picked by name.
MikroTik's own binary API on ports 8728 and 8729 is not HTTP, and no Oversight sensor can use it. REST is the HTTP route.
Which route for what
| To watch | SNMP | REST (RouterOS 7) |
|---|---|---|
| A link going down, or flapping | Yes | Yes |
| Temperatures and voltages | Yes | Yes |
| Power supply failed | Yes, once proved on the device (see below) | Yes, as ok or fail |
| CPU load | Per core | Whole box |
| BGP and OSPF sessions | No: RouterOS has no routing MIBs | Yes |
| WireGuard and IPsec peers | IPsec count only | Yes, per peer |
| PoE ports, SFP optics, LTE signal | Yes | Not reliably |
| A reboot | Yes, from uptime | As text only |
A Cloud Hosted Router (CHR) has no temperature, fan, voltage or power supply readings at all; that is the hypervisor's job. Watch a CHR for links, CPU and routing only.
Setting up SNMP
On the router, add a read-only community that answers only the probe groups' addresses, then shut the factory public community, which answers anyone:
/snmp community add name=<long-random-string> addresses=<probe-address>/32 read-access=yes write-access=no /snmp set enabled=yes contact="NOC" location="Rack 3" /snmp community set [find name=public] disabled=yes
Add one addresses entry per probe group address, comma separated. If the router refuses to disable public, set its addresses to 127.0.0.1/32 instead. Never enable write access: SNMP write can reboot the router and run scripts.
MikroTik recommends SNMP v3, but Oversight's probes use v2c, so make up for it: the address restriction above, a firewall rule allowing UDP 161 in the input chain from the probes only, and polling over a management network or VPN where there is one.
In Oversight, save an SNMP credential with version 2c and the community, on the device or anywhere above it. See Credentials.
Setting up REST
REST rides on the router's HTTPS web service, so it needs a certificate, the service switched on and limited to the probes, and a user that can read but not change anything.
/certificate add name=oversight-https common-name=<router-name-or-address> days-valid=3650 /certificate sign oversight-https /ip service set www-ssl certificate=oversight-https disabled=no address=<probe-address>/32 tls-version=only-1.2 /user group add name=oversight-ro policy=read,api,rest-api /user add name=oversight group=oversight-ro password=<strong-password> address=<probe-address>/32
- Make a group of your own. The built-in
readgroup also carries reboot, sniff and sensitive rights.read,api,rest-apiis the set known to work; some versions manage with less, but that is the safe one. - Before RouterOS 7.12 the user's address limit is ignored for REST (CVE-2023-41570). The service's
addressand the firewall are what protect it there, so upgrade if you can. - Enabling
www-sslalso serves WebFig on that port. The address limit covers both. - A self-signed certificate means turning off Verify the certificate on the sensor. If the router has a public name,
/certificate/add-acmegets it a Let's Encrypt certificate and verification can stay on.
In Oversight, save an HTTP(S) credential with the username and password, and give every REST sensor these settings:
| Setting | Value |
|---|---|
| Scheme and port | https, 443 (or the port you gave www-ssl) |
| Method | GET |
| Headers | Authorization: Basic {{basicauth}} |
| Verify the certificate | Off for a self-signed certificate |
| Interval | 60 seconds or more. On small single-core routers REST polling is noticeably heavy. |
And give every REST sensor the guard rule SV01 not equal to 200 gives CRIT, so a refused login or a wrong path raises an alarm rather than reading OK with nothing in it.
Finding the numbers
SNMP reads one value per sensor, and anything that belongs to a port or a sensor chip needs its row number on the end of the OID.
- Ports are numbered by ifIndex. On the router,
/interface print oidlists every interface with its OIDs, row number included. PoE, SFP optics and LTE tables use the same number. Check again after adding interfaces or a major upgrade. - Health readings on RouterOS 7 (and late 6) come from one table,
mtxrGaugeTable, and its row number is a fixed code for the kind of reading, not a position. The same code means the same thing on every model that has that reading.
| Row | Reading | Unit and multiplier |
|---|---|---|
| 17 | cpu-temperature | Whole degrees C, multiplier 1 |
| 14 | temperature (board) | Whole degrees C |
| 7101, 7102 | board-temperature1, 2 | Whole degrees C |
| 50, 51 | sfp-temperature, switch-temperature | Whole degrees C |
| 13 | voltage | Tenths of a volt, multiplier 0.1 |
| 16 | poe-out-consumption | Tenths of a watt, multiplier 0.1 |
| 54 | fan-state | A status code |
| 7001, 7002 | fan1-speed, fan2-speed | RPM |
| 7401, 7402 | psu1-state, psu2-state | A status code |
A model only has the rows it has sensors for. To see which, walk 1.3.6.1.4.1.14988.1.1.3.100.1.2 once from any workstation: it lists the name of each row. In Oversight, pick MIKROTIK-MIB and mtxrGaugeValue from the MIB list, then put the row number on the end of the OID.
SNMP sensors worth having
In the table, MT stands for 1.3.6.1.4.1.14988.1.1, and <if> for the port's ifIndex.
| What | OID | Settings | Rules |
|---|---|---|---|
| Link up | 1.3.6.1.2.1.2.2.1.8.<if> (ifOperStatus) | GAUGE, as read | SV01 not equal to 1 gives CRIT. Down ports read 2 on some builds and 6 on others, so test for not 1. |
| Link flaps | MT.14.1.1.90.<if> (mtxrInterfaceStatsLinkDowns) | COUNTER32, change per interval | At least 1 gives WARN, at least 3 gives CRIT |
| CPU temperature | MT.3.100.1.3.17 | GAUGE, unit C | Greater than 80 gives WARN, greater than 95 gives CRIT |
| Board temperature | MT.3.100.1.3.14 | GAUGE, unit C | Greater than 65 gives WARN, greater than 75 gives CRIT |
| Input voltage | MT.3.100.1.3.13 | GAUGE, multiplier 0.1, unit V | Outside 10% either side of the supply's nominal voltage gives WARN. An unused input reads 0: leave it alone. |
| Power supply | MT.3.100.1.3.7401 and .7402 | GAUGE | Prove the codes first (below), or use REST. |
| Rebooted | 1.3.6.1.2.1.1.3.0 (sysUpTime) | GAUGE, multiplier 0.01, unit s | Less than 600 gives WARN: it restarted in the last ten minutes |
| CPU load | 1.3.6.1.2.1.25.3.3.1.2.<n>, n the core, from 1 | GAUGE, unit % | Greater than 95 gives WARN. One sensor per core; REST gives the whole box in one. |
| PoE port | MT.15.1.1.3.<if> (mtxrPOEStatus) | GAUGE; labels come from the MIB | On a port feeding a known device: not equal to 3 (poweredOn) gives CRIT. Elsewhere: equal to (4,9,10,13) gives CRIT. |
| SFP signal received | MT.19.1.1.10.<if> (mtxrOpticalRxPower) | GAUGE, multiplier 0.001, unit dBm | Note the reading when the link is commissioned. 3 dB below it gives WARN, 6 dB below gives CRIT. |
| SFP lost signal | MT.19.1.1.3.<if> (mtxrOpticalRxLoss) | GAUGE | Equal to 1 gives CRIT |
| LTE signal (RSRP) | MT.16.1.1.4.<if of lte1> | GAUGE, unit dBm | Less than -110 gives WARN, less than -120 gives CRIT |
| IPsec tunnels up | MT.20.1.0 (mtxrIkeSACount) | GAUGE | Less than the number of tunnels you expect gives CRIT |
The temperature and signal thresholds are sensible starting points, not MikroTik's figures; MikroTik only publishes its overheating cut-out (105 C). Adjust them once you know what the kit normally reads.
Do not alarm on fan speed. Fans stop when the unit is cool, and a healthy router can read 0 RPM for days. Watch temperature and fan-state (row 54, 0 when healthy) instead.
Prove the power supply codes on the device
The health table gives power supply state as a number, and the published mapping (0 ok, 1 failed) has been contradicted by at least one real device. Before relying on rows 7401 and 7402, compare what the sensor reads with /system health print, which says ok or fail in words, and if you can, pull one supply while commissioning and watch the number change. Or use the REST sensor below, which reads the words and avoids the question.
RouterOS 6
Older routers have health as separate objects rather than the table: MT.3.10.0 (temperature), MT.3.11.0 (processor temperature) and MT.3.8.0 (voltage), all in tenths, so multiplier 0.1. Power supply state is MT.3.15.0 and MT.3.16.0, where 1 means ok: not equal to 1 gives CRIT. Some of these survive on RouterOS 7 and some do not, which is why the table is the one to use there.
REST sensors worth having
Every value in a RouterOS REST reply is a string, even numbers. A plain number such as "45" still goes into a numeric slot, but "true", "ok", "established" and durations such as "2d20h12m20s" are text. The neat way round that for anything that is simply up or down is to ask for only the healthy items and count them with JSONCOUNT and the expression ., meaning the whole reply: 1 means up, 0 means down or gone.
A filtered request (?name=...) returns a list, so paths start 0.; a request by name (/rest/interface/ether1) returns the item itself.
| What | Path | Extraction | Rules |
|---|---|---|---|
| Link up | /rest/interface?name=ether1&running=true | SV03, JSONCOUNT, expression . | SV03 equal to 0 gives CRIT |
| CPU load | /rest/system/resource?.proplist=cpu-load | SV03, JSON, cpu-load, unit % | Greater than 80 gives WARN, greater than 95 gives CRIT |
| Power supply | /rest/system/health?name=psu1-state | SD03, JSON, 0.value | SD03 not equal to ok gives CRIT. A second sensor for psu2-state. |
| CPU temperature | /rest/system/health?name=cpu-temperature | SV03, JSON, 0.value, unit C | Greater than 80 gives WARN, greater than 95 gives CRIT |
| BGP sessions up | /rest/routing/bgp/session?established=true | SV03, JSONCOUNT, expression . | Less than the number of sessions you expect gives CRIT |
| One BGP session | /rest/routing/bgp/session?name=<session>&established=true | SV03, JSONCOUNT, expression . | SV03 equal to 0 gives CRIT |
| OSPF neighbours | /rest/routing/ospf/neighbor?state=Full | SV03, JSONCOUNT, expression . | Less than the number expected gives CRIT |
| IPsec peer | /rest/ip/ipsec/active-peers?remote-address=<peer>&state=established | SV03, JSONCOUNT, expression . | SV03 equal to 0 gives CRIT |
| WireGuard peer | /rest/interface/wireguard/peers?name=<peer> | SD03, REGEX, /"last-handshake":"([^"]*)"/ | SD03 matching /[hdw]/ gives CRIT: no handshake for an hour or more |
| LTE link up | /rest/interface/lte?name=lte1&running=true | SV03, JSONCOUNT, expression . | SV03 equal to 0 gives CRIT |
Plus SV01 not equal to 200 gives CRIT on every one. A few notes on the table:
- BGP session names are not always what you typed: RouterOS often adds
-1to the connection's name. Read the real name from/routing/bgp/session printfirst. Session entries linger for a while after a session drops, which is why the filter asks forestablished=truerather than counting every row. - BGP fields with a dot in their name, such as
remote.address, cannot be read by an Oversight JSON path, which uses dots to separate names. Filter and count instead, as above. - WireGuard peers can be filtered by
namefrom RouterOS 7.15; before that, filter oncomment=orpublic-key=. A peer that has never shaken hands has nolast-handshakeat all, so the slot is empty and the sensor cannot tell: pair it with a ping across the tunnel. With a keepalive set, a healthy peer's handshake is never more than two or three minutes old. /rest/system/resourceshould return a single item, but MikroTik's own example shows a list. Ifcpu-loadreads nothing, tick Store raw response, look at what came back, and use0.cpu-loadif it is a list.
A starting set
Not every reading needs a sensor. For a RouterOS 7 router doing routing, these catch what actually goes wrong:
- Link up on each uplink (SNMP), and link flaps on the same ports.
- CPU temperature (SNMP row 17).
- Power supplies, on models with two (REST,
okorfail). - CPU load for the whole box (REST).
- BGP sessions established, or OSPF neighbours Full (REST).
- WireGuard or IPsec peers, where the site depends on them.
- Rebooted (SNMP uptime).
For a CRS switch, drop the routing sensors and add PoE state on ports feeding cameras or access points, and received signal and lost signal on fibre uplinks. For an LTE router, add the lte1 link and its RSRP. On RouterOS 6, use SNMP throughout. At the default 60 seconds from one probe group, an SNMP sensor costs about £1.73 a month and a REST sensor about £3.46.
Things that catch people out
- Tenths and wholes. In the health table, temperatures are whole degrees but volts, amps and watts are tenths. The old RouterOS 6 objects are tenths throughout. A voltage reading of 237 is 23.7 V.
- An SFP reading of exactly 0 dBm usually means the module reports nothing (a direct attach cable or copper module), not a perfect signal. RouterOS 7.24 also lists empty SFP cages in the optics table;
MT.19.1.1.13.<if>is 1 where a module is actually fitted. - PoE state 5 (shortCircuit) has been seen on idle, unused ports of an RB5009. Only alarm on ports with something plugged in.
- Health readings are approximate, in MikroTik's own words, and meant to warn of trouble coming. Some older PoE-powered CRS models show a fixed 16 V whatever the input.
- Polling cost. Health readings are expensive for some models to produce: MikroTik warns that the CCR2004-16G-2S+PC is one. Keep health sensors at 60 seconds or more. Oversight asks for one value at a time and never walks the table, which is the gentle way.
- Permission changes are slow to take. A REST user's rights can take several minutes to change after you alter the user or group. If a login is refused straight after setting it up, wait and try again before changing anything.
MIBs
MIKROTIK-MIB, IF-MIB, HOST-RESOURCES-MIB and SNMPv2-MIB are all loaded already, so everything above can be picked from the MIB list or typed as a number. RouterOS does not implement the BGP or OSPF MIBs, so there is nothing more to load for routing.