How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message, or ask a question in your own words.
NVIDIA (Mellanox): Onyx, Cumulus Linux and SONiC switches
NVIDIA's Spectrum Ethernet switches, formerly Mellanox, can run any of three operating systems on the same hardware: Onyx (also called MLNX-OS), Cumulus Linux or SONiC. How to watch a switch depends on which it runs, and the model number does not tell you: an SN2700 or SN4600C can have any of them. So find out first.
Which one is it?
Read the switch's description by SNMP, 1.3.6.1.2.1.1.1.0, or log in and look at the prompt:
| Operating system | Description reads like | Prompt |
|---|---|---|
| Onyx / MLNX-OS | Onyx,MSN2410,SWv3.6.6102 or Mellanox MSN2700,MLNX-OS,SWv3.6.4112 | switch (config) # |
| Cumulus Linux | Cumulus-linux 5.9.2 (Linux Kernel ...) | cumulus@switch:~$ |
| SONiC | SONiC Software Version: SONiC... | admin@sonic:~$ |
In short: on Cumulus Linux 5, use its REST API, which reports everything including MLAG and an overall health verdict. On Onyx, use SNMP for hardware, ports and BGP, and its JSON API for MLAG only. Neither needs any NVIDIA MIB loaded.
Cumulus Linux: the NVUE API
From Cumulus Linux 5.6 the NVUE REST API is on by default, on port 8765, and takes a username and password in each request. It returns the switch's live state as JSON.
Setting it up
nv set system api state enabled nv set system api port 8765 nv set system api listening-address <eth0-address> nv set system aaa user oversight role nvue-monitor nv set system aaa user oversight password nv config apply
nvue-monitor is Cumulus's read-only role: it can show, not change. Two things to know:
- The API refuses everyone with 403 until the
cumulususer's password has been changed from its default, even though Oversight uses its own account. - On Cumulus 5.5 and earlier the API is off and switching it on means editing the web server by hand. Treat those switches as SNMP only (below).
In Oversight, save an HTTP(S) credential with the username and password, and give each sensor https, port 8765, GET, the header Authorization: Basic {{basicauth}}, and Verify the certificate off, since the switch makes its own certificate at first boot. Paths start /nvue_v1. Add SV01 not equal to 200 gives CRIT to each.
Never add ?rev=applied to a path: that returns the configuration, not the live state, so a failed power supply or a downed port would look exactly as configured.
Sensors worth having
| What | Path | Extraction | Rules |
|---|---|---|---|
| Overall health | /nvue_v1/system/health | SD03, JSON, status | SD03 not equal to OK gives CRIT. Cumulus 5.10 and later. |
| Any power supply not ok | /nvue_v1/platform/environment/psu | SD03, REGEX, below | SD03 matching /\w/ gives CRIT |
| Any fan not ok | /nvue_v1/platform/environment/fan | SD03, REGEX, below | SD03 matching /\w/ gives CRIT |
| A port up | /nvue_v1/interface/swp51 | SD03, JSON, link.oper-status | SD03 not equal to up gives CRIT. Cumulus 5.9 and later. |
| A BGP peer | /nvue_v1/vrf/default/router/bgp/neighbor/10.0.0.1, or /neighbor/swp51 for an unnumbered peer | SD03, JSON, state | SD03 not equal to established gives CRIT |
| MLAG peer | /nvue_v1/mlag | SD03, JSON, peer-alive; SD04, JSON, backup-active | SD03 not equal to True gives CRIT; SD04 not equal to True gives WARN |
| Disk filling | /nvue_v1/system/disk/usage | SV03, REGEX, /"\/":\s*\{[^}]*"used-percent":\s*"(\d+)%"/, unit % | At least 90 gives WARN, at least 95 gives CRIT. Cumulus 5.12 and later. |
| Restarted | /nvue_v1/system | SV03, JSON, uptime, unit s | Less than 600 gives WARN |
The power supply and fan sensors use one REGEX that finds the first item whose state is anything but ok or absent, and gives nothing at all when every item is healthy:
/^(?:(?=[\s\S]*?"state":\s*"(?!ok"|absent")(\w+)"))?/
That way one sensor covers every fan, without knowing how the model names them. absent is let through on purpose: a single-supply switch reports its empty bay as absent, and would otherwise alarm for ever. Where both supplies must be fitted, remove |absent" from the expression.
Some notes:
- Case matters. Rules compare text exactly. NVUE uses lower case (
ok,up,established), but MLAG saysTrueand the health checkOK. If a rule fires on a healthy switch, tick Store raw response and compare the case. - Overall health covers the switching chip, the hardware and the system's own processes in one verdict, which makes it the best single sensor. Where its healthy word is not exactly
OKon your release, the stored response will show it. - Temperature.
/nvue_v1/platform/environment/temperature(5.9 and later;/sensorbefore that) lists each sensor withcurrent,maxandcrit, all in degrees C. The switch's own health verdict already watches these, so a sensor of your own is only needed for an earlier warning: read the switching chip'scurrent, and set WARN at its ownmax. Do not use one temperature for every sensor: a healthy SN4600's switching chip runs at about 71 C, while its CPU cores sit in the 40s. - A known fault on Cumulus 5.11.3 to 5.17.0 makes the MLAG JSON come back as plain text. If it affects the API too, the MLAG sensor finds nothing to read and cannot tell; it does not raise a false alarm. Upgrading clears it.
Cumulus Linux: SNMP
For switches on 5.5 or earlier, or where the API is not wanted. The command changed at 5.11:
# Cumulus Linux 5.11 and later nv set system snmp-server state enabled nv set system snmp-server listening-address <eth0-address> vrf mgmt nv set system snmp-server readonly-community <community> access <probe-address>/32 nv config apply # Cumulus Linux 5.3 to 5.10 nv set service snmp-server enable on nv set service snmp-server listening-address <eth0-address> vrf mgmt nv set service snmp-server readonly-community <community> access <probe-address>/32 nv config apply
On 5.0 to 5.2 there are no NVUE commands for SNMP: add agentAddress <eth0-address>@mgmt and rocommunity <community> <probe-address>/32 to /etc/snmp/snmpd.conf and restart snmpd. On 5.3 and later, do not also edit that file: NVIDIA says not to mix the two.
- Without a listening address, SNMP answers only the switch itself.
- The management VRF is on by default, and SNMP does not cross VRFs, which is why the listening address says
vrf mgmt. - Switching SNMP off later restarts the routing service, which may interrupt traffic. Worth knowing before trying it during commissioning.
Find the part and port numbers once:
snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.47.1.1.1.1.7 snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.31.1.1.1.1
Cumulus numbers parts in ranges: temperature sensors from 100000001, fans from 100011001, power supplies from 110000001. The first walk shows which exist.
| What | OID | Settings | Rules |
|---|---|---|---|
| Power supply | 1.3.6.1.2.1.99.1.1.1.5.<supply>, such as ...5.110000001 | GAUGE; 1 ok | Not equal to 1 gives CRIT |
| Fan | 1.3.6.1.2.1.99.1.1.1.5.<fan> | GAUGE; 1 ok | Not equal to 1 gives CRIT |
| Temperature sensor failed | 1.3.6.1.2.1.99.1.1.1.5.<sensor> | GAUGE; 1 ok | Not equal to 1 gives CRIT |
| Temperature | 1.3.6.1.2.1.99.1.1.1.4.<sensor> | GAUGE, multiplier 0.1, unit C | Per sensor, at the limit the API or nv show platform environment temperature gives for it |
| Port up | 1.3.6.1.2.1.2.2.1.8.<port> | GAUGE | Not equal to 1 gives CRIT |
| BGP peer, numbered, default VRF | 1.3.6.1.2.1.15.3.1.2.<peer address> | GAUGE; 6 established | Not equal to 6 gives CRIT |
| BGP peer, unnumbered or another VRF | 1.3.6.1.4.1.40310.7.3.1.1.2.<row> | GAUGE; 6 established | Not equal to 6 gives CRIT |
For the second BGP row, walk 1.3.6.1.4.1.40310.7.3.1.1.25 once: it names the interface each row belongs to, such as swp57. The standard BGP table only holds numbered IPv4 peers in the default VRF, so on a fabric of unnumbered peers it is empty.
SNMP cannot tell you about MLAG, disk space or the overall health verdict. For those, use the API.
Onyx
SNMP
SNMP is on by default on Onyx, with the read-only community public. Set your own, and limit SNMP to the management port:
snmp-server enable snmp-server community <community> ro snmp-server listen enable snmp-server listen interface mgmt0 configuration write
Onyx has no per-address limit on a community, so restrict SNMP with the management access lists as well. In Oversight, save an SNMP credential with version 2c and the community.
Find the part numbers once:
snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.47.1.1.1.1.2 snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.47.1.1.1.1.5
The first names each part (Onyx names them in the description, such as FAN1/FAN/F1), the second says what each is: 6 a power supply, 7 a fan, 8 a sensor. The numbers are built from the module and sensor, so NVIDIA's own example, 501020021, is fan module 1's first fan sensor.
| What | OID | Settings | Rules |
|---|---|---|---|
| Fan | 1.3.6.1.2.1.99.1.1.1.5.<fan> | GAUGE; 1 ok | Not equal to 1 gives CRIT |
| Temperature sensor failed | 1.3.6.1.2.1.99.1.1.1.5.<sensor> | GAUGE; 1 ok | Not equal to 1 gives CRIT |
| Temperature | 1.3.6.1.2.1.99.1.1.1.4.<sensor> | GAUGE, multiplier 0.1, unit C | Per sensor; a switching chip runs far warmer than a CPU |
| Power supply | 1.3.6.1.2.1.131.1.1.1.3.<supply> | GAUGE; 3 enabled, 2 disabled | Equal to 2 gives CRIT |
| Port up | 1.3.6.1.2.1.2.2.1.8.<port> | GAUGE | Not equal to 1 gives CRIT |
| BGP peer | 1.3.6.1.2.1.15.3.1.2.<peer address> | GAUGE; 6 established | Not equal to 6 gives CRIT |
NVIDIA documents the entity state MIB on Onyx for fans and temperatures, not power supplies, though other monitoring tools read supplies from it. Walk 1.3.6.1.2.1.131.1.1.1.3 once and check the supplies have a row before relying on that sensor.
MLAG through the JSON API
Onyx has no MIB for MLAG, so it comes from Onyx's JSON API. Current releases (3.10 LTS and MLNX-OS 3.12) accept the login and the command in the same request, which is what makes it usable:
web https enable json-gw enable configuration write
| Setting | Value |
|---|---|
| Scheme, port | https, 443 |
| Path | /admin/launch?script=rh&template=json-request&action=json-login |
| Method | POST |
| Headers | Content-Type: application/json |
| Request body | {"username":"{{basicuser}}","password":"{{basicpassword}}","cmd":"show mlag"} |
| Verify the certificate | Off |
| Interval | 300 seconds: each request is a fresh login |
| Reading | Slot and path | Rule |
|---|---|---|
| Command ran | SD03, JSON, status | Not equal to OK gives CRIT: the login or command failed |
| MLAG state | SD04, JSON, data.Operational status | Not equal to Up gives CRIT |
| Peer link | SD05, JSON, data.MLAG IPLs Summary.1.0.Operational State | Not equal to Up gives CRIT |
| MLAG ports inactive | SV03, JSON, data.MLAG Ports Status Summary.Inactive | Greater than 0 gives WARN |
Spaces in the names are fine in a path. Check the healthy word's case once with Store raw response: Up is expected, but has only been seen as Down. Change the default admin and monitor passwords at commissioning, or every login is answered with a request to set them. Onyx 3.7 and earlier need a separate login step first, which a single request cannot make; watch MLAG on those only through the peer link's ports by SNMP.
SONiC
SONiC answers SNMP once a community is added (sudo config snmp community add <community> RO, and sudo config snmpagentaddress add -v mgmt <mgmt-address> for the management VRF), and ports and BGP are then read as for Cumulus above. Its other interfaces need more than one request, so SNMP is the route. Walk the switch once to see which hardware readings it offers.
A starting set
| Switch | Sensors |
|---|---|
| Cumulus Linux 5.10 and later | Overall health; power supplies; fans; each uplink and MLAG peer-link member up; each BGP peer; MLAG peer; disk filling (5.12 and later) |
| Cumulus Linux, SNMP only | Each power supply and fan; the switching chip's temperature sensor; uplinks up; BGP peers |
| Onyx | Fans, temperature sensors and power supplies (SNMP); uplinks and peer-link ports up; BGP peers; MLAG (JSON API) |
At the default 60 seconds from one probe group, an API sensor costs about £3.46 a month and an SNMP sensor about £1.73.
MIBs
None are needed to poll: everything above is standard MIBs typed as numbers, apart from Cumulus's BGP table. To pick them by name, import through Files: ENTITY-MIB, ENTITY-SENSOR-MIB, ENTITY-STATE-MIB and BGP4-MIB, and for Cumulus's unnumbered BGP peers, CUMULUS-SNMP-MIB then CUMULUS-BGPVRF-MIB. NVIDIA publishes these at docs.nvidia.com/networking-ethernet-software/mibs/. Use NVIDIA's copy of ENTITY-MIB, which needs fewer others than the newer edition. No Mellanox private MIB is needed.