Articles setup, updated 2026-09-28

NVIDIA (Mellanox): Onyx, Cumulus Linux and SONiC switches

NVIDIA's Spectrum Ethernet switches, formerly Mellanox, can run any of three operating systems on the same hardware: Onyx (also called MLNX-OS), Cumulus Linux or SONiC. How to watch a switch depends on which it runs, and the model number does not tell you: an SN2700 or SN4600C can have any of them. So find out first.

Which one is it?

Read the switch's description by SNMP, 1.3.6.1.2.1.1.1.0, or log in and look at the prompt:

Operating systemDescription reads likePrompt
Onyx / MLNX-OSOnyx,MSN2410,SWv3.6.6102 or Mellanox MSN2700,MLNX-OS,SWv3.6.4112switch (config) #
Cumulus LinuxCumulus-linux 5.9.2 (Linux Kernel ...)cumulus@switch:~$
SONiCSONiC Software Version: SONiC...admin@sonic:~$

In short: on Cumulus Linux 5, use its REST API, which reports everything including MLAG and an overall health verdict. On Onyx, use SNMP for hardware, ports and BGP, and its JSON API for MLAG only. Neither needs any NVIDIA MIB loaded.

Cumulus Linux: the NVUE API

From Cumulus Linux 5.6 the NVUE REST API is on by default, on port 8765, and takes a username and password in each request. It returns the switch's live state as JSON.

Setting it up

nv set system api state enabled
nv set system api port 8765
nv set system api listening-address <eth0-address>
nv set system aaa user oversight role nvue-monitor
nv set system aaa user oversight password
nv config apply

nvue-monitor is Cumulus's read-only role: it can show, not change. Two things to know:

  • The API refuses everyone with 403 until the cumulus user's password has been changed from its default, even though Oversight uses its own account.
  • On Cumulus 5.5 and earlier the API is off and switching it on means editing the web server by hand. Treat those switches as SNMP only (below).

In Oversight, save an HTTP(S) credential with the username and password, and give each sensor https, port 8765, GET, the header Authorization: Basic {{basicauth}}, and Verify the certificate off, since the switch makes its own certificate at first boot. Paths start /nvue_v1. Add SV01 not equal to 200 gives CRIT to each.

Never add ?rev=applied to a path: that returns the configuration, not the live state, so a failed power supply or a downed port would look exactly as configured.

Sensors worth having

WhatPathExtractionRules
Overall health/nvue_v1/system/healthSD03, JSON, statusSD03 not equal to OK gives CRIT. Cumulus 5.10 and later.
Any power supply not ok/nvue_v1/platform/environment/psuSD03, REGEX, belowSD03 matching /\w/ gives CRIT
Any fan not ok/nvue_v1/platform/environment/fanSD03, REGEX, belowSD03 matching /\w/ gives CRIT
A port up/nvue_v1/interface/swp51SD03, JSON, link.oper-statusSD03 not equal to up gives CRIT. Cumulus 5.9 and later.
A BGP peer/nvue_v1/vrf/default/router/bgp/neighbor/10.0.0.1, or /neighbor/swp51 for an unnumbered peerSD03, JSON, stateSD03 not equal to established gives CRIT
MLAG peer/nvue_v1/mlagSD03, JSON, peer-alive; SD04, JSON, backup-activeSD03 not equal to True gives CRIT; SD04 not equal to True gives WARN
Disk filling/nvue_v1/system/disk/usageSV03, REGEX, /"\/":\s*\{[^}]*"used-percent":\s*"(\d+)%"/, unit %At least 90 gives WARN, at least 95 gives CRIT. Cumulus 5.12 and later.
Restarted/nvue_v1/systemSV03, JSON, uptime, unit sLess than 600 gives WARN

The power supply and fan sensors use one REGEX that finds the first item whose state is anything but ok or absent, and gives nothing at all when every item is healthy:

/^(?:(?=[\s\S]*?"state":\s*"(?!ok"|absent")(\w+)"))?/

That way one sensor covers every fan, without knowing how the model names them. absent is let through on purpose: a single-supply switch reports its empty bay as absent, and would otherwise alarm for ever. Where both supplies must be fitted, remove |absent" from the expression.

Some notes:

  • Case matters. Rules compare text exactly. NVUE uses lower case (ok, up, established), but MLAG says True and the health check OK. If a rule fires on a healthy switch, tick Store raw response and compare the case.
  • Overall health covers the switching chip, the hardware and the system's own processes in one verdict, which makes it the best single sensor. Where its healthy word is not exactly OK on your release, the stored response will show it.
  • Temperature. /nvue_v1/platform/environment/temperature (5.9 and later; /sensor before that) lists each sensor with current, max and crit, all in degrees C. The switch's own health verdict already watches these, so a sensor of your own is only needed for an earlier warning: read the switching chip's current, and set WARN at its own max. Do not use one temperature for every sensor: a healthy SN4600's switching chip runs at about 71 C, while its CPU cores sit in the 40s.
  • A known fault on Cumulus 5.11.3 to 5.17.0 makes the MLAG JSON come back as plain text. If it affects the API too, the MLAG sensor finds nothing to read and cannot tell; it does not raise a false alarm. Upgrading clears it.

Cumulus Linux: SNMP

For switches on 5.5 or earlier, or where the API is not wanted. The command changed at 5.11:

# Cumulus Linux 5.11 and later
nv set system snmp-server state enabled
nv set system snmp-server listening-address <eth0-address> vrf mgmt
nv set system snmp-server readonly-community <community> access <probe-address>/32
nv config apply

# Cumulus Linux 5.3 to 5.10
nv set service snmp-server enable on
nv set service snmp-server listening-address <eth0-address> vrf mgmt
nv set service snmp-server readonly-community <community> access <probe-address>/32
nv config apply

On 5.0 to 5.2 there are no NVUE commands for SNMP: add agentAddress <eth0-address>@mgmt and rocommunity <community> <probe-address>/32 to /etc/snmp/snmpd.conf and restart snmpd. On 5.3 and later, do not also edit that file: NVIDIA says not to mix the two.

  • Without a listening address, SNMP answers only the switch itself.
  • The management VRF is on by default, and SNMP does not cross VRFs, which is why the listening address says vrf mgmt.
  • Switching SNMP off later restarts the routing service, which may interrupt traffic. Worth knowing before trying it during commissioning.

Find the part and port numbers once:

snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.47.1.1.1.1.7
snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.31.1.1.1.1

Cumulus numbers parts in ranges: temperature sensors from 100000001, fans from 100011001, power supplies from 110000001. The first walk shows which exist.

WhatOIDSettingsRules
Power supply1.3.6.1.2.1.99.1.1.1.5.<supply>, such as ...5.110000001GAUGE; 1 okNot equal to 1 gives CRIT
Fan1.3.6.1.2.1.99.1.1.1.5.<fan>GAUGE; 1 okNot equal to 1 gives CRIT
Temperature sensor failed1.3.6.1.2.1.99.1.1.1.5.<sensor>GAUGE; 1 okNot equal to 1 gives CRIT
Temperature1.3.6.1.2.1.99.1.1.1.4.<sensor>GAUGE, multiplier 0.1, unit CPer sensor, at the limit the API or nv show platform environment temperature gives for it
Port up1.3.6.1.2.1.2.2.1.8.<port>GAUGENot equal to 1 gives CRIT
BGP peer, numbered, default VRF1.3.6.1.2.1.15.3.1.2.<peer address>GAUGE; 6 establishedNot equal to 6 gives CRIT
BGP peer, unnumbered or another VRF1.3.6.1.4.1.40310.7.3.1.1.2.<row>GAUGE; 6 establishedNot equal to 6 gives CRIT

For the second BGP row, walk 1.3.6.1.4.1.40310.7.3.1.1.25 once: it names the interface each row belongs to, such as swp57. The standard BGP table only holds numbered IPv4 peers in the default VRF, so on a fabric of unnumbered peers it is empty.

SNMP cannot tell you about MLAG, disk space or the overall health verdict. For those, use the API.

Onyx

SNMP

SNMP is on by default on Onyx, with the read-only community public. Set your own, and limit SNMP to the management port:

snmp-server enable
snmp-server community <community> ro
snmp-server listen enable
snmp-server listen interface mgmt0
configuration write

Onyx has no per-address limit on a community, so restrict SNMP with the management access lists as well. In Oversight, save an SNMP credential with version 2c and the community.

Find the part numbers once:

snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.47.1.1.1.1.2
snmpwalk -v2c -c <community> <switch> 1.3.6.1.2.1.47.1.1.1.1.5

The first names each part (Onyx names them in the description, such as FAN1/FAN/F1), the second says what each is: 6 a power supply, 7 a fan, 8 a sensor. The numbers are built from the module and sensor, so NVIDIA's own example, 501020021, is fan module 1's first fan sensor.

WhatOIDSettingsRules
Fan1.3.6.1.2.1.99.1.1.1.5.<fan>GAUGE; 1 okNot equal to 1 gives CRIT
Temperature sensor failed1.3.6.1.2.1.99.1.1.1.5.<sensor>GAUGE; 1 okNot equal to 1 gives CRIT
Temperature1.3.6.1.2.1.99.1.1.1.4.<sensor>GAUGE, multiplier 0.1, unit CPer sensor; a switching chip runs far warmer than a CPU
Power supply1.3.6.1.2.1.131.1.1.1.3.<supply>GAUGE; 3 enabled, 2 disabledEqual to 2 gives CRIT
Port up1.3.6.1.2.1.2.2.1.8.<port>GAUGENot equal to 1 gives CRIT
BGP peer1.3.6.1.2.1.15.3.1.2.<peer address>GAUGE; 6 establishedNot equal to 6 gives CRIT

NVIDIA documents the entity state MIB on Onyx for fans and temperatures, not power supplies, though other monitoring tools read supplies from it. Walk 1.3.6.1.2.1.131.1.1.1.3 once and check the supplies have a row before relying on that sensor.

MLAG through the JSON API

Onyx has no MIB for MLAG, so it comes from Onyx's JSON API. Current releases (3.10 LTS and MLNX-OS 3.12) accept the login and the command in the same request, which is what makes it usable:

web https enable
json-gw enable
configuration write
SettingValue
Scheme, porthttps, 443
Path/admin/launch?script=rh&template=json-request&action=json-login
MethodPOST
HeadersContent-Type: application/json
Request body{"username":"{{basicuser}}","password":"{{basicpassword}}","cmd":"show mlag"}
Verify the certificateOff
Interval300 seconds: each request is a fresh login
ReadingSlot and pathRule
Command ranSD03, JSON, statusNot equal to OK gives CRIT: the login or command failed
MLAG stateSD04, JSON, data.Operational statusNot equal to Up gives CRIT
Peer linkSD05, JSON, data.MLAG IPLs Summary.1.0.Operational StateNot equal to Up gives CRIT
MLAG ports inactiveSV03, JSON, data.MLAG Ports Status Summary.InactiveGreater than 0 gives WARN

Spaces in the names are fine in a path. Check the healthy word's case once with Store raw response: Up is expected, but has only been seen as Down. Change the default admin and monitor passwords at commissioning, or every login is answered with a request to set them. Onyx 3.7 and earlier need a separate login step first, which a single request cannot make; watch MLAG on those only through the peer link's ports by SNMP.

SONiC

SONiC answers SNMP once a community is added (sudo config snmp community add <community> RO, and sudo config snmpagentaddress add -v mgmt <mgmt-address> for the management VRF), and ports and BGP are then read as for Cumulus above. Its other interfaces need more than one request, so SNMP is the route. Walk the switch once to see which hardware readings it offers.

A starting set

SwitchSensors
Cumulus Linux 5.10 and laterOverall health; power supplies; fans; each uplink and MLAG peer-link member up; each BGP peer; MLAG peer; disk filling (5.12 and later)
Cumulus Linux, SNMP onlyEach power supply and fan; the switching chip's temperature sensor; uplinks up; BGP peers
OnyxFans, temperature sensors and power supplies (SNMP); uplinks and peer-link ports up; BGP peers; MLAG (JSON API)

At the default 60 seconds from one probe group, an API sensor costs about £3.46 a month and an SNMP sensor about £1.73.

MIBs

None are needed to poll: everything above is standard MIBs typed as numbers, apart from Cumulus's BGP table. To pick them by name, import through Files: ENTITY-MIB, ENTITY-SENSOR-MIB, ENTITY-STATE-MIB and BGP4-MIB, and for Cumulus's unnumbered BGP peers, CUMULUS-SNMP-MIB then CUMULUS-BGPVRF-MIB. NVIDIA publishes these at docs.nvidia.com/networking-ethernet-software/mibs/. Use NVIDIA's copy of ENTITY-MIB, which needs fewer others than the newer edition. No Mellanox private MIB is needed.