How to monitor particular kit with Oversight: which sensor to use, where on the device to set up access, which user function and extraction rows read the response, and how to write the rules that turn it into an alarm. Search for a vendor, a product or an error message. A misspelt word still finds its article when nothing matches exactly.
Docker's own Engine API has no bearer tokens, no API keys and no user accounts. The only way to reach it across a network is a TLS client certificate, and a certificate that can list containers can equally create a privileged one with the host's filesystem mounted inside it. There is no read-only credential to issue. That is why Oversight does not talk to Docker directly.
Instead a small container, the Oversight Docker API, runs on the Docker host. It holds the local socket, answers three questions over plain HTTP, and requires a bearer token on every request. It is roughly two hundred lines of Go on an empty image with no shell and no package manager, so the thing holding the socket is as small as it can reasonably be made.
Installing it on the Docker host
Everything below is run on the host you want to watch, as a user who can run docker. Pick the image for that host's architecture: uname -m reports x86_64 for amd64 and aarch64 for arm64.
mkdir -p /root/oversight && cd /root/oversight
curl -fsSLO https://oversight.im/docker/compose.yml
curl -fsSLO https://oversight.im/docker/dockerapi-0.4.0-amd64.tar.gz
docker load < dockerapi-0.4.0-amd64.tar.gz
The tarball carries both the version tag and latest, which is what the compose file refers to, so nothing needs editing to match a version number.
compose.yml must be edited before it will start
This is not optional and the service will not come up until it is done. The file is commented throughout, so read it rather than this article for the detail, but three things need your attention:
- TOKEN. It ships as
CHANGE-MEand the service refuses to start on anything under sixteen characters. Generate one per host withopenssl rand -hex 24and do not reuse it anywhere else. This token can stop and restart containers, so treat it as control of the host, not as a read-only password. - The published port. It ships bound to
10.0.0.1, which is deliberately an address your host almost certainly does not have, so that it cannot come up on a public interface by accident. Change it to the address the Oversight probe reaches. A bare9021:9021publishes it on every interface the host has. - The docker group id. The container runs as nobody and needs the host's docker group to read the socket. Write it to a
.envfile beside the compose file, rather than exporting it, because compose substitutes it at parse time and an exported variable is lost on the nextupfrom a fresh shell.
echo "DOCKER_GID=$(getent group docker | cut -d: -f3)" > .env
docker compose up -d
docker compose logs --no-log-prefix
The log should read oversight-dockerapi 0.4.0 listening on :9021, socket /var/run/docker.sock. Check it answers before going near Oversight, substituting your own token:
curl -s -H "Authorization: Bearer YOUR-TOKEN" http://127.0.0.1:9021/health
If that returns "daemon":"down" with a permission error, the group id in .env is not the one that owns the socket. Check it with stat -c '%g' /var/run/docker.sock and put that number in directly. Some hosts, Synology among them, have no docker group at all and the socket is owned by root, in which case getent returns nothing and the fallback of 999 is wrong.
Upgrading it later
On the host, in the directory holding compose.yml:
curl -fsSL https://oversight.im/docker/upgrade.sh -o upgrade.sh && sh upgrade.sh
It works out the architecture, fetches the current build, checks it against the published checksum, loads it and recreates the container. Where the installed version is already the current one it says so and stops, unless FORCE=1 is set.
It never writes to compose.yml, because that file holds this host's token and published address. Where the published compose.yml has changed, because a new version has added a setting, it tells you and prints the command to compare the two.
What it provides
Three endpoints, all requiring Authorization: Bearer <token>. Anything without a valid token is answered with 401 and never reaches Docker.
| Request | What comes back |
|---|---|
GET /health | The daemon as a whole: daemon (up or down), version, containers, running, paused, stopped, images, and any warnings the daemon reports about itself. |
GET /containers | Every container that exists, as a set of names: 1 for running, 0 for anything else. Nothing more. |
GET /containers?name=<name> | The same, narrowed to one container. |
POST /container/<name>/startPOST /container/<name>/stopPOST /container/<name>/restart | Acts on one container. Returns ok true when it worked, with already true where the container was in that state to begin with, so asking twice is not an error. |
Two things worth knowing about the readings. Health only exists where the image declares a healthcheck, so health is empty on most containers rather than healthy, and that is not a fault. And a container that has been removed is not in the list at all, rather than appearing as stopped, so a count of stopped containers will never notice a container that a compose stack has torn down. That is what total is for.
A sensor for the daemon as a whole
Build an HTTPS sensor on the device that represents the Docker host. First hold the token: on the Credentials panel of that device add a credential of type HTTPS and put the token in the Bearer token field. It is sealed on save and never shown again, and the sensor refers to it as a variable rather than carrying a copy.
| Setting | What to put |
|---|---|
| Scheme | http. The service speaks plain HTTP; it holds no certificate of its own. |
| Port | 9021, or whatever you published it on. |
| Path | /health |
| Method | GET |
| Headers | Authorization: Bearer {{bearer}} |
| Interval | 60 seconds is ample. The call is one request to a local socket and costs the host almost nothing, but there is no benefit in asking faster than you would act. |
Then the extraction rows, which say what goes in each slot:
| Slot | Name | Method | Expression |
|---|---|---|---|
| SV01 | Running | JSON | running |
| SV02 | Stopped | JSON | stopped |
| SV03 | Containers | JSON | containers |
| SD01 | Daemon | JSON | daemon |
| SD02 | Version | JSON | version |
Useful rules on those: Stopped greater than 0 for a container that has fallen over, and Containers not equal to the number you expect, which is the one that catches a container removed rather than stopped. Count the service's own container when you set that number: it appears in its own list.
A sensor for named containers
One HTTPS sensor with the path /containers covers every container on the host. The response names them, and nothing else:
{"element-call-jwt":0,"element-call-livekit":1,"pihole":1,"vaultwarden":1}
So the extraction expression is the container's name, and the value is already the number a rule tests. Twenty numeric slots means up to twenty containers on one sensor and one reading:
| Slot | Name | Method | Expression |
|---|---|---|---|
| SV01 | LiveKit | JSON | element-call-livekit |
| SV02 | JWT service | JSON | element-call-jwt |
| SV03 | Pi-hole | JSON | pihole |
| SV04 | Vaultwarden | JSON | vaultwarden |
Each slot is 1 when that container is running and 0 when it is not, so the rule is equal to 0 on the ones that matter, with whatever severity and action that container deserves. Put only the containers you would act on in slots; there is no value in a slot for something you would not get out of bed for.
A container that has been removed is not in the set at all, so its expression extracts nothing and the sensor goes UNKNOWN rather than reading 0. That is deliberate. A container someone has torn down is a different problem from one that has stopped, and it should not quietly look the same.
?name= narrows the response to a single container. The name is already the expression, so there is no need for it when reading slots; it is there for a caller that wants one container rather than the estate.
When Docker itself stops
This service runs inside the Docker it reports on. Stopping the daemon stops every container on the host, including this one, so the sensor does not report a stopped daemon: it fails to connect. That is the signal, and it is the right one, but on its own it cannot tell a stopped daemon from a dead host or a broken network path.
Give the Docker sensor a dependency on a ping of the same host. With the ping up and the Docker sensor failing, Docker is down; with the ping down as well, the host is gone and the Docker alarm is suppressed as a consequence of it rather than raised on its own. Without the dependency both situations page you identically.
The daemon slot still earns its place. It reports down, rather than the sensor failing, in the cases where the daemon has gone but its containers have not: dockerd killed rather than stopped cleanly, or a host running with live-restore enabled. Enabling live-restore is good practice on a production host anyway, since it keeps containers up across a daemon upgrade, and it has the side effect of letting this sensor report a stopped daemon properly instead of going dark.
If it does not work
| What you see | What it is |
|---|---|
| HTTP 401 from the sensor | The token does not match. Check the credential resolves on the device (the sensor's setup names which object it came from) and that the header reads exactly Authorization: Bearer {{bearer}}. |
| Connection refused | The published port in compose.yml is not an address the probe can reach, or the container did not start. docker compose logs on the host says which. |
"daemon":"down" with a permission error | The docker group id in .env is wrong for that host. |
| Sensor goes UNKNOWN | An extraction found nothing: the named container has been removed, or its name is misspelt in the expression. Container names are exact, and the list from docker ps -a --format '{{.Names}}' on the host is what to check against. |
| HTTP 400, bad container name | The name in the path is not a name Docker would accept. Names are letters, digits, and then underscore, dot or hyphen. |