The queue service is a read-only dashboard for the NATS server that HWS dispatches work through, and this page also covers running that NATS server itself as a managed microVM behind app-lb.
What it is
Several HWS components share one nats-server with JetStream: CI's job queue, function-runner work queues, and a control plane's sandbox stream. The queue binary is a small host process that shows what is happening on that server:
- Queue depth: every account's streams, how many messages each holds, and how much is pending for each consumer.
- Throughput: messages and bytes per second in and out, charted over time.
- Connected clients: who is attached, from where, and with what subscriptions.
- Logs: a live tail of nats-server's log file.
It also flags the two conditions that usually mean a queue is stuck:
| Flag | Meaning |
|---|---|
no consumer |
A WorkQueue or Interest stream (messages leave only on ack) with no consumer at all. Nothing will ever drain it |
stalled |
A consumer holding work whose ack floor has not moved for longer than its redelivery deadline allows |
Both roll up into a banner at the top of the page.
It runs as a static (proxy_pass) deployment behind app-lb, the same way app-obs does.
What it does not do
- It never connects to NATS as a client. It only makes HTTP
GETrequests to nats-server's monitoring port (/varz,/connz,/jsz). It cannot publish, subscribe, bind consumers, or shut the server down, and it needs no NATS credential. - It persists nothing. Chart history and the log tail live in memory and are lost on restart. For long-term log retention, send nats-server's logs to app-obs.
The monitoring port is unauthenticated, which is why it should stay on loopback. The queue dashboard is how you see that data from outside the host, behind app-lb's sign-in.
Install and run
cargo build --release --locked --manifest-path queue/Cargo.toml
sudo install -m0755 queue/target/release/queue /usr/local/bin/queue
sudo useradd --system --no-create-home --shell /usr/sbin/nologin queue
sudo install -d -o queue -g queue /var/lib/queue
To run it under supervisord, install queue/deploy/supervisor/queue.conf:
sudo cp queue/deploy/supervisor/queue.conf /etc/supervisor/conf.d/
sudo supervisorctl reread && sudo supervisorctl update
The unit starts after nats-server (priority=200). If it loses that race, the first scrape fails and the next one succeeds.
For the log panel, the queue user needs read access to nats-server's log file, for example usermod -aG nats queue or a group-readable /var/log/nats.
To try it locally against a NATS server with monitoring on port 8222:
QUEUE_NATS_LOG_FILE=/var/log/nats/nats-server.log \
cargo run --manifest-path queue/Cargo.toml
# then open http://127.0.0.1:9700/dashboard
nats-server must have its HTTP monitoring listener enabled (http: 127.0.0.1:8222 in its config).
Configuration
Configuration is environment-only; there is no config file and no CLI flags.
| Variable | Default | Meaning |
|---|---|---|
QUEUE_NATS_MONITOR_URL |
http://127.0.0.1:8222 |
Base URL of nats-server's HTTP monitoring listener |
QUEUE_API_ADDR |
127.0.0.1:9700 |
Where the dashboard binds |
QUEUE_API_TOKEN |
unset | Bearer token for the dashboard and its JSON. Unset leaves them open |
QUEUE_NATS_LOG_FILE |
unset | nats-server's log file. Unset disables the log panel |
QUEUE_POLL_SECS |
5 |
Scrape interval, and the resolution of every rate |
QUEUE_HISTORY_POINTS |
720 |
Samples kept for charts (one hour at the default interval) |
QUEUE_LOG_LINES |
2000 |
Log lines held in memory |
QUEUE_LOG_PRIME_BYTES |
65536 |
How much of an existing log file to read at startup |
QUEUE_MAX_CLIENTS |
256 |
Maximum clients pulled from /connz |
QUEUE_REQUEST_TIMEOUT_SECS |
4 |
Deadline for one monitoring request |
QUEUE_UI_COOKIE_DOMAIN |
HEYO_UI_COOKIE_DOMAIN |
Parent domain for the shared light/dark theme cookie |
QUEUE_UI_COOKIE_NAME |
HEYO_UI_COOKIE_NAME, else heyo_theme |
Theme cookie name |
RUST_LOG |
Log filter; the shipped unit uses info,queue=debug |
Values that would make the process useless are corrected with a warning instead of refused:
QUEUE_POLL_SECS=0becomes 1 second.QUEUE_HISTORY_POINTS=0becomes 2.QUEUE_LOG_LINES=0becomes 1.- A request timeout at or above the poll interval is capped at 80% of the interval (minimum 500 ms), so scrapes never overlap.
Unparseable numbers are ignored with a warning and the default is used.
The log panel needs a shared filesystem
NATS writes its own diagnostics (permission violations, slow-consumer disconnects, stream restore failures) only to its log file. Every panel except the log works against a remote monitoring port; the log panel works only when queue and nats-server share a filesystem.
The file is followed, not re-read. At each end-of-file the path is re-checked: a new inode is treated as a rotation and a shorter file as a truncation, and both reopen from the start. This keeps the tail working through supervisord's rename-based rotation.
Lines that don't match nats-server's format ([pid] date time [LVL] message) are kept whole with no level, and are never hidden by a severity filter, since those are usually panics and stack traces.
HTTP API
| Route | Auth | Returns |
|---|---|---|
GET /dashboard |
token | The dashboard page (/ redirects here) |
GET /api/overview |
token | One scrape of the whole server: server info, throughput rates, totals, per-account streams and consumers, clients, and chart history |
GET /api/logs |
token | The log tail |
GET /healthz |
open | ok. Does not check NATS |
GET /__ui/{path} |
open | Shared stylesheet, theme script, and fonts |
"token" means Authorization: Bearer <QUEUE_API_TOKEN> is required when the variable is set.
/healthz deliberately does not probe NATS. If it did, a NATS outage would take the dashboard out of rotation exactly when you need it.
/api/overview reports connected: false and an error string when the last scrape failed.
/api/logs parameters
| Parameter | Meaning |
|---|---|
since |
Sequence number; returns only newer lines. The dashboard polls with this |
level |
Severity floor: trace, debug, info, warn, error, fatal (warn means warn and worse). An unknown level is 400 |
q |
Case-insensitive substring |
limit |
Default 200, maximum 5000. Applied after level and q, so "last 20 errors" means 20 errors |
Register with app-lb
queue/examples/queue.json is the app-lb deployment spec that ships with the service:
{
"id": "queue",
"routes": [{ "host": "queue.example.com" }],
"upstreams": ["127.0.0.1:9700"],
"health": { "path": "/healthz", "timeout_secs": 2 },
"auth": {
"client_id": "REPLACE.apps.googleusercontent.com",
"client_secret": { "secret": "google", "key": "client_secret" },
"allowed_domains": ["example.com"],
"public_paths": ["/healthz", "/__ui/"],
"cookie_domain": "example.com",
"forward_identity": true
}
}
heyctl apply -f queue/examples/queue.json
- Replace the host, client id, allowed domains, and cookie domain. Store the Google client secret with
heyctl create secret google --from-stdin client_secret. See app-lb auth. /__ui/must be public, or the page renders unstyled behind the sign-in redirect.- Setting
cookie_domainto the same parent domain as your other HWS dashboards lets one sign-in cover all of them; settingHEYO_UI_COOKIE_DOMAINto the same value shares the theme choice. forward_identityfills in the name in the top bar. It is display only; the page is read-only.
This page discloses every account. Behind the gate it shows all accounts' stream names, subjects, depths, and consumer names, and every connected client's address, name, and subscriptions. Consider that before pointing it at a server shared with tenants. The process logs a warning at startup when QUEUE_API_TOKEN is unset.
Running NATS as a managed microVM
app-lb/examples/nats/ contains two ways to run nats-server with JetStream inside a Firecracker VM whose lifecycle app-lb owns.
| Variant | Files | Use it for |
|---|---|---|
| Managed | Dockerfile.managed, managed.json, image/start-managed.sh, image/managed.conf |
A broker whose lifecycle and state are independent of CI. Mandatory authentication, workspace-backed state |
| Standalone | image/Dockerfile, image/init.sh, image/nats-server.conf, image/preflight.sh, build-image.sh, nats.json |
A simple single-host broker for development or experiments |
Things that apply to both
app-lb cannot proxy the NATS protocol. app-lb is an HTTP proxy, and NATS is a raw TCP protocol in which the server speaks first. So the deployments split the ports:
| Port | Reached by | How |
|---|---|---|
8222 (HTTP monitoring) |
app-lb | vm.port; the health check probes NATS's own /healthz |
4222 (client protocol) |
NATS clients | vm.open_ports, directly at the VM's guest IP |
Find the address from the admin API. addr is built from vm.port, so take the host part and use 4222:
NATS_IP=$(curl -s localhost:9090/deployments/nats | jq -r '.vms[0].addr' | cut -d: -f1)
export CLOUD_NATS_URL="nats://$NATS_IP:4222"
The guest IP is stable for the life of a sandbox but can change when the VM is recreated, so clients should re-read it rather than bake it in.
Exactly one replica. JetStream here is a single server with a file store, not a cluster. Two replicas would be two independent brokers. Both specs set max_replicas: 1, warm_pool: 0, and idle_action: "retain" (a retired VM is stopped with its disk kept, not destroyed). Scaling out means a real NATS cluster, which these examples don't provide.
Routes are empty. "routes": [] keeps the unauthenticated monitoring port off the proxy. To expose it deliberately, add a route with an auth block:
heyctl set routes nats --host nats.internal.example.com
heyctl set routes nats --none # withdraw
Don't expose 4222 or the monitoring port to untrusted networks. Cross-host client access needs a private encrypted network or TLS, not just an open firewall port.
Managed variant
managed.json is a template, not a ready deployment:
{
"id": "nats-managed",
"routes": [],
"vm": {
"driver": "firecracker",
"image": "REPLACE_WITH_VERIFIED_NATS_2_11_17_IMAGE",
"port": 8222,
"open_ports": [4222],
"size_class": "small",
"disk_size_gb": 20,
"ttl_seconds": 0,
"start_command": "setsid nohup /opt/nats/start.sh </dev/null >/tmp/nats-boot.log 2>&1 &",
"env_from": [
{"secret": "nats-managed", "key": "token", "as": "NATS_TOKEN"}
],
"workspace": {
"path": "/workspace",
"store": "https://REPLACE_WITH_PRIVATE_ARTIFACT_STORE",
"ref": "workspace-nats-managed",
"auth": {"secret": "nats-artifacts", "key": "api-key"}
}
},
"build": {
"repo": "https://github.com/Heyo-Computer/heyo-public.git",
"ref": "REPLACE_WITH_CI_VERIFIED_COMMIT",
"dockerfile": "app-lb/examples/nats/Dockerfile.managed",
"context": ".",
"image_size_mb": 512
},
"scaling": {
"min_replicas": 0,
"max_replicas": 1,
"warm_pool": 0,
"scale_to_zero_after_secs": 0,
"boot_timeout_secs": 180,
"idle_action": "retain"
},
"health": {"path": "/healthz?js-enabled-only=true", "timeout_secs": 5}
}
How it works:
- The image contains NATS 2.11.17. PID 1 only prepares the guest; app-lb mounts the workspace and injects
NATS_TOKENfrom the secret store before running/opt/nats/start.sh. start.shrefuses to start unlessNATS_TOKENis set,/workspaceis a separate mounted filesystem,/workspace/.managed-statecontainsnats-state-v1, and/workspace/jetstreamis a real directory on that filesystem. It never formats a disk or creates missing state.- The launcher JSON-quotes the token into
NATS_CONFIG_TOKENso NATS's config parser treats it as an opaque string. Store the raw token; don't pre-quote it. - JetStream data and the NATS log (
/workspace/nats.log, rotated at 16 MB) live on the workspace.vm.workspaceis what carries state across VM replacement;disk_size_gbalone does not.
To activate it:
- Build the image through CI or
build, and setvm.imageto the verified image. - Create the
nats-managedsecret (keytoken) and thenats-artifactssecret holding the store's API key. See heyosecret. - Point
workspace.storeat a private artifacts store and use a dedicatedref. Never reuse CI's workspace or snapshot tag. - Seed the workspace with the marker file and an empty
jetstreamdirectory (or a restored store). - Set
min_replicas: 1.
Limits to be aware of:
- Replacing the broker has a stop, capture, restore gap. It is not a replicated cluster or a zero-downtime upgrade.
- Workspace capture is crash-consistent storage recovery, not a JetStream backup. Before a planned migration, stop producers and consumers and take a JetStream network backup with consumers, then restore and compare stream and consumer state before switching clients. Never copy a live JetStream directory as a backup.
- The marker file guards initialization; it does not prove a restore succeeded.
Test the image locally from the repository root (Docker and the nats CLI required):
docker build --platform linux/amd64 -f app-lb/examples/nats/Dockerfile.managed \
-t heyo-nats-managed-test:2.11.17 .
NATS_TEST_IMAGE=heyo-nats-managed-test:2.11.17 \
python3 app-lb/examples/nats/test_managed.py -v
These tests cover authentication, startup rejection, backup and restore with pending acks, and restart. They don't boot Firecracker or exercise app-lb workspace capture.
Standalone variant
cd app-lb/examples/nats
./build-image.sh # -> <heyvm images dir>/nats.ext4
heyctl apply -f nats.json
heyctl rollout status nats
heyctl exec nats -- /opt/nats/preflight.sh
build-image.sh wraps heyvm mvm build with this directory as the context. It needs docker, mke2fs, heyvm, and fakeroot (unless run as root). Optional environment: IMAGE_NAME (default nats), SIZE_MB, DNS_SERVER.
How it differs from the managed variant:
nats.jsonuses adisk_size_gb: 20data disk, not a workspace.init.shformats/dev/vdbif it has no filesystem, mounts it at/workspace, and refuses to start nats-server if/workspaceis not a mount. The disk belongs to one sandbox; it survives stop and resume (idle_action: "retain") and image rebuilds, but not VM recreation.- JetStream lives at
/workspace/jetstream; the log is at/workspace/log/nats-server.log. - Authentication is off by default. Add it by writing a config fragment to the data disk, which
init.shincludes when present:
heyctl exec nats -- sh -c 'mkdir -p /workspace/nats && cat > /workspace/nats/auth.conf <<EOF
authorization { token: "<token>" }
EOF'
heyctl restart nats
ttl_seconds: 86400is a backstop that stops the VM if app-lb dies and stops renewing it.
The Dockerfile pins a nats-server patch release. Bump it deliberately and check the release notes for store-format changes, since the binary is rebuilt but the store on disk is not.
Operating NATS
heyctl get deployments
heyctl exec nats -- /opt/nats/preflight.sh # standalone: durability and listeners
heyctl shell nats
curl -s "http://$NATS_IP:8222/jsz?streams=1" | jq .
curl -s "http://$NATS_IP:8222/varz" | jq '{uptime, connections, in_msgs, out_msgs}'
Point the queue dashboard at a VM-hosted broker with QUEUE_NATS_MONITOR_URL=http://<guest ip>:8222. The log panel won't work in that setup, since the log file is inside the guest.
Troubleshooting
| Symptom | Check |
|---|---|
| Dashboard says it can't reach the server | QUEUE_NATS_MONITOR_URL is wrong, or nats-server has no http: listener |
| Log panel says it is disabled | Set QUEUE_NATS_LOG_FILE |
| Log panel empty after setting the path | The queue user can't read the file, or nats-server logs somewhere else |
| Rates show nothing | Rates need two scrapes; wait one poll interval |
A stream is flagged no consumer |
A worker that should consume it isn't running or hasn't created its consumer |
A consumer is flagged stalled |
Its worker is attached but not acking; check the worker's logs |
/api/logs?level=... returns 400 |
Use trace, debug, info, warn, error, or fatal |
| NATS clients can't connect through the app-lb hostname | Expected; connect directly to <guest ip>:4222 |
| Managed broker won't start | Check /tmp/nats-boot.log in the guest for which start.sh precondition failed |
For direct inspection, the monitoring port is the source of truth:
curl -s 'http://127.0.0.1:8222/jsz?accounts=1&streams=1&consumers=1&config=1' | jq
curl -s http://127.0.0.1:9700/api/overview | jq '.accounts[].streams[]'
Stream names, message counts, and consumer numbers should agree between the two; rates exist only in the dashboard.