Compare commits

...
1 Commits
Author SHA1 Message Date
JesseMarkowitz 10d8a988ca DEVELOPMENT.md: keep inference-host logs in one directory
CI / Backend tests (push) Canceled after 0s
CI / Frontend lint + build (push) Canceled after 0s
CI / Docker image builds (push) Canceled after 0s
2026-10-08 05:34:53 -04:00
+9 -6
View File
@@ -471,29 +471,32 @@ logging on that host first and stop it only when the run has finished.**
Run each command in its own terminal on the inference host. `tee` writes each Run each command in its own terminal on the inference host. `tee` writes each
line as it arrives, so what happened in the seconds before a crash or a forced line as it arrives, so what happened in the seconds before a crash or a forced
reboot survives on disk. reboot survives on disk. Every log goes into one directory, so a campaign's
evidence stays together and is easy to archive or remove afterwards:
```bash ```bash
mkdir -p "$HOME/inference-host-logs"
# Power, temperature, utilisation and PCIe link state, once a second # Power, temperature, utilisation and PCIe link state, once a second
nvidia-smi --query-gpu=timestamp,pcie.link.gen.current,pcie.link.width.current,power.draw,temperature.gpu,utilization.gpu \ nvidia-smi --query-gpu=timestamp,pcie.link.gen.current,pcie.link.width.current,power.draw,temperature.gpu,utilization.gpu \
--format=csv -l 1 | tee "$HOME/gpu-link-$(date +%F-%H%M).csv" --format=csv -l 1 | tee "$HOME/inference-host-logs/gpu-link-$(date +%F-%H%M).csv"
# The driver's own sampling: power, utilisation, clocks, memory, ECC and throttling # The driver's own sampling: power, utilisation, clocks, memory, ECC and throttling
nvidia-smi dmon -s pucvmet -d 5 | tee "$HOME/gpu-dmon-$(date +%F-%H%M).log" nvidia-smi dmon -s pucvmet -d 5 | tee "$HOME/inference-host-logs/gpu-dmon-$(date +%F-%H%M).log"
# Kernel and Ollama messages, live. The `+` is an OR: `journalctl -k -u ollama` # Kernel and Ollama messages, live. The `+` is an OR: `journalctl -k -u ollama`
# asks for messages that are both kernel messages and the ollama unit's, which # asks for messages that are both kernel messages and the ollama unit's, which
# is none, and writes an empty log. # is none, and writes an empty log.
journalctl -f -o short-iso _TRANSPORT=kernel + _SYSTEMD_UNIT=ollama.service \ journalctl -f -o short-iso _TRANSPORT=kernel + _SYSTEMD_UNIT=ollama.service \
| tee "$HOME/ollama-kernel-watch-$(date +%F-%H%M).log" | tee "$HOME/inference-host-logs/ollama-kernel-watch-$(date +%F-%H%M).log"
``` ```
If the GPU faults, find the moment and then read what the card was doing just If the GPU faults, find the moment and then read what the card was doing just
before it: before it:
```bash ```bash
grep -iE 'xid|fallen off|nvrm' "$HOME"/ollama-kernel-watch-*.log grep -iE 'xid|fallen off|nvrm' "$HOME"/inference-host-logs/ollama-kernel-watch-*.log
awk -F', ' 'NR>1 && $4+0 > max {max=$4+0; at=$1} END {print "peak W", max, "at", at}' "$HOME"/gpu-link-*.csv awk -F', ' 'NR>1 && $4+0 > max {max=$4+0; at=$1} END {print "peak W", max, "at", at}' "$HOME"/inference-host-logs/gpu-link-*.csv
``` ```
A fault that follows sustained draw at the card's power limit points to power A fault that follows sustained draw at the card's power limit points to power