Files
knowledge-wiki/channels/1519072621130547220.md
T

60 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Channel Wiki: #jarvis-jr-truenas
_Channel ID: 1519072621130547220_
_Last sync: 2026-08-27 03:03 UTC_
## Purpose
TrueNAS, self-hosted infrastructure, registry, proxy, and platform setup discussions.
## Key Context
- Dedicated channel skill: `home-v1--truenas`.
- Authoritative private Gitea repository: https://gitea.lego-cloud.eu/home-v1-skills-code-agent/home-v1--truenas (`test` default branch).
- Hermes loads the skill directly from `/opt/data/home-v1-skills-code-agent/home-v1--truenas` through `skills.external_dirs`, so repository updates and runtime behavior share one source tree.
- TrueNAS WebSocket access expects runtime variables `HL_V1_TRUENAS_URL` and `HL_V1_TRUENAS_API_KEY`; values must remain in memory and must never be printed or committed.
- The skill owns reusable `scripts/truenas_rpc.py` JSON-RPC logic and `scripts/truenas_health.py` read-only health summaries, with offline tests and Gitea Actions validation.
- **CONFIRMED delivery (2026-08-26):** initial skill commit `634e0e63290cb4a8c44b961c89d0583cc2e7612f` is on the private repository's `test` branch; authenticated remote-SHA/artifact readback passed, all 10 offline tests passed, and Gitea Actions task `1398` passed. A fresh Hermes loader reported the skill available and discovered both scripts.
- **CONFIRMED access check (2026-08-26):** both expected keys exist in Bitwarden Secrets Manager and a short-lived authenticated WebSocket health query succeeded without changing TrueNAS. Strict TLS verification failed because the configured URL uses an IP address whose identity does not match the certificate; diagnostic-only `--insecure` access succeeded. The durable fix is a certificate-matching DNS URL or correct CA/hostname path—not permanent TLS disablement.
- **Live read-only observation, not yet investigated:** TrueNAS reported `26.0.0-BETA.2`; pool `lego-cloud-v1-2025-11-28` was `ONLINE` and `healthy: true` but also reported `status_code: CORRUPT_DATA`, with 3 active alerts. The setup request did not authorize investigation or remediation, and no action was taken.
- Keycloak hostname: `keycloak.lego-cloud.eu`; historical discussion referenced the `master` realm.
- Gitea organization URL: https://gitea.lego-cloud.eu/home-v1
- Repository referenced for shared access: https://gitea.lego-cloud.eu/home-v1/arda-v1-code-agent
- Historical work investigated Harbor installation on TrueNAS SCALE using a custom-app YAML.
- The external Nginx reverse proxy already existed in front of TrueNAS; Harbor planning should not add an unnecessary duplicate Nginx layer.
- Services exposed through the existing HTTPS path still require explicit internal/service ports and routing.
## Behavioral Decisions
- Teach and train without negativity or pointing out what others lack.
- Operate autonomously enough that RootAtSkic does not need to micromanage routine investigation.
- Keep `home-v1--truenas` continuously updated when durable TrueNAS methods, safeguards, API quirks, or reusable procedures are confirmed; maintain logic as tested scripts in its Gitea repository rather than accumulating one-off inline commands.
## Active Topics
- TrueNAS application deployment patterns.
- Harbor/container-registry architecture and certificate/routing requirements.
- **CONFIRMED capacity objective (RootAtSkic, 2026-08-17):** maximize this TrueNAS host's memory. For the documented W680D4U-2L2T/G5 / i9-14900KS platform, the target is `192 GB ECC` (`4 × 48 GB DDR5 ECC UDIMM`), not a smaller diagnostic kit. Procurement and validation constraints are recorded below.
- **OPEN investigation request (RootAtSkic, 2026-08-18):** review the attached Hermes container startup log and identify improvements, but **do not change anything yet**. No response or approved implementation was present in the fetched history; RootAtSkic followed up twice asking whether Hermes was alive. Source attachment: Discord message `1539151473189978142`, `message.txt`.
## Hermes startup-log investigation — 2026-08-18 (read-only)
- **CONFIRMED log observations:** container setup exited successfully and the supervised gateway started, but setup spent `424.296399` seconds between “Changing hermes UID to 568” and “Changing hermes GID to 568” (`425.036112` seconds from the UID step through setup completion). The cause is **UNKNOWN** from this log alone; recursive ownership work is a candidate to investigate, not a confirmed diagnosis.
- The lifecycle ledger reported the previous gateway life as an unclean exit consistent with `SIGKILL / OOM / VM death`, while recording `suspected_oom=False`, about 297,604 KiB RSS, abundant available memory and no swap use at its last heartbeat. The exact exit cause remains **UNKNOWN**.
- Startup marked two interrupted cron executions `unknown` after restart. It also logged repeated operational friction: `hermes: command not found` in tool shells, `discord.server_actions` configured with unsupported action `all`, unavailable web-dependent tools because `check_web_api_key` returned false, Git commands launched outside a repository, and one command held for security approval.
- **CONFIRMED security warning:** the API server was bound to `0.0.0.0` while `terminal.backend` was `local`/unsandboxed, meaning trusted-network firewalling or a sandboxed terminal backend should be evaluated. This is an investigation finding only; RootAtSkic explicitly prohibited configuration changes at this stage.
## Live DIMM inventory — 2026-08-25 (read-only)
- Host SMBIOS Type 17 reports **96 GiB installed** as **2 × 48 GiB G.Skill DDR5 UDIMMs**, part `F5-5200J4040A48G`, dual-rank. The DIMMs advertise 4800 MT/s in SMBIOS and are configured at 4400 MT/s.
- The board manual identifies the physical channels/slots as `DDR5_A1`, `DDR5_A2`, `DDR5_B1`, and `DDR5_B2`; its priority-two-DIMM configuration is **`DDR5_A2` + `DDR5_B2`**. This matches RootAtSkic's recollection of one DIMM in channel A and one in channel B.
- Firmware exposes those two occupied records under Intel internal locator names `Controller0-ChannelA-DIMM1` and `Controller1-ChannelA-DIMM1`; these strings must not be presented as the motherboard's physical channel labels. Their SMBIOS/SPD serial identifiers are `FC5AF1F5` and `765377F8`. These eight-hex-digit IDs may differ from a longer serial/barcode printed on each DIMM label.
- **OPEN verification concern (RootAtSkic, 2026-08-25):** one reported eight-character identifier contains letters while the other contains only numerals. The fetched window contains no response resolving that concern. Treat both values only as unverified SMBIOS-reported hexadecimal identifiers—not confirmed physical-label serial numbers—until raw records/SPD or DIMM labels are independently checked. Source: Discord message `1541875756626354266`.
- SMBIOS does not expose module manufacturing week/year. Determining production date requires readable SPD EEPROM data or inspection of the physical DIMM labels; neither was available from the container during this check.
### Source anchor
- Latest processed human message: `1542256118284157158` (2026-08-26 19:35 UTC).
## TrueNAS Stability Incident — 2026-08-17
- TrueNAS `26.0.0-BETA.2` on `nas.lego-cloud.eu` underwent full-host boots at 11:02:52, 12:16:18, 12:32:13, and 13:52:25 Europe/Vilnius time. Docker/containerd restarts were consequences, not the initial cause.
- Recursive recovery of `/var/lib/systemd/pstore` proved that all observed Aug 17 outages followed explicit kernel panics. Fifteen panic records are persisted from Aug 4–17, all using kernel `6.18.23-production+truenas` and OpenZFS `2.4.1-1`.
- A repeated signature across earlier boots faults on the same impossible address `0x04000034` in `__mutex_lock`, called through `dbuf_find` / `dbuf_hold_impl` in ZFS. The latest panic occurred at 16:01:52 Europe/Vilnius and rebooted at about 16:03:50: process `rg` under UID 10000 (verified as the Hermes user) triggered an OverlayFS xattr read that faulted on non-canonical pointer `0x7fff8a3979c65828` in `nvt_lookup_name_type -> nvlist_lookup_common -> zpl_xattr_get_sa [zfs]`. A prior panic also involved Hermes `rg` and the same ZFS xattr/nvlist area. Hermes should avoid `search_files`/ripgrep on this TrueNAS-hosted container until the kernel/OpenZFS issue is fixed, using bounded `read_file` and Python traversal instead. Other panics involve ZFS ARC/dbuf/list corruption and kernel page accounting. This makes a TrueNAS 26/OpenZFS/kernel regression or use-after-free a strong primary candidate; broad userspace segfaults leave CPU/RAM corruption as a material secondary possibility.
- Current pool and Apps recovered healthy, with no storage or memory-capacity pressure. BMC reports no current power, fan, or overload fault; historical `last_power_event` is `ac failed` without a timestamp. EDAC counters are 0 corrected / 0 uncorrected on both controllers, and no MCE/APEI event appears in panic captures.
- CPU is an Intel i9-14900KS running microcode `0x12F` (current Intel mitigation level). Stable boot environment `25.10.3` remains available; `kdump_enabled` is false.
- Best discriminator is booting `25.10.3`: if panics stop under the same workload, the beta kernel/OpenZFS stack is implicated; if they continue, proceed to offline RAM/CPU testing and conservative BIOS defaults. Preserve pstore files for a TrueNAS bug report.
- TrueNAS support feedback requested hardware elimination before SA-xattr investigation: several full MemTest86+ passes, verify latest board BIOS and Intel-default CPU power/voltage settings, test known-good ECC UDIMMs if possible, and attach all 15 pstore records. A 15-record archive was prepared at `/opt/data/tmp/truenas-pstore-15-panics-2026-08-17.tar.gz` (SHA-256 `ff82f38c265931f6132ea377d4f14aa25741d32fcb3e2a268148b1614b817b81`). The files contain the private hostname/pool name, so attach them to the support ticket rather than publishing indiscriminately.
- ECC-memory investigation: target is the board maximum, **192 GB ECC = 4 × 48 GB DDR5 ECC UDIMM**. ASRock Rack officially supports 288-pin DDR5 1.1 V ECC/non-ECC **UDIMMs only** on W680D4U-2L2T/G5, maximum 48 GB/DIMM with 14th-gen Core. Four 48 GB 2Rx8 DIMMs are 2DPC dual-rank and should operate at the board-supported 3600 MT/s; 5600 is the DIMM rating, not expected installed speed. Intel officially lists ECC support for i9-14900KS. The exact 48 GB QVL module is SMART `SR6G7UD5385MB01`; specification-compatible but not individually QVL-listed alternatives include Kingston `KSM56E46BD8KM-48HA`, Micron `MTC20C2085S1EC56BR`, and Samsung `M323R6GA3BB0-CWM`. Buy four identical modules from one revision/lot; never RDIMM/LRDIMM. No Lithuanian retailer or EU distributor page checked on 2026-08-17 verified four exact 48 GB modules in stock, so the QVL SMART part likely needs special ordering. For lower capacities, exact QVL SKUs include Kingston `KSM48E40BD8KM-32HM`/`KSM48E40BS8KM-16HM`, Micron `MTC20C2085S1EC48BA1`, and Samsung `M324R4GA3BB0-CQK0L`; do not confuse Samsung `M323R4GA3BB0-CQK0L` (QVL non-ECC) with ECC.