51 lines
8.7 KiB
Markdown
51 lines
8.7 KiB
Markdown
# Channel Wiki: #jarvis-jr-truenas
|
||
_Channel ID: 1519072621130547220_
|
||
_Last sync: 2026-08-26 03:03 UTC_
|
||
|
||
## Purpose
|
||
TrueNAS, self-hosted infrastructure, registry, proxy, and platform setup discussions.
|
||
|
||
## Key Context
|
||
- Keycloak hostname: `keycloak.lego-cloud.eu`; historical discussion referenced the `master` realm.
|
||
- Gitea organization URL: https://gitea.lego-cloud.eu/home-v1
|
||
- Repository referenced for shared access: https://gitea.lego-cloud.eu/home-v1/arda-v1-code-agent
|
||
- Historical work investigated Harbor installation on TrueNAS SCALE using a custom-app YAML.
|
||
- The external Nginx reverse proxy already existed in front of TrueNAS; Harbor planning should not add an unnecessary duplicate Nginx layer.
|
||
- Services exposed through the existing HTTPS path still require explicit internal/service ports and routing.
|
||
|
||
## Behavioral Decisions
|
||
- Teach and train without negativity or pointing out what others lack.
|
||
- Operate autonomously enough that RootAtSkic does not need to micromanage routine investigation.
|
||
|
||
## Active Topics
|
||
- TrueNAS application deployment patterns.
|
||
- Harbor/container-registry architecture and certificate/routing requirements.
|
||
- **CONFIRMED capacity objective (RootAtSkic, 2026-08-17):** maximize this TrueNAS host's memory. For the documented W680D4U-2L2T/G5 / i9-14900KS platform, the target is `192 GB ECC` (`4 × 48 GB DDR5 ECC UDIMM`), not a smaller diagnostic kit. Procurement and validation constraints are recorded below.
|
||
- **OPEN investigation request (RootAtSkic, 2026-08-18):** review the attached Hermes container startup log and identify improvements, but **do not change anything yet**. No response or approved implementation was present in the fetched history; RootAtSkic followed up twice asking whether Hermes was alive. Source attachment: Discord message `1539151473189978142`, `message.txt`.
|
||
|
||
## Hermes startup-log investigation — 2026-08-18 (read-only)
|
||
- **CONFIRMED log observations:** container setup exited successfully and the supervised gateway started, but setup spent `424.296399` seconds between “Changing hermes UID to 568” and “Changing hermes GID to 568” (`425.036112` seconds from the UID step through setup completion). The cause is **UNKNOWN** from this log alone; recursive ownership work is a candidate to investigate, not a confirmed diagnosis.
|
||
- The lifecycle ledger reported the previous gateway life as an unclean exit consistent with `SIGKILL / OOM / VM death`, while recording `suspected_oom=False`, about 297,604 KiB RSS, abundant available memory and no swap use at its last heartbeat. The exact exit cause remains **UNKNOWN**.
|
||
- Startup marked two interrupted cron executions `unknown` after restart. It also logged repeated operational friction: `hermes: command not found` in tool shells, `discord.server_actions` configured with unsupported action `all`, unavailable web-dependent tools because `check_web_api_key` returned false, Git commands launched outside a repository, and one command held for security approval.
|
||
- **CONFIRMED security warning:** the API server was bound to `0.0.0.0` while `terminal.backend` was `local`/unsandboxed, meaning trusted-network firewalling or a sandboxed terminal backend should be evaluated. This is an investigation finding only; RootAtSkic explicitly prohibited configuration changes at this stage.
|
||
|
||
## Live DIMM inventory — 2026-08-25 (read-only)
|
||
- Host SMBIOS Type 17 reports **96 GiB installed** as **2 × 48 GiB G.Skill DDR5 UDIMMs**, part `F5-5200J4040A48G`, dual-rank. The DIMMs advertise 4800 MT/s in SMBIOS and are configured at 4400 MT/s.
|
||
- The board manual identifies the physical channels/slots as `DDR5_A1`, `DDR5_A2`, `DDR5_B1`, and `DDR5_B2`; its priority-two-DIMM configuration is **`DDR5_A2` + `DDR5_B2`**. This matches RootAtSkic's recollection of one DIMM in channel A and one in channel B.
|
||
- Firmware exposes those two occupied records under Intel internal locator names `Controller0-ChannelA-DIMM1` and `Controller1-ChannelA-DIMM1`; these strings must not be presented as the motherboard's physical channel labels. Their SMBIOS/SPD serial identifiers are `FC5AF1F5` and `765377F8`. These eight-hex-digit IDs may differ from a longer serial/barcode printed on each DIMM label.
|
||
- **OPEN verification concern (RootAtSkic, 2026-08-25):** one reported eight-character identifier contains letters while the other contains only numerals. The fetched window contains no response resolving that concern. Treat both values only as unverified SMBIOS-reported hexadecimal identifiers—not confirmed physical-label serial numbers—until raw records/SPD or DIMM labels are independently checked. Source: Discord message `1541875756626354266`.
|
||
- SMBIOS does not expose module manufacturing week/year. Determining production date requires readable SPD EEPROM data or inspection of the physical DIMM labels; neither was available from the container during this check.
|
||
|
||
### Source anchor
|
||
- Latest processed human message: `1541875756626354266` (2026-08-25 18:23 UTC).
|
||
|
||
## TrueNAS Stability Incident — 2026-08-17
|
||
- TrueNAS `26.0.0-BETA.2` on `nas.lego-cloud.eu` underwent full-host boots at 11:02:52, 12:16:18, 12:32:13, and 13:52:25 Europe/Vilnius time. Docker/containerd restarts were consequences, not the initial cause.
|
||
- Recursive recovery of `/var/lib/systemd/pstore` proved that all observed Aug 17 outages followed explicit kernel panics. Fifteen panic records are persisted from Aug 4–17, all using kernel `6.18.23-production+truenas` and OpenZFS `2.4.1-1`.
|
||
- A repeated signature across earlier boots faults on the same impossible address `0x04000034` in `__mutex_lock`, called through `dbuf_find` / `dbuf_hold_impl` in ZFS. The latest panic occurred at 16:01:52 Europe/Vilnius and rebooted at about 16:03:50: process `rg` under UID 10000 (verified as the Hermes user) triggered an OverlayFS xattr read that faulted on non-canonical pointer `0x7fff8a3979c65828` in `nvt_lookup_name_type -> nvlist_lookup_common -> zpl_xattr_get_sa [zfs]`. A prior panic also involved Hermes `rg` and the same ZFS xattr/nvlist area. Hermes should avoid `search_files`/ripgrep on this TrueNAS-hosted container until the kernel/OpenZFS issue is fixed, using bounded `read_file` and Python traversal instead. Other panics involve ZFS ARC/dbuf/list corruption and kernel page accounting. This makes a TrueNAS 26/OpenZFS/kernel regression or use-after-free a strong primary candidate; broad userspace segfaults leave CPU/RAM corruption as a material secondary possibility.
|
||
- Current pool and Apps recovered healthy, with no storage or memory-capacity pressure. BMC reports no current power, fan, or overload fault; historical `last_power_event` is `ac failed` without a timestamp. EDAC counters are 0 corrected / 0 uncorrected on both controllers, and no MCE/APEI event appears in panic captures.
|
||
- CPU is an Intel i9-14900KS running microcode `0x12F` (current Intel mitigation level). Stable boot environment `25.10.3` remains available; `kdump_enabled` is false.
|
||
- Best discriminator is booting `25.10.3`: if panics stop under the same workload, the beta kernel/OpenZFS stack is implicated; if they continue, proceed to offline RAM/CPU testing and conservative BIOS defaults. Preserve pstore files for a TrueNAS bug report.
|
||
- TrueNAS support feedback requested hardware elimination before SA-xattr investigation: several full MemTest86+ passes, verify latest board BIOS and Intel-default CPU power/voltage settings, test known-good ECC UDIMMs if possible, and attach all 15 pstore records. A 15-record archive was prepared at `/opt/data/tmp/truenas-pstore-15-panics-2026-08-17.tar.gz` (SHA-256 `ff82f38c265931f6132ea377d4f14aa25741d32fcb3e2a268148b1614b817b81`). The files contain the private hostname/pool name, so attach them to the support ticket rather than publishing indiscriminately.
|
||
- ECC-memory investigation: target is the board maximum, **192 GB ECC = 4 × 48 GB DDR5 ECC UDIMM**. ASRock Rack officially supports 288-pin DDR5 1.1 V ECC/non-ECC **UDIMMs only** on W680D4U-2L2T/G5, maximum 48 GB/DIMM with 14th-gen Core. Four 48 GB 2Rx8 DIMMs are 2DPC dual-rank and should operate at the board-supported 3600 MT/s; 5600 is the DIMM rating, not expected installed speed. Intel officially lists ECC support for i9-14900KS. The exact 48 GB QVL module is SMART `SR6G7UD5385MB01`; specification-compatible but not individually QVL-listed alternatives include Kingston `KSM56E46BD8KM-48HA`, Micron `MTC20C2085S1EC56BR`, and Samsung `M323R6GA3BB0-CWM`. Buy four identical modules from one revision/lot; never RDIMM/LRDIMM. No Lithuanian retailer or EU distributor page checked on 2026-08-17 verified four exact 48 GB modules in stock, so the QVL SMART part likely needs special ordering. For lower capacities, exact QVL SKUs include Kingston `KSM48E40BD8KM-32HM`/`KSM48E40BS8KM-16HM`, Micron `MTC20C2085S1EC48BA1`, and Samsung `M324R4GA3BB0-CQK0L`; do not confuse Samsung `M323R4GA3BB0-CQK0L` (QVL non-ECC) with ECC.
|