lint / yaml (push) Successful in 18s
lint / terraform (pull_request) Successful in 35s
lint / terraform (push) Successful in 30s
lint / yaml (pull_request) Successful in 17s
lint / ansible (push) Failing after 19m37s
lint / ansible (pull_request) Failing after 19m55s
748 lines
59 KiB
Markdown
748 lines
59 KiB
Markdown
# Changelog
|
||
|
||
What changed in this homelab, when, and why. Newest first.
|
||
|
||
**Conventions**
|
||
- One `##` section per date. No version numbers — services here are deployed
|
||
continuously and independently, so there is nothing to tag.
|
||
- Lead with a one-line summary, then a table of what changed by area.
|
||
- **Incidents get their own subsection**, including anything self-inflicted. The
|
||
post-mortem is the point: what broke, why, and whether it was latent.
|
||
- `Carried forward` lists known gaps left open on purpose, so they do not get
|
||
silently forgotten.
|
||
- Agent-facing traps and procedures do **not** belong here — they go in
|
||
`CLAUDE.md` under *Working rules* / *Environment constraints*.
|
||
|
||
---
|
||
|
||
## 2026-09-10
|
||
|
||
**Flux 的 deployment、drift repair 和 scoped prune 闭环验证完成。**
|
||
|
||
| area | change |
|
||
|---|---|
|
||
| GitOps | PR #18 合并后,Flux 自行发现 revision `f257a2a` 并删除已在 `prune: true` 下重新进入 inventory 的测试 ConfigMap;未发送 reconcile annotation,`http-echo` Deployment/Service 保持 Ready,root 继续 `prune: false` |
|
||
| Helm migration | 选定 `gitea-actions` 作为第一个 Flux HelmRelease adoption:它不承载 Git、入口、DNS、证书、数据库或 secrets controller。live StatefulSet 与 Git 都使用 regular DinD,但 Helm 保存的 release values/manifest 仍是失败的 rootless 配置;接管先固定 chart `0.1.1` 并验证 live Pod spec 不变,升级另开 PR |
|
||
| Helm adoption stage | 为 `gitea-actions` 加入固定 chart `0.1.1` 的 HelmRepository、values ConfigMap 和 `suspend: true` HelmRelease;第一阶段只让 Flux 登记对象,确认 source 与固定 chart render 后再解除 suspend,失败策略使用 `RetryOnFailure` 以避免回滚到 stored rootless manifest |
|
||
| Helm adoption activate | 第一阶段合并后 Flux source 与子 Kustomization 均 Ready,Helm release 仍为 revision 1,runner Pod 未 rollout;再次确认固定 chart 的完整 render 对 live 集群为零差异后,第二阶段移除 `suspend`,允许 Flux 修正 Helm 存储状态并开始 drift detection |
|
||
| Gitea adoption stage | 开始用 Flux 接管关键 `gitea` release:固定现有 chart `12.5.3`,以 `suspend: true` 登记 HelmRelease、source、values ConfigMap 和现有 HTTPRoute,子 Kustomization 保持 `prune: false`;现有 OIDC Secret 继续只引用不覆盖,其尚未进入 OpenBao/ESO 的缺口独立跟踪 |
|
||
| Gitea adoption activate | 第一阶段合并后 source、子 Kustomization 和 HTTPRoute 均 Ready,Helm release 仍为 revision 14,Gitea Pod 未重建或重启;再次确认固定 chart 对 live 业务资源零差异后移除 `suspend`,允许 Flux 修正 Helm 存储状态并启用 drift detection |
|
||
| Gitea upgrade plan | 规划两跳升级:chart `12.6.0` + 显式 Gitea `1.26.4`,再到 chart `12.7.0` + 显式 Gitea `1.27.3`;每个 minor 都先以 suspended desired state 合并、停机建立 CNPG/PVC 一致回滚点,再用独立 PR 激活。当前 CNPG 无连续备份、local-path PVC 无 snapshot class,因此禁止无备份直接触发数据库 migration |
|
||
| Gitea 1.26 preparation | 将第一跳目标写入 Git:chart 固定为 `12.6.0`、rootless 镜像显式固定为 `1.26.4`,同时重新设置 HelmRelease `suspend: true`;该准备 revision 合并后只更新 desired state,不触发 Pod replacement 或数据库 migration |
|
||
| Gitea 1.26 backup | 预拉取 `1.26.4-rootless` 后,在 HelmRelease suspended 状态将 Gitea scale 到 0;生成并校验 539355-byte CNPG custom dump、2609188-byte PVC tar 和 2848667-byte GPG encrypted bundle,随后恢复旧版 `1.25.5` 并验证内外 API 与 Flux source。按明确决定不上传 OCI,本阶段接受只有节点本地回滚点的风险 |
|
||
| Gitea 1.26 activation | 停机一致备份门槛完成后,激活变更只移除 HelmRelease 的 `suspend`;chart `12.6.0`、显式 `1.26.4-rootless` image、values、数据库与 PVC 均保持已 review 的准备状态 |
|
||
| Gitea 1.26 result | Flux 以 Helm revision 16 成功完成 chart `12.6.0` / Gitea `1.26.4-rootless` 的 Recreate upgrade 和 migration 323–330;Pod 内/统一域名 API、临时 branch push/delete、Flux source 及 main/smoke 的全部 CI jobs 均通过,Pod 在约 15 分钟采样中保持零重启,Authelia OIDC init 同步与浏览器交互式管理员登录也已确认成功 |
|
||
| Gitea 1.27 preparation | 预拉取 `1.27.3-rootless` 并将第二跳 desired state 原子设置为 chart `12.7.0`、显式 image `1.27.3` 和 `suspend: true`;合并只暂停并登记目标,不执行 migration,激活前必须从当前 1.26.4 数据建立新的配套回滚点 |
|
||
| Gitea 1.27 activation | 按明确决定跳过新的 1.26.4 数据库/PVC 备份,激活变更只移除 HelmRelease 的 `suspend`;接受 migration 失败后不能无损回退到 1.26.4 的风险,现有 1.25.5 本地备份仅能作为会丢失第一跳后状态的灾难恢复点 |
|
||
| CI runner network | 修复 Actions job 容器访问 GitHub 超时:k3s Pod MTU 为 1450,而 DinD 动态 bridge 默认为 1500;为 Docker daemon 固定 `--mtu=1450`。隔离测试证明相同 curl 镜像在默认 bridge 超时、在 MTU 1450 bridge 下访问 GitHub 与 API 均约 0.1 秒成功 |
|
||
|
||
### Incident: Gitea 备份后的恢复命令被 stdin 校验阻塞
|
||
|
||
最初把本地 custom-format dump 通过 `kubectl exec -i` 输送给 CNPG Pod 内的
|
||
`pg_restore --list`;远端 stdin 没有正常结束,组合脚本因此停在校验步骤,尚未执行
|
||
后面的 scale-up。Deployment 保持预期的 0,没有失败 Pod 或数据写入。发现后终止
|
||
会话、先恢复 Gitea,再把 dump 临时复制到 CNPG 可写数据卷完成校验并立即删除。
|
||
旧版 Gitea 恢复后内外 API 和 Flux source 均正常;后续 runbook 不再把 stdin 管道与
|
||
恢复命令放进同一个 shell transaction。
|
||
|
||
`Carried forward`: complete the two-stage zero-change `gitea` HelmRelease
|
||
adoption, migrate its remaining manual OIDC Secret to OpenBao/ESO, then upgrade
|
||
Gitea and add credential-free PR plan output before ordering the remaining Helm
|
||
migrations by dependency and blast radius. Root Flux prune remains disabled
|
||
until brownfield ownership is audited.
|
||
|
||
## 2026-09-09
|
||
|
||
**Recorded the brownfield GitOps/IaC redesign before changing live infrastructure.**
|
||
|
||
| area | change |
|
||
|---|---|
|
||
| docs | Added `docs/homelab-gitops-redesign.md` as the durable record of control-plane boundaries, ingress and DNS consolidation, certificates, OCI recovery, CI placement, state recovery and phased adoption. It requires zero-change adoption before mutation and keeps PVE/OpenBao/Git recovery independent of k3s |
|
||
| OCI | Read-only discovery found the likely authoritative lost-root state in `oci-k8s-free-tier-tfstate/terraform.tfstate`: Terraform 1.15.8 state format 4, serial 249, covering the existing VM and public network. No state content or credential was committed; the bucket currently lacks versioning |
|
||
| secrets | Recorded that ESO 2.8.0, five ExternalSecrets and the scoped OpenBao Kubernetes-auth path already exist; the next gate is live recovery testing and migration of any remaining manual Secrets |
|
||
| Terraform | Recorded Gitea 1.27 State Registry as the preferred candidate for local roots after version and recovery testing; the OCI recovery root remains in OCI Object Storage to avoid a home-control-plane dependency loop |
|
||
| cleanup | Removed the retired NapCat tree, the Contour and Kanidm archive trees, and seven generated Terraform plan files before establishing the clean Git baseline; plans may embed complete state and remain globally ignored |
|
||
| CI | Added a review-first Gitea Actions runner bootstrap: official actions chart 0.1.1, pinned runner 2.3.0, one persistent instance-scoped Kubernetes runner with capacity four, plus an ESO reference to its registration token in OpenBao. The first deployment proved that rootlesskit is blocked by the node's AppArmor unprivileged-userns policy; because the chart requires privileged DinD in either mode, the reviewed fix uses regular DinD instead of weakening the host-wide policy. The runner image intentionally carries neither `uv` nor Terraform: Terraform uses its versioned setup action, while `uv` is pinned and installed from official PyPI because the nested job network reaches PyPI but times out against the GitHub API queried by `setup-uv`. Ansible installs only `ansible-core` in its tool venv and puts declared Galaxy collections in a shared path visible to ansible-lint; installing the `ansible` meta-package had made Galaxy falsely skip that shared installation |
|
||
| identity | Declared the Samba AD `gitea-admins` group with `panxiao81` as its initial member. Gitea already maps this OIDC group to site administrators; the local `gitea_admin` account remains as break-glass access |
|
||
| docs | Reconciled the redesign and CI status with reality: the Gitea remote, instance-scoped runner, OpenBao-projected registration token, green Stage 1 and Flux bootstrap are live; off-site mirroring and recovery verification remain pending |
|
||
| 协作规范 | 在 `AGENTS.md` 中明确:homelab 向 `git.ddupan.top` 提交的 commit message、PR、issue 与项目文档默认优先使用中文,同时保留必要的英文技术标识符 |
|
||
| k3s | 在本机逐级从 `v1.33.6+k3s1` 升级到 `v1.33.13+k3s2`、`v1.34.11+k3s1`、`v1.35.8+k3s1`,最终到 `v1.36.4+k3s1`;每一级均建立 SQLite/server 冷备份并验证节点、工作负载、PVC、Gateway、DNS 与 Gitea。k3s 每次重启都会覆盖 CoreDNS 的手工 `serve_stale`,已按 `platform/k3s/Corefile.desired` 恢复 |
|
||
| GitOps | Flux `v2.9.5` 的四个核心 controller 已上线;集群内只读 Gitea source 与 `prune: false` 的 root Kustomization 均在合并 revision `aaa54a1` 上 Ready,完成了首个 pull reconciliation 闭环 |
|
||
| cleanup | 已把 `bao-acme` HTTP-01 solver 改到 Envoy Gateway 的明文 listener,并删除不再承载流量的 Contour namespace、provisioner、RBAC、GatewayClass 和全部 `projectcontour.io` CRD;Envoy Gateway、证书、DNS 与 Gitea 复查正常 |
|
||
| GitOps canary | 加入由 Flux 部署到独立 `gitops-canary` namespace 的 `http-echo` Deployment 和 Service;历史 Contour HTTPRoute 明确排除在 Kustomization 之外,初始保持 `prune: false` |
|
||
| prune 验证 | 为 `http-echo` 加入无业务依赖的 `flux-prune-canary` ConfigMap;先在 `prune: false` 下确认 Flux inventory,后续通过独立 PR 删除并仅为 canary 开启 prune |
|
||
| prune 验证第二阶段 | 第一阶段已确认 `flux-prune-canary` 带 Flux ownership 标签并进入 `http-echo` inventory;从 Git 删除该测试对象,同时仅为 `http-echo` 开启 `prune: true`,root 继续保持 `prune: false` |
|
||
| prune 验证修正 | 第二阶段证明“同一 revision 开启 prune 并删除旧对象”不会回收该对象:Flux 已从 inventory 移除它,但 live ConfigMap 保留。将 ConfigMap 在已经生效的 `prune: true` 下重新纳管,下一 revision 只做删除 |
|
||
| prune 验证最终阶段 | 已确认 `prune: true` 生效且测试 ConfigMap 重新进入 Flux inventory;本次只从 Git 删除该对象,不修改 canary 或 root 的 prune 设置,用于完成精确垃圾回收验证 |
|
||
|
||
### Incident: prune 启用与对象删除放在同一 revision
|
||
|
||
测试把 `http-echo` 从 `prune: false` 改为 `true` 的同时从 Git 删除测试
|
||
ConfigMap。Flux 按新 revision 更新了 inventory,但没有删除按旧设置管理的 live
|
||
对象,导致 ConfigMap 成为 inventory 之外的残留。没有业务影响。修正方式是在
|
||
`prune: true` 已经生效后先重新纳管对象,再用下一 revision 单独删除。
|
||
|
||
`Carried forward`: re-verify OpenBao/ESO recovery and remaining Secret inventory;
|
||
configure an off-site Git mirror; plan the Gitea upgrade beyond 1.25.5;
|
||
take an encrypted independent OCI state copy before enabling bucket versioning;
|
||
reconstruct the missing root to a zero-change plan; add a low-risk Flux canary workload;
|
||
then move Tunnel origins to Envoy one hostname at a time.
|
||
|
||
## 2026-08-15
|
||
|
||
**`retro-pdc` (NT4) has a floppy drive again — Proxmox does not offer one, so it
|
||
comes in through `args`.**
|
||
|
||
| area | change |
|
||
|---|---|
|
||
| proxmox | VM 102 gained `args: -drive if=floppy,format=raw,file=/mnt/pve/laptop/template/iso/retro-pdc-fda.img` and a blank FAT12 1.44 MB image to go with it. PVE exposes no floppy in the UI *or* the config schema, and it launches QEMU with `-nodefaults`, so before this the guest had no A: at all — `info block` listed only `drive-ide0`/`drive-ide2`. Verified after a power cycle: `floppy0 … retro-pdc-fda.img (raw)`. Swapping the medium works live (`eject floppy0` / `change floppy0 <path>` over `qm monitor`) — the block node id changed, so it really re-inserted; only an `args` edit needs the VM stopped and started |
|
||
| proxmox | **`roles/pve_floppy`** — a ~250-line stdlib-only web UI for the thing PVE has no UI for: attach/detach a floppy drive (edits `args`) and swap the medium live (`change floppy0` over the monitor). Its own play in `site.yml`, on **pve1 only** — the baseline play is `serial: 1`, which would have put three copies of the same UI on the LAN. One instance is enough because `pvesh` proxies to whichever node owns the VM, verified pve1→pve3 for `config`, `monitor` *and* the root-only `--args` write. It shells out as root instead of using an API token because it has to: PVE gates `args` on a literal `$authuser eq 'root@pam'` (`PVE::API2::Qemu`, the `# catches args, lock, etc.` branch) and a token's authuser is `root@pam!name`. No stop/start buttons — the PVE UI has those, and this app has no authentication |
|
||
| proxmox | The floppy UI got **authentication, without one line of PAM or LDAP code**: it posts the login form to PVE's own `/access/ticket`, so it accepts every realm the cluster has — `pam` for node-local accounts and `ad` for Samba AD over verified LDAPS (`pve_auth`) — and holds no bind DN, no bind password and no realm config of its own. It talks to **pvedaemon on `127.0.0.1:85`** rather than `pvesh create /access/ticket` so the password never appears in a process argv, and rather than pveproxy:8006 so there is no TLS-to-self dance. Authentication alone is not enough: a session also needs **`Sys.Modify` on `/`** in PVE's ACL (i.e. `pve-admins-ad`), because pointing a VM at an arbitrary host file is a host-level act, not a VM-level one. Sessions are HMAC-signed cookies (HttpOnly, SameSite=Strict, Secure), 8 h, with the signing key generated per process — a restart logs everyone out, deliberately. Serves TLS with the node's own ACME cert and **refuses to start without one** (`--insecure` to override): a login form on cleartext HTTP is worse than no service. Failed logins say only *"login failed"*, so the app cannot be used to enumerate AD |
|
||
| proxmox | The procedure is in `CLAUDE.md`. **Not** codified in `pve_vm`: VM 102 is not in `pve_vms` yet, and that gap is already tracked with the rest of the retro-pdc hardware pinning |
|
||
|
||
Two things learned power-cycling it. **NT4 ignores ACPI**: `qm shutdown 102 --timeout 60`
|
||
returned *"VM quit/powerdown failed - got timeout"* and the guest never moved, so it has to be
|
||
shut down from inside (`sendkey ctrl-esc`, `sendkey u`, `sendkey s`, `sendkey ret` over
|
||
`qm monitor`) and then `qm stop`ped once it shows *现在可以关掉电源*. And the shutdown dialog
|
||
defaults to **restart**, not power off — a guest restart keeps the same QEMU process, so the new
|
||
`args` would never have taken effect.
|
||
|
||
`Carried forward`: the guest side is unverified — NT4 came back to the Ctrl+Alt+Del
|
||
login screen and nothing here has its password. `A:` should hold a `README.TXT`
|
||
written to the image from the host.
|
||
|
||
The floppy UI **is deployed** on pve1 and serves TLS on 8088; the realm list it renders
|
||
comes back from the cluster (`ad`, `pam`, `pve`) and a wrong password gets *"login
|
||
failed"* through a real pvedaemon round trip. What is still unverified is everything
|
||
past a *successful* login — the VM table and the four actions have only ever run against
|
||
a stubbed `pvesh`, because no password for any realm exists in the session that built it.
|
||
No DNS record was needed: `pve1.ad.ddupan.top` already resolves, which is also why the
|
||
node's own ACME cert is the right one to serve.
|
||
|
||
### A modem emulator, so retro guests can dial an ISP
|
||
|
||
Retro guests expect dial-up, and there is no PSTN here. Three options were weighed before
|
||
writing anything:
|
||
|
||
| option | verdict |
|
||
|---|---|
|
||
| 86Box's built-in Hayes modem (since 4.2, and retrolab runs 6.0) | **Use it for 86Box guests.** Adapter `[COM] Standard Hayes-compliant Modem`, phonebook file maps a dialled number to `host:port`, non-zero listening port makes it answer. ⚠ Turn **Telnet emulation off** — PPP frames start `FF 03` and telnet IAC eats them. Its built-in internet mode (dial `0.0.0.0`) is **SLIP, not PPP**, so it needs the guest hacked into SLIP; not what a period PC did |
|
||
| `tcpser` for the QEMU guests | **Rejected.** Not packaged past bionic (source build), and its only socket DTE transport is **ip232**, which is not 8-bit transparent: `ip232_write` doubles every `0xFF`, `ip232_read` steals `FF 00`/`FF 01` for DTR, and the modem injects `FF <flags>` for DCD/RI. Read the source rather than assuming — `-serial tcp:` to it would corrupt every PPP frame and every ZMODEM block. Its `-p` port is the *phone line* side, not the serial side, so QEMU's telnet chardev cannot be pointed at it either |
|
||
| `roles/retro_modem/files/atmodem.py` | **Written.** ~420 lines, stdlib asyncio, no deps. Listens on a **unix socket** (PVE's `qm set -serial0 socket` plugs straight in) or TCP, so no socat + pty sandwich. `ATD` either opens a TCP connection or hands the raw line to **pppd** — an ISP terminal server, which is what the machines are actually dialling. `--line <port>` gives it a phone number: an inbound TCP connection rings the guest, which answers with `ATA` or automatically once it has set `S0`. That is the half **NT4's RAS needs to receive calls**, and it is also how one retro guest dials another |
|
||
|
||
Two non-obvious bits, both commented in place. Extended commands are `&X`/`%X`, so
|
||
searching for `D` to find the dial command fires on the `&D2` in every dialer's init
|
||
string and dials "2" — consume the prefix first. And unknown commands answer `OK` on
|
||
purpose: that is what makes an emulated modem work with dialers nobody has tested against.
|
||
|
||
`S0` auto-answer needed the DTE read to be interruptible: a guest that has set `S0` sends
|
||
*nothing* while waiting for a call, so the modem sits blocked reading the serial port and
|
||
could never decide to pick up on its own. The read now races an answer event.
|
||
|
||
`--selftest` is the acceptance test. It checks what ip232 gets wrong — dial through the
|
||
phonebook, round-trip all 256 byte values unchanged, escape with `+++` — then takes an
|
||
inbound call both ways, by `ATA` and by `S0`. It caught a real bug: when the far end
|
||
dropped, the outbound pump kept reading the serial port forever, so after `NO CARRIER` the
|
||
modem never returned to command mode and silently swallowed every later command. Both
|
||
directions now end the call.
|
||
|
||
**Busy signal, and the direction bug it flushed out.** A second caller now gets refused by
|
||
the kernel at `connect()`, because an engaged modem *closes its listening socket* until it
|
||
hangs up — a ringing line counts as engaged too. Accepting and then closing would have been
|
||
worse than nothing: the caller's modem would report `CONNECT` and immediately `NO CARRIER`,
|
||
which is a phantom call, not a busy tone. The dialling side maps `ConnectionRefusedError`
|
||
to **`BUSY`** and everything else to `NO CARRIER` (order matters — it subclasses `OSError`).
|
||
|
||
Writing that exposed a wrong assumption in the usage: **PVE's `-serial0 socket` leaves QEMU
|
||
listening**, so atmodem has to dial *into* the VM, not wait for it. Added `--connect`
|
||
alongside `--listen`; 86Box and plain TCP still want the listening side. Verified against a
|
||
stand-in listener — the guest end sees `OK` come back.
|
||
|
||
One process is one modem on one line, deliberately: that is what a modem is. An ISP's T1
|
||
into a rack of them is N processes on N ports. A single number in front of the rack is a
|
||
hunt group, i.e. a dispatcher, and that is **not** built.
|
||
|
||
**Existing AT libraries were checked and rejected**, so this does not get re-litigated:
|
||
almost everything on PyPI (`attila`, `python-gsmmodem`, `modem-cmd`, `esp_modem`) is the
|
||
**DTE** side — it *sends* AT to a real modem. `AT-Command-Emulator` is DCE but GSM
|
||
(`AT+CMGS`, SMS), so no dialling and no data mode. The one real match,
|
||
[`tcpatmodem`](https://github.com/stblassitude/tcpatmodem) (MIT, PyPI), is a DCE-side
|
||
interpreter — but its DTE side is **stdin/stdout only**, it cannot answer (`RING` appears
|
||
solely as an entry in its result-code table; no `bind`/`listen`/`accept` in the source),
|
||
and it was last touched in January 2019. Nor can its interpreter be lifted out on its own:
|
||
its dispatch table has no `&`, `%` or `\` entry and falls through to `ERROR`, with `a` and
|
||
`z` wired to `command_error` outright — so `AT&F` and `ATZ`, the first things every dialer
|
||
sends, both fail.
|
||
|
||
The deeper reason nothing is reusable: **commands and responses are different grammars.**
|
||
A command line is a run of concatenated commands with no separator (`AT&F&C1&D2S0=0X4E1V1`)
|
||
that cannot be tokenised without knowing the command set, and `D` swallows the rest of the
|
||
line; a response is line-oriented `\r\n<verb>\r\n` with `+CMD: <params>`. DTE libraries
|
||
parse the second, because a DTE never reads the first — it writes it. The one shared piece
|
||
is the `+CMD=<params>` parameter grammar, which is the part retro dial-up does not use. Adopting it would mean a dependency, a
|
||
socket↔stdio bridge, and a fork for the answering half, to replace a ~90-line parser. The
|
||
AT parsing is the commodity part; the socket DTE, the answering and the pppd hand-off are
|
||
not, and nothing off the shelf has them. Complete Hayes DCE implementations do exist —
|
||
DOSBox-X `serialmodem.cpp`, 86Box `net_modem.c`, tcpser `modem_core.c` — but all are C
|
||
welded into their host emulator. Read them if a dialer misbehaves.
|
||
|
||
**Verified end to end against a real `pppd`, not just the self-test.** A client `pppd`
|
||
dialled through atmodem into the `pppd` atmodem spawned for the `ppp` phonebook target —
|
||
i.e. the whole ISP chain, with no retro guest involved. From syslog:
|
||
|
||
```
|
||
send (ATDT5551212^M) / expect (CONNECT) / ATDT5551212^M^M / CONNECT / -- got it
|
||
Serial connection established. Using interface ppp0 / ppp1
|
||
PAP peer authentication succeeded for retro Remote message: Login ok
|
||
local IP address 10.62.0.1 remote IP address 10.62.0.2
|
||
```
|
||
|
||
Both ends came up (`ppp0` 10.62.0.1 ↔ `ppp1` 10.62.0.2), 11 frames and ~400 bytes each
|
||
way. That is LCP, PAP, IPCP and IPv6CP all negotiating across the emulated modem with
|
||
`asyncmap 0` in `/etc/ppp/options`, which is the strongest 8-bit-cleanliness proof
|
||
available — and it exercises the `ppp` target and the asyncio-subprocess pipe path, which
|
||
nothing had run before. Note `pppd notty` forks its own *charshunt* onto a pty internally
|
||
(`Connect: ppp0 <--> /dev/pts/6`); our pipes feed that fine. The test appended one line to
|
||
`/etc/ppp/pap-secrets` and restored the file from backup afterwards, verified identical.
|
||
|
||
**Then a real dialer found a real bug: `ATDT;`.** Driven against **VM 103 (Win98 SE)** on
|
||
pve3, whose `serial0: socket` atmodem attached to directly. Win98's Standard Modem opens
|
||
every call like this:
|
||
|
||
```
|
||
DTE> ATZ DCE< OK
|
||
DTE> ATE0V1&C1&D2S0=0 DCE< OK <- the init string the design was betting on
|
||
DTE> ATM1X4 DCE< OK <- unknown commands, answered OK not ERROR
|
||
DTE> ATDT; DCE< OK <- was NO CARRIER; that killed every call
|
||
DTE> ATDT5551212 DCE< CONNECT <- Win98 only sends digits after that OK
|
||
```
|
||
|
||
A trailing `;` means *"dial, then return to command state"*, and Windows TAPI **dials in
|
||
stages** — a bare `ATDT;` first, digits second. Answering `NO CARRIER` to the opener made
|
||
Win98 give up before it ever sent a number, which looked exactly like a phonebook miss and
|
||
was not one. `dial()` now accumulates staged digits and answers `OK`; the self-test replays
|
||
the whole Win98 sequence verbatim so it cannot regress.
|
||
|
||
The two design calls that looked arbitrary are the two that carried it: consuming `&`/`%`
|
||
prefixes before looking for `D` (or `&D2` dials "2"), and answering `OK` to unknown
|
||
commands (or `M1X4` aborts the dial). tcpatmodem would have failed on line 2.
|
||
|
||
**Result: Windows 98 is on the network over an emulated modem.** ISP was `pppd` on the
|
||
laptop, reached over the LAN from pve3, so nothing was installed on the node:
|
||
|
||
```
|
||
call from ('192.168.10.9', 57456)
|
||
ppp0 UNKNOWN 10.62.0.1 peer 10.62.0.2/32
|
||
64 bytes from 10.62.0.2: icmp_seq=1 ttl=128 time=12.8 ms (4/4, ttl=128 = Windows)
|
||
```
|
||
|
||
Also learned: Win98 does **not** drive PVE's USB tablet, so QMP `input-send-event` abs
|
||
clicks are accepted by QEMU and ignored by the guest — until the guest installs USB HID.
|
||
And Win98's own dialling properties prepend the location's outside-line digit and country
|
||
code (`0 5551212`, canonical `86-5551212`), which `digits()` cannot match; untick
|
||
**使用区号与拨号属性** in the connection's properties.
|
||
|
||
**And then Win98 dialled NT4.** Two atmodems on pve3 — one on VM 103's `serial0` with the
|
||
phonebook, one on VM 102's `serial0` with `--line 6102` — turn `5551102` into a call from
|
||
the Win98 guest to `retro-pdc`'s Remote Access Server (RETRO001, 1 port, 正在运行):
|
||
|
||
```
|
||
WIN98 DTE> ATDT; DCE< OK
|
||
DTE> ATDT5551102 DCE< CONNECT
|
||
NT4 DTE> ATH / AT / ATE0V1 / AT / ATS0=0 <- RAS initialising the port
|
||
DCE< RING <- our modem rings it
|
||
DTE> ATA <- RAS answers by hand
|
||
DCE< CONNECT
|
||
```
|
||
|
||
That validates the whole answering half (`--line` → `RING` → `ATA`) against a real NT 4.0
|
||
RAS, and settles a design guess: **RAS sets `S0=0` and answers manually on `RING`**, so the
|
||
`ATA` path is the one that carries, and `S0` auto-answer is there for DOS-era software.
|
||
RAS's init sequence is a *third* dialer handled without changes. Everything above the
|
||
modem — PPP and NT4 domain authentication against RETRO — was left to the operator, who
|
||
was at the keyboard by then.
|
||
|
||
**RAS then found the second real bug: `+++ATH` as one write.** NT4 hangs up by sending the
|
||
escape and the command glued together, and `_dte_to_peer` matched only a bare `b"+++"` — so
|
||
the whole string was forwarded to the far end as data and the line could never be dropped.
|
||
A bare `+++` still waits out its trailing guard; `+++` followed by a command is
|
||
unambiguous, so it escapes at once and the remainder goes to the command reader through a
|
||
small pushback buffer. In the logs this showed up as `+++ATH` → `ERROR` (that part is
|
||
correct — in *command* mode real modems error too; the bug was the data-mode path).
|
||
|
||
Reading the rest of that log is a lesson in not blaming the layer you just wrote. RAS
|
||
answered three calls cleanly, each running 45–75 s before **RAS** hung up — a failure above
|
||
the modem (PPP/auth), matching 端口状态 showing 线路未连接 with zero bytes counted. And the
|
||
eight unanswered `RING`s were Win98 hanging up and **redialling in the same second**
|
||
(`ATH` … `ATDT5551102` both at 22:02:55), before RAS had re-armed its port — the period-
|
||
accurate equivalent of redialling before the far end's modem has reset. RAS re-initialises
|
||
with `AT`/`ATZ`/`ATE0V1`/`ATS0=0` when it recovers.
|
||
|
||
**Dialling an address directly was broken, and only asking about it found it.** `D` takes
|
||
an optional **T**one/**P**ulse modifier, which was never stripped — so `ATDT192.168.10.1:23`
|
||
tried to resolve the host `T192.168.10.1` and returned `NO CARRIER`. The phonebook path hid
|
||
it completely, because lookups go through `digits()`. Strip exactly one modifier, never
|
||
`lstrip()`, or `ATDTtelnet.example.com` loses its `t`. Now covered by the self-test.
|
||
|
||
**`telnet:` targets.** A raw TCP pipe is wrong for a real telnetd: dialling `192.168.10.1:23`
|
||
delivered `\xff\xfb\x01\xff\xfb\x03login: ` to the guest — `IAC WILL ECHO, IAC WILL SGA`
|
||
rendered as `ÿû☺ÿû♥` before the prompt — and an un-doubled `0xFF` corrupts any 8-bit
|
||
transfer. A `telnet:host[:port]` phonebook target now wraps the peer in a ~45-line telnet
|
||
client: it swallows IAC sequences, answers `DO ECHO`/`DO SGA` and refuses everything else,
|
||
un-escapes `IAC IAC`, and doubles `0xFF` outbound. Same dial through it now yields exactly
|
||
`\r\nCONNECT\r\nlogin: `.
|
||
|
||
It is **opt-in per entry** for the same reason 86Box's telnet toggle has to be turned off:
|
||
enabling IAC handling on the PPP or guest-to-guest numbers would corrupt them, since there
|
||
`0xFF` is data — `FF 03` starts every PPP frame.
|
||
|
||
**Redesigned into a switchboard, which came out smaller than what it replaced.** Dialling
|
||
a VM cannot be another target type: the answering guest needs a *modem* to hear `RING` and
|
||
reply `ATA`, so wiring a caller straight to its serial socket hands RAS raw bytes and it
|
||
never picks up. Only a process holding both ends can ring one on behalf of the other. So
|
||
one process now owns N lines (`--vm 102:6102 --vm 103`), and `vm:102` is an internal call:
|
||
|
||
| before, 2 guests | after |
|
||
|---|---|
|
||
| 2 processes, 1 TCP port, 2 logs | 1 process, 1 log |
|
||
| a `--line` port allocated per VM | internal routing by vmid |
|
||
| busy = open/close a listener | busy = does that line have a call |
|
||
| hunt group impossible | falls out of the line table |
|
||
|
||
The internal hop is a `socket.socketpair()`, so every path below it — the pump, 8-bit
|
||
cleanliness, `+++`, `NO CARRIER` — is the same validated code that carries an external
|
||
call. Per-line TCP ports stay, because **86Box lives on retrolab**, a different host, and
|
||
has to reach a line over the network. Lines also reattach on their own now: a guest reboot
|
||
takes the chardev peer with it, and a switchboard needing a restart after every VM reboot
|
||
is not a service.
|
||
|
||
**The phonebook is the API — there isn't one.** It hot-reloads on mtime change, so editing
|
||
a number no longer restarts the modem or drops a live call. That single change removes any
|
||
need for a daemon protocol: the file lives on **pmxcfs** (`/etc/pve/retro-phonebook`), so
|
||
`pve_floppy` running on **pve1** edits exactly what the switchboard on **pve3** reads, with
|
||
no IPC, no API and no second service. The UI gained a phonebook textarea rather than
|
||
becoming a new app, so it inherits the PVE ticket auth, the AD realm, TLS on the node cert
|
||
and the `Sys.Modify` gate that were already there. (pmxcfs mtime has 1-second granularity —
|
||
two edits inside one second would be missed. Irrelevant for human or UI edits.)
|
||
|
||
Targets are an **allowlist, not a blocklist**, and that is the security boundary: a
|
||
phonebook entry is something the modem *acts on* — `ppp` and `ssh:` make it spawn a process
|
||
— so a free-form target would be remote command execution wearing a phone number. `exec:`
|
||
was deliberately never added for that reason, and the UI's self-test asserts what it
|
||
*refuses* (`exec:`, shell metacharacters, non-numeric numbers), not just what it accepts.
|
||
|
||
Also added: `ssh:user@host[:port]` (the same subprocess shape as `ppp`, so nearly free).
|
||
|
||
### The floppy UI was taking ~22 seconds a page; now 4 cold, 0 warm
|
||
|
||
Measured before changing anything, which is the whole story: **every `pvesh` costs ~1.9 s**
|
||
(Perl startup plus a cluster round trip), and a monitor query *forwarded to another node* is
|
||
**3.6 s**. The page made one call for the VM list, one per VM for its config, and one per
|
||
running VM for its monitor — 1 + 4 + 4 calls for four guests.
|
||
|
||
| fix | why it works |
|
||
|---|---|
|
||
| Read configs off **pmxcfs** (`/etc/pve/.vmlist`, `/etc/pve/nodes/<node>/qemu-server/<vmid>.conf`) | Exactly the data `pvesh get .../config` returns, already replicated to every node, at file-read speed. Removed 5 of the 9 calls |
|
||
| Ask the monitor **only about VMs that have a floppy** | "What is in the drive" is meaningless for a VM with no drive. Two thirds of the monitor calls were asking anyway |
|
||
| Run the survivors **in parallel** | `ThreadingHTTPServer` already gives each request a thread; fanning out inside it makes N round trips cost about one |
|
||
| **Cache**, TTL 30 s, cleared by every action | At the 5 s I first wrote, every click still missed — the TTL has to be longer than a human's click interval to ever be warm |
|
||
|
||
Result: 14.4 s just for the list+config calls became a file read, and the page went
|
||
**~22 s → 4.04 s cold, 0.000 s warm**, same data. The remaining 4 s is one forwarded
|
||
monitor call and is the floor for `pvesh`; beating it needs the API over HTTP with a
|
||
retained ticket, i.e. server-side session state this app deliberately does not have (its
|
||
sessions are stateless signed cookies). Not worth it for a page that is now instant in use.
|
||
|
||
Config parsing has its own test: snapshots are appended as `[name]` sections after the live
|
||
config so parsing must stop at the first one, and the split is on the **first** colon only
|
||
because an `args` value is full of them.
|
||
|
||
The login page was its own 1.9 s, before any of that: `realms()` ran a `pvesh` on every
|
||
unauthenticated hit to list something that changes when an auth domain is added, i.e.
|
||
never. Cached for the process lifetime — 2.01 s cold, then **0.034 s**.
|
||
|
||
Deployed with `--tags floppy`; a second run reports `changed=0`.
|
||
|
||
#### Incident — one login in eight was silently rejected (latent since the app was written)
|
||
|
||
The deploy failed on `pve_floppy`'s own self-check, which then passed on a rerun. Chasing
|
||
the flake rather than re-running found a real bug in session cookies, not in the test:
|
||
|
||
```python
|
||
return base64.urlsafe_b64encode(msg + b"|" + _mac(msg)).decode() # sign
|
||
msg, sig = raw.rsplit(b"|", 1) # verify
|
||
```
|
||
|
||
The signature is the **raw 32-byte HMAC digest**, and 32 random bytes contain `0x7C` — the
|
||
byte for `|` — about **12 %** of the time (`1 - (255/256)**32`). When they did, `rsplit`
|
||
split the token *inside its own signature*, `compare_digest` failed, and the user was
|
||
bounced back to the login page. Random, unreproducible, and it had been there since the app
|
||
was written — never noticed because, as recorded above, nothing past a *successful* login
|
||
had ever been exercised.
|
||
|
||
Fix: sign with `_mac(msg).hex()`, which cannot contain the separator. The regression test
|
||
does 300 round trips rather than one, because a single round trip passed ~88 % of the time
|
||
and that is exactly how this survived having a test at all. Verified 20/20 self-test runs on
|
||
the node after deploying, and the service is confirmed running the new binary — the failed
|
||
run had installed the file but died before its restart handler, leaving the old code live.
|
||
|
||
Not yet done: DTR-drop hangup, S-registers beyond `S0`, and the `retro_modem` Ansible role
|
||
— the switchboard runs from `/tmp` on pve3 and the `pve_floppy` change is not deployed.
|
||
|
||
### retronet's WINS now points at the NT4 PDC, not the production DC
|
||
|
||
`retro-pdc` came up on its static **10.61.0.5** with the WINS service installed, so
|
||
retronet's DHCP `wins-server` moved from `192.168.10.5` (the Samba DC) to it —
|
||
`vyos_router` defaults, applied and saved, second run `changed=0`. Verified end to end
|
||
rather than assumed: TCP/42 and 139 open, `nmblookup -A` shows the box holding
|
||
`RETRO<1b>` / `RETRO<1d>` / `..__MSBROWSE__.` (it is the domain master browser), a
|
||
recursive WINS query through it resolves `RETRO01<00>` → 10.61.0.5, and kea's generated
|
||
config carries `netbios-name-servers: 10.61.0.5` for the subnet. That removes retronet's
|
||
last dependency on the production DC. Note the NetBIOS domain is **`RETRO`**, host
|
||
**`RETRO01`** — not `RETRONET`, which is only the VNet/shared-network name.
|
||
|
||
⚠ **`option wins-server` is a multi-value node**, so the role's `set` line *added* a second
|
||
server rather than replacing the first — both were live, and both were written to
|
||
`config.boot` by the play's `save: true`. Removed with an explicit `delete`. The role is
|
||
set-lines-only by design and can never remove a stale value; changing any multi node needs
|
||
a one-off delete against the live box. Trap recorded in `CLAUDE.md`.
|
||
|
||
---
|
||
|
||
## 2026-07-28
|
||
|
||
**The repo got git history for the first time; a stage-1 lint gate; and
|
||
`auth.ddupan.top` now resolves on the LAN instead of via Cloudflare.**
|
||
|
||
| area | change |
|
||
|---|---|
|
||
| repo | First commit ever — 490 files, ~11.8 MB. Root `.gitignore` added; ~13 GB of ISOs, WinPE images, HF model blobs, `node_modules` and vendored netboot menus excluded, along with nine secret-bearing files. Later the same day `blocky/logs/` joined them — Blocky writes one query log per day, the first had already been committed, so it is ignored **and** `git rm --cached`d. It stays in the history of `fde9ff2`: every DNS query the LAN made that day, no credentials |
|
||
| ci | Stage-1 lint: `yamllint`, `ansible-lint`, `terraform fmt`/`validate`. Config tuned against a real run — 293 YAML errors down to 0, 300 Ansible findings down to 34 |
|
||
| proxmox | Added `ansible/requirements.yml`. It had never existed, while the other two Ansible projects declared theirs — so `ansible.netcommon`, `vyos.vyos`, `ansible.posix` and `community.general` were undeclared and a fresh checkout could not reproduce the environment |
|
||
| cert-manager | `certificate-auth-ddupan.yaml` — LE cert for `auth.ddupan.top`. The `*.ad.ddupan.top` wildcard cannot cover it: different zone, one label shallower |
|
||
| envoy-gateway | Second HTTPS listener `https-auth`, SNI-selected. The existing `*.ad.ddupan.top` listener is untouched |
|
||
| authelia | `httproute.yaml` routes `auth.ddupan.top` to the `authelia` Service |
|
||
| k3s | CoreDNS answers `auth.ddupan.top` with the gateway (192.168.10.127) and suppresses AAAA |
|
||
| gitea | Actions enabled (`ENABLED=true`, `DEFAULT_ACTIONS_URL=github`). No runner deployed yet, so nothing executes |
|
||
| victoriametrics | 25 whitespace fixes, each verified to parse to an identical document. Two alert-rule files were deliberately **not** fixed — their trailing spaces sit inside `\|` literal block scalars and are part of the alert text |
|
||
| docs | `CHANGELOG.md` (this file) and `docs/cicd.md`, the CI/CD design. `docs/superpowers/` retired — its only work, the 2026-04-18 CloudNativePG migration, shipped and has been healthy 101 days |
|
||
| secrets | All four secret-bearing configs returned to git. **cloudflared**: the credentials Secret and `config.yml` ConfigMap were *dead* — deleted rather than externalised (see below); the tunnel token moved to a gitignored `secret.yaml` with a committed template. **gitea**: DB password moved to a `gitea-db` Secret, injected via `gitea.additionalConfigFromEnvs` as `GITEA__DATABASE__PASSWD`. **litellm-gateway**: compose now interpolates `${POSTGRES_PASSWORD}` from its already-gitignored `.env`. **authelia**: all seven pieces of secret material moved into Secrets via `secret.existingSecret` + `secret.additionalSecrets`, which also took the OIDC signing key out of a plaintext ConfigMap (see below) |
|
||
| external-secrets | **ESO 2.8.0 deployed**, pulling all four Secrets from OpenBao. The operator authenticates with its own ServiceAccount JWT via bao's Kubernetes auth backend, so no credential is stored in the cluster. Policy scoped to `kv/k8s/*` read-only — narrower than the human `admin` policy |
|
||
| openbao | Kubernetes auth backend **enabled and configured for the first time** — it had never existed. Two stale defaults fixed: `openbao_k8s_host` pointed at `192.168.10.10` (nothing listens there), and `openbao_addr` used `127.0.0.1`, which now fails TLS because bao's Let's Encrypt cert has a DNS SAN only |
|
||
| terraform | **All four roots migrated from local state to SeaweedFS S3** (`tfstate` bucket), with native `use_lockfile` locking. Reached over a new LAN route (`s3.ad.ddupan.top`) rather than `obj.ddupan.top`, so state does not depend on the WAN |
|
||
| seaweedfs | S3 identities moved out of `values.yaml` into OpenBao via ESO, plus a least-privilege `terraform` identity scoped to the state bucket |
|
||
| blocky | **Staged, not deployed.** LAN resolver + ad-blocker + split-horizon DNS, as a compose stack on the laptop. Fills a real gap: there is nowhere today to put a LAN record for a `ddupan.top` name — the DC is authoritative only for `ad.ddupan.top`, and the NEC IX has no static-host feature |
|
||
| gitea | LAN route added: cert for `git.ddupan.top`, a third gateway listener (`https-git`) and an HTTPRoute. Serves HTTP 200 in 32ms, so `git push` no longer has to leave the LAN. **DNS not yet switched** — nothing resolves it locally until Blocky or a DC zone lands |
|
||
| smtp-relay | **DKIM signing enabled for `ddupan.top`** — mail relayed via M365 was landing in Junk. The signing config existed but had never been switched on (`Enabled: False`, `Status: CnameMissing`), so outbound mail carried only the tenant's `*.onmicrosoft.com` signature, which does not align with `ddupan.top`. No DNS change was needed |
|
||
| tailscale | **The PVE SDN subnets are now advertised** — the laptop, the tailnet's only subnet router, offered `192.168.10.0/24` and nothing else, so `10.60.0.0/24` (labnet) and `10.61.0.0/24` (retronet) were unreachable from the tailnet even though the laptop has had OSPF routes to both all along. Recorded in `tailscale/subnet-routes.sh` rather than left as shell history, since the whole failure mode is forgetting the step. retronet is included deliberately: quarantining it from the LAN is the point, quarantining it from the tailnet just forces a second VPN. **Both still need approving in the admin console** |
|
||
| retrolab | **86Box mouse capture fixed over RDP.** xorgxrdp's pointer declares relative axes (`REL_X`/`REL_Y`, range `-1..-1`) but posts absolute screen coordinates through them — `xf86PostMotionEvent(dev, TRUE, …)`. 86Box's XInput2 backend trusts the declared mode and fed those absolutes in as movement deltas, pinning the emulated pointer in a corner the moment you clicked to capture. 86Box only ever exempts pointers *by device name* (`TigerVNC pointer`, `Virtual core XTEST pointer`), so `retro_desktop` now renames the xrdp pointer to `TigerVNC pointer` in `/etc/X11/xrdp/xorg.conf`. Also gave the role's channel-read task `check_mode: false`, without which every `--check` run asserted that audio was broken |
|
||
|
||
The hostname deliberately did not change. Issuer, redirect URIs and cookie
|
||
domain all remain `auth.ddupan.top`, so no OIDC client needed re-registering —
|
||
only the network path moved. The public route (Cloudflare → tunnel →
|
||
`cloudflared` → authelia) still works and terminates at the same Service.
|
||
|
||
### Incident — retrolab logins came up with no window manager (self-inflicted)
|
||
|
||
No title bars, no Applications menu, so no way to log out — reported as "I can't
|
||
logout now", and it came back after the move to pve3 because the cause is
|
||
persistent, not transient.
|
||
|
||
`xfce4-session` restores exactly the client list in
|
||
`~/.cache/sessions/xfce4-session-retrolab:10`, and that list had **Count=4:
|
||
xfsettingsd, xfce4-panel, Thunar, xfdesktop** — no `xfwm4`. Once the WM is
|
||
missing from a saved session, every later login is WM-less too.
|
||
|
||
Self-inflicted, and the recovery caused the relapse: xfwm4 had died earlier
|
||
(cause unknown, `.xsession-errors` shows an older `Another compositing manager
|
||
is running on screen 0`), and it was restarted over SSH with `xfwm4 --replace`.
|
||
That process has no `SESSION_MANAGER` in its environment — it logs "Failed to
|
||
connect to session manager" — so the next session save did not record it.
|
||
|
||
Fixed on the host: restarted the WM, set
|
||
`xfconf-query -c xfce4-session -p /general/SaveOnExit -n -t bool -s false`, and
|
||
deleted the stale session file. Left as a documented trap in
|
||
`roles/retro_desktop/tasks/main.yml` rather than a task — see the comment there
|
||
for why automating it costs more than it saves. The distro's failsafe session
|
||
does not help: xrdp never offers the greeter that selects it.
|
||
|
||
### retrolab moved from pve1 to pve3 — 86Box was CPU-starved
|
||
|
||
86Box emulation stuttered and the emulated Sound Blaster glitched. pve1 is an
|
||
**i3-6100U, 2 cores / 4 threads at 2.3 GHz**, and it also carries `vyos-rtr`;
|
||
it was sitting at load 2.0 with 50% CPU. 86Box's recompiler is effectively
|
||
single-threaded, so it wanted clock, not cores.
|
||
|
||
pve3 was the target rather than pve2 for a storage reason, not a CPU one — both
|
||
are **Ryzen 5 PRO 2400GE (4c/8t, 3.2 GHz)** and both idle, but LINSTOR holds an
|
||
**UpToDate replica on pve3 and only a Diskless one on pve2**, so on pve2 every
|
||
block would have crossed the network to another node's disk.
|
||
|
||
`cpu: host` was already set, which is most of the performance win but also
|
||
forced the migration to be **offline**: pve1 is Intel, pve2/pve3 are AMD, and a
|
||
live migration would have handed the running kernel a different feature set.
|
||
Shutdown, `qm migrate` (2 seconds — nothing to copy, shared DRBD), start. The
|
||
guest now reports the Ryzen. vm:101 is not an HA resource, so nothing else
|
||
needed rearranging.
|
||
|
||
Verified after the move, because this was the first time either SDN VNet had to
|
||
leave a node: the DHCP reservation still resolves (`10.60.0.10`, VLAN 100), and
|
||
`br-retro` reaches the VyOS gateway `10.61.0.1` on VLAN 110 — both now crossing
|
||
the physical 1G LAN to reach vyos on pve1 instead of staying inside one host.
|
||
`xrdp` and the laptop's `/mnt/iso` NFS mount came back on their own.
|
||
|
||
Win98 took a hard power-off (ScanDisk on next boot): the guest OS ignored ACPI
|
||
shutdown until its timeout, and the desktop session could not be driven to shut
|
||
the emulator down cleanly first.
|
||
|
||
### 86Box's Win98 guest could not DHCP — the emulated cable was unplugged
|
||
|
||
Symptom: Win98 on retronet got only an APIPA address, `winipcfg` renew failed
|
||
instantly with "DHCP 服务器不存在". Everything downstream of the guest was
|
||
healthy and measured that way: `br-retro` up with `enp6s19` **and** `tap0`
|
||
enslaved and forwarding, an address temporarily added to `br-retro` pinged the
|
||
VyOS gateway `10.61.0.1`, and kea was listening on `10.61.0.1:67` with the
|
||
`RETRONET` pool configured.
|
||
|
||
The tell was `ip -s link show tap0`: **RX 0 packets, ever** — TX counted the
|
||
frames the bridge flooded *toward* the guest, but 86Box had never written a
|
||
single frame *from* it. Not a fabric problem at all.
|
||
|
||
Cause: `net_01_link = 2` in `~/86Box VMs/98/86box.cfg`. In 86Box
|
||
`NET_LINK_DOWN = (1 << 1)`, so the value means the NIC's link is **down** —
|
||
86Box's own "unplug the cable" toggle, reachable by clicking the network icon
|
||
in its status bar. The default is `504` (every speed/duplex bit set); deleting
|
||
the line restores it. The guest driver was fine throughout: it read its MAC
|
||
(`00:E0:4C:CB:7A:58`) off the emulated PROM and bound TCP/IP normally.
|
||
|
||
Fixed by removing the line and restarting the emulator. Win98 now holds
|
||
**10.61.0.107** from the `RETRONET` pool, first lease that segment has ever
|
||
handed out.
|
||
|
||
Consequence, and the reason the VM does not boot unattended: the NIC's boot ROM
|
||
is enabled (`bios = 1` under `[Realtek RTL8029AS #1]`), and with the link up
|
||
Etherboot 5.4.4 now runs a DHCP loop at every boot instead of failing
|
||
instantly. It never accepts kea's reply — the reply is on the wire, addressed
|
||
to the card, and Etherboot still prints `No IP address` — so it retries
|
||
indefinitely and the machine never reaches the hard disk. Press **Q** at
|
||
`Boot from (N)etwork or (Q)uit?` to skip it; set `bios = 0` if PXE on retronet
|
||
is not wanted. Left as-is: enabling that ROM looks deliberate.
|
||
|
||
### Incident — retrolab's desktop stranded again (needrestart, second occurrence)
|
||
|
||
Same failure as 2026-07-25, different trigger. **unattended-upgrades** upgraded
|
||
`libc6` at 06:28:37, and needrestart restarted `xrdp-sesman` at 06:28:55.
|
||
sesman came back with an empty session table and could no longer reattach the
|
||
running `:10` display, so every reconnect started a *new* one — and
|
||
`xfce4-session` refuses to run twice for the same user, so each died in about a
|
||
second (`Window manager (pid 102992, display 11) exited quickly (1 secs)`). The
|
||
desktop and its 86Box Win98 VM kept running, just permanently unreachable.
|
||
|
||
The 2026-07-25 fix was `NEEDRESTART_MODE: l` in `retrolab.yml`, which only ever
|
||
covered *our* playbook runs. Ubuntu's automatic upgrades were never in scope,
|
||
which is why the same thing happened again eight hours before anyone noticed.
|
||
|
||
Recovered by killing `:10` outright (86Box included — no way to save it, the
|
||
session could not be reached to shut it down). Fixed properly with
|
||
`/etc/needrestart/conf.d/50-xrdp.conf` pinning `qr(^xrdp)` to `0`, deployed by
|
||
the `retro_desktop` role. This is the mechanism needrestart already uses for
|
||
`gdm`, `sddm` and `xdm` — xrdp-sesman is the same class of service and simply
|
||
was not on the list. Trade accepted: sesman runs against the old libc until the
|
||
host reboots.
|
||
|
||
Watch for this on any other host that grows a long-lived xrdp session.
|
||
|
||
### Incident — Gitea down ~12 minutes (self-inflicted trigger, latent cause)
|
||
|
||
Enabling Actions required a `helm upgrade`, and the chart's `strategy: Recreate`
|
||
kills the old pod before starting the new one. The new pod never came up.
|
||
|
||
The cause was **not** the config change. `configure-gitea` is an **init**
|
||
container running `gitea admin auth update-oauth`, which *fetches*
|
||
`autoDiscoverUrl` before Gitea will start. That URL was unreachable, so the init
|
||
container exited non-zero and the pod crash-looped. Any restart — node reboot,
|
||
eviction, chart bump — would have done the same. Rolling back would not have
|
||
helped, because the rollback also restarts the pod.
|
||
|
||
Restored by temporarily commenting out the `oauth:` block (the auth source
|
||
stays in Gitea's DB; commenting only stops the init-time sync), then permanently
|
||
by moving `auth.ddupan.top` onto the LAN.
|
||
|
||
### Incident — a dead VPN tunnel masquerading as a bad ISP
|
||
|
||
`[email protected]` reported `active running` and its interface was
|
||
`UP`, but the tunnel was dead — 100% loss to its own gateway, 29,156 dropped TX
|
||
packets. Its **58 split-tunnel routes stayed installed**, blackholing Cloudflare
|
||
(`104.21/16`, `172.67/16`), Fastly (`151.101/16`), Microsoft `13.107.x`, AWS
|
||
CloudFront and Akamai. `github.com` is not in that route set, which is why it
|
||
kept working and made the failure look like selective CDN blocking.
|
||
|
||
This was the real cause of the Gitea outage above, of `pypi.org` being
|
||
unreachable, and — because pods use the host routing table — of the same
|
||
blackhole applying cluster-wide. `pve1` was unaffected throughout, having no
|
||
`tun0`. Fixed by restarting the service.
|
||
|
||
### DKIM: the CNAMEs were right all along
|
||
|
||
The published CNAMEs matched `Selector1CNAME`/`Selector2CNAME` exactly, yet
|
||
`ddupan1.d-v1.dkim.mail.microsoft` was **NXDOMAIN** — which reads as a wrong
|
||
tenant label and sends you hunting for the "real" value. It is not.
|
||
**Microsoft creates the tenant host only when signing is enabled**, so the
|
||
target cannot resolve before `Set-DkimSigningConfig -Enabled $true`. The
|
||
NXDOMAIN was the expected pre-enable state, not a fault. After enabling, the
|
||
zone answered `NOERROR` and both selectors served 2048-bit keys immediately.
|
||
|
||
Corollary: `Status: CnameMissing` on a config that has never been enabled does
|
||
not mean your DNS is wrong. Enable it and re-check before touching DNS.
|
||
|
||
**Verified end-to-end**, headers of a test message received at an external
|
||
Outlook.com account: `dkim=pass (signature was verified) header.d=ddupan.top`,
|
||
`spf=pass`, `dmarc=pass`, `compauth=pass reason=100`.
|
||
|
||
**It still landed in Junk** — `X-MS-Exchange-Organization-SCL: 5`
|
||
(`X-Message-Delivery` decodes to `SCL=6`), `dest:J`, `RF:JunkEmail`. Note the
|
||
split: the tenant-side outbound stamp was `SCL:1`, so the score came from the
|
||
*receiving consumer* filter. Authentication is a precondition for good
|
||
placement, not a guarantee of it — the remainder is reputation (`ddupan.top`
|
||
has no sending history and relays via a shared M365 outbound pool,
|
||
`52.101.228.88`) plus content (the test messages were one-line bodies with
|
||
"test" in the subject, no charset, no `MIME-Version` — a worst case for
|
||
scoring). Nothing further to configure; it needs real traffic, time, and
|
||
"not junk" marks.
|
||
|
||
### Discovered — IPv6 broken host-wide on the laptop (not fixed)
|
||
|
||
`Connect-ExchangeOnline -Device` hung with no output. The cause was not the
|
||
module: **`br0` has no global IPv6 address**, because
|
||
`net.ipv6.conf.all.forwarding=1` (needed for libvirt/k3s) makes the kernel
|
||
default `accept_ra` to `0`, so SLAAC never runs — while NetworkManager still
|
||
installed a v6 default route. The only global v6 address on the box belongs to
|
||
`tun0`, so source selection hands it to routes that egress `br0`. Packets leave
|
||
the LAN wearing the VPN's address and nothing returns; the socket sits in
|
||
`SYN-SENT`.
|
||
|
||
This is **not** the known dead-tunnel trap. The VPN was healthy — its gateway
|
||
pinged, v4 through it worked. `ip route get` says `dev br0` and looks innocent;
|
||
the tell is the **source address**, not the device. Anything that resolves AAAA
|
||
and does not fall back fast hangs the same way — `.NET` does not do Happy
|
||
Eyeballs, which is why `curl` masks the fault entirely.
|
||
|
||
Worked around per-process with `DOTNET_SYSTEM_NET_DISABLEIPV6=1`. Not fixed at
|
||
host level: the fix is `net.ipv6.conf.br0.accept_ra=2`, which changes IPv6
|
||
behaviour for k3s, libvirt and NFS on the lab's single point of failure and
|
||
deserves its own change window.
|
||
|
||
### The Authelia OIDC signing key was in a ConfigMap, not a Secret
|
||
|
||
Externalising `authelia/values.yaml` turned up a live exposure rather than a
|
||
git-hygiene problem. The chart's `files/configuration.oidc.jwk.yaml` branches on
|
||
how the key is supplied: `key.path` reads it from a mounted file at runtime,
|
||
but **`key.value` inlines it directly into the ConfigMap**. values.yaml used
|
||
`value:`, so the RSA key that signs every ID token for `auth.ddupan.top` was
|
||
sitting in plaintext in a ConfigMap — readable by anything with `get configmap`
|
||
in that namespace, and not encrypted at rest the way a Secret can be.
|
||
|
||
Fixed by moving all seven pieces of secret material into Kubernetes Secrets
|
||
(`secret.existingSecret` for six, `secret.additionalSecrets` for the JWKS key)
|
||
and referencing them by `path:`. The Secrets were built from the live
|
||
chart-generated Secret, so **no key material changed** — verified afterwards by
|
||
the JWKS endpoint still serving `kid=main` with the same modulus, meaning no
|
||
issued token was invalidated and nobody was logged out.
|
||
|
||
`authelia/values.yaml` is now committed. That was the last of the four configs
|
||
gitignored for embedded secrets.
|
||
|
||
### Incident — Authelia down ~5 minutes on the first attempt
|
||
|
||
The first upgrade put the JWKS key as a seventh key inside the `existingSecret`.
|
||
The chart projects that volume with an explicit `items:` list containing only
|
||
the six keys it generates, so the extra key was stored but **never mounted**.
|
||
Authelia died on `open /secrets/internal/…jwks.main.pem: no such file or
|
||
directory`, which cascaded into every other option appearing "required" because
|
||
the whole config template had failed to render.
|
||
|
||
Rolled back first to restore SSO, then fixed with `secret.additionalSecrets`,
|
||
which mounts a second Secret at `/secrets/<name>`. Two lessons: `helm template`
|
||
is not sufficient on its own — it happily rendered a config referencing a file
|
||
no volume projected — so the check that matters is cross-referencing every
|
||
`/secrets/...` reference against the rendered volumes' `items:`. And the
|
||
existingSecret volume mounts at `/secrets/internal`, not `/secrets/<name>`.
|
||
|
||
### Live S3 credentials were committed in the initial commit — now rotated
|
||
|
||
**Rotated 2026-07-28.** The leaked `anvAdmin` key is dead: verified denied
|
||
against the live endpoint. `anvReadOnly` and `terraform` were never exposed and
|
||
were left alone.
|
||
|
||
Rotating it broke `research-auto`, which turned out to be using the cluster-wide
|
||
admin key as its own S3 credentials. That dependency was invisible from this
|
||
repo — the app lives in `~/research-auto` and its Secret had been applied ad hoc,
|
||
with no owner references and its whole `k8s/` directory untracked. Finding it
|
||
needed a scan of every Secret in the cluster for the leaked key, not a grep of
|
||
this repo.
|
||
|
||
Fixed properly rather than by re-pointing it at the new admin key: `research-auto`
|
||
now has its own SeaweedFS identity scoped to the `research` bucket, written into
|
||
`~/research-auto/k8s/secrets.yaml` (gitignored, alongside the existing
|
||
`secrets.example.yaml` template) and applied. Both deployments were restarted —
|
||
these are env vars, so running pods keep the old value until recreated.
|
||
|
||
The orphaned `seaweedfs-s3-secret`, which the chart stopped generating once
|
||
`existingConfigSecret` was set but which still held the dead key, was deleted.
|
||
A cluster-wide scan now finds the leaked key in no Secret at all.
|
||
|
||
### Original exposure
|
||
|
||
`seaweedfs/values.yaml` carried the `anvAdmin` accessKey/secretKey inline and went
|
||
into git with the very first commit. They are still in history.
|
||
|
||
They survived three separate secret scans. The reason is instructive: the scan
|
||
regex looked for `secret[:=]`, and the key is written **`secretKey:`** — the word
|
||
"secret" is followed by "Key", not a colon. Together with the earlier `PASSWD:`
|
||
miss (case) and the `values.yaml`/`auth.json` misses (filename, not content),
|
||
that is three different ways the same class of scan fails.
|
||
|
||
Now externalised: the identities live in OpenBao at `kv/k8s/seaweedfs-s3`, ESO
|
||
syncs them, and the chart reads `filer.s3.existingConfigSecret` instead of
|
||
rendering credentials from values. **The leaked `anvAdmin` key still needs
|
||
rotating** — externalising stops it getting worse, it does not undo history.
|
||
|
||
### The cloudflared config was dead, not secret-bearing
|
||
|
||
Externalising `cloudflared/cloudflared.yaml` turned out to be the wrong fix: the
|
||
embedded ConfigMap and credentials Secret were **not in use at all**. Three
|
||
independent proofs — the config routed `idm.ddupan.top` to keycloak (retired
|
||
2026-07-10); it pointed `auth.ddupan.top` at `authelia:9091`, which 502s, while
|
||
Terraform had corrected that to `:80` and auth demonstrably works; and the
|
||
credentials volume mounted `subPath: <uuid>.json` against a Secret whose key was
|
||
`credentials-file`, so that mount never resolved.
|
||
|
||
The tunnel is token-managed and its ingress rules come from the Cloudflare API
|
||
via `cloudflared/terraform`. Confirmed on restart, which logged
|
||
`Updated to new configuration` carrying exactly the Terraform-managed rules, with
|
||
`authelia:80`. So both documents were deleted instead of being re-plumbed.
|
||
|
||
### Carried forward
|
||
|
||
- The OpenBao **PostgreSQL secrets engine** is the next step beyond static values:
|
||
Gitea and Authelia both read their credentials only at startup, so short-TTL
|
||
dynamic credentials would break them. Static roles (stable username, scheduled
|
||
password rotation) plus something to restart the consumer is the shape that fits.
|
||
`gitea.extraEnvSourceFile` and Authelia's `path:` indirection already read from
|
||
files, which is what an OpenBao agent-injector writes.
|
||
- The `cloudflared-tunnel` Secret still carried the dead `credentials-file` key
|
||
until today: `kubectl apply` MERGES, so removing it from the manifest did not
|
||
remove it from the cluster. Removed with a JSON patch. Worth remembering whenever
|
||
a key is dropped from a Secret.
|
||
- `.terraform.lock.hcl` is ignored in `openbao/` and `netbox/` but committed in
|
||
the other two roots. That is backwards — provider versions should be pinned.
|
||
- Gitea's Ingress declares no class, and the only classes present are `contour`
|
||
(retired) and `tailscale`. Gitea is therefore reachable only via the Cloudflare
|
||
tunnel, i.e. it depends on the WAN.
|
||
- No Actions runner deployed; CI substrate undecided.
|
||
- **IPv6 is broken on the laptop** (see above). Worked around per-process only;
|
||
`net.ipv6.conf.br0.accept_ra=2` still needs applying deliberately.
|
||
- `ddupan.top` still publishes SPF `~all` and DMARC `p=none`. Both should harden
|
||
(`-all`, `p=quarantine`) once a few days of aggregate reports confirm DKIM
|
||
passes — hardening before that would quarantine the lab's own mail.
|