Author SHA1 Message Date
panxiao81 57b13387cb ci: 输出 kind 节点准备调试日志
kind-on-kata-smoke / smoke (push) Failing after 23m50s
2026-09-14 16:16:36 +00:00
panxiao81 1deb024621 ci: 移除错误的容器内磁盘检查
kind-on-kata-smoke / smoke (push) Failing after 13m13s
2026-09-14 16:02:13 +00:00
panxiao81 311a506982 ci: 记录 kind guest 临时盘容量
kind-on-kata-smoke / smoke (push) Failing after 8s
2026-09-14 16:00:51 +00:00
panxiao81 a0d38fc9ab ci: 验证 kind 使用 guest-local overlay2
kind-on-kata-smoke / smoke (push) Failing after 10m59s
2026-09-14 15:54:28 +00:00
panxiao81 fa293ffc80 ci: 重试 kind 工具下载
kind-on-kata-smoke / smoke (push) Failing after 16m43s
2026-09-14 14:53:46 +00:00
panxiao81 77df686981 ci: 向 kind node 提供 guest kmsg
kind-on-kata-smoke / smoke (push) Failing after 28s
2026-09-14 14:51:37 +00:00
panxiao81 9e6515406f ci: 将 guest Docker 暴露给 kind job
kind-on-kata-smoke / smoke (push) Failing after 5m45s
2026-09-14 14:42:10 +00:00
panxiao81 f6c216ce93 ci: 验证 Kata 内运行 kind 集群
kind-on-kata-smoke / smoke (push) Failing after 15s
2026-09-14 14:38:14 +00:00
panxiao81 e0e8213b8f Merge pull request '部署 zot:接入 SeaweedFS、SPIRE 并统一 S3 凭据来源' (#55) from feat/zot-spire-s3 into main
yaml / yaml (push) Successful in 11s
ansible / collection-test (push) Successful in 50s
ansible / lint (push) Successful in 11m30s
Reviewed-on: #55
2026-09-14 13:38:17 +00:00
panxiao81 5c2b575a4f feat(zot): 接入 SeaweedFS 与 SPIRE 并统一 S3 凭据来源
yaml / yaml (pull_request) Successful in 20s
ansible / collection-test (pull_request) Successful in 59s
ansible / lint (pull_request) Successful in 10m32s
2026-09-14 13:26:54 +00:00
panxiao81 67dfe1ddcb Merge pull request 54: 记录 SPIRE 与 OpenBao workload identity 用法 2026-09-14 10:33:14 +00:00
panxiao81 ed86738c4b 文档:记录 SPIRE 与 OpenBao workload identity 用法 2026-09-14 10:31:36 +00:00
panxiao81 8c27a7286b Merge pull request 53: 配置 SPIRE JWT-SVID 登录 OpenBao
terraform / validate (push) Successful in 1m5s
2026-09-13 16:01:52 +00:00
panxiao81 c9987122e6 配置 SPIRE JWT-SVID 登录 OpenBao
terraform / validate (pull_request) Successful in 1m12s
2026-09-13 15:56:08 +00:00
panxiao81 9139f6f1f1 Merge pull request 52: 暴露 SPIRE OIDC discovery endpoint
yaml / yaml (push) Successful in 14s
ansible / collection-test (push) Successful in 1m8s
ansible / lint (push) Successful in 2m10s
2026-09-13 15:50:21 +00:00
panxiao81 b6ad65d768 暴露 SPIRE OIDC discovery endpoint
ansible / collection-test (pull_request) Successful in 2m17s
yaml / yaml (pull_request) Successful in 28s
ansible / lint (pull_request) Successful in 4m21s
2026-09-13 15:44:18 +00:00
panxiao81 d256682d11 Merge pull request 49: 建立 DNS source of truth
yaml / yaml (push) Successful in 20s
ansible / collection-test (push) Successful in 2m3s
ansible / lint (push) Successful in 3m17s
2026-09-13 15:42:12 +00:00
panxiao81 bdacc03e74 Merge pull request 51: 部署 SPIRE workload identity 基础设施
yaml / yaml (push) Successful in 18s
2026-09-13 15:35:53 +00:00
panxiao81 22192a3c68 部署 SPIRE workload identity 基础设施
yaml / yaml (pull_request) Successful in 13s
2026-09-13 15:33:20 +00:00
panxiao81 e2be2ca045 fix(dns): 声明公网隧道 CNAME 目标
yaml / yaml (pull_request) Successful in 18s
ansible / collection-test (pull_request) Successful in 54s
ansible / lint (pull_request) Successful in 5m13s
2026-09-10 18:04:02 +00:00
panxiao81 b390565146 ci(ansible): 并行测试本地 collection
ansible / lint (pull_request) Successful in 2m40s
yaml / yaml (pull_request) Successful in 16s
ansible / collection-test (pull_request) Successful in 54s
2026-09-10 17:48:19 +00:00
panxiao81 97021109ac feat(dns): 建立统一清单并接管 Samba 静态记录
yaml / yaml (pull_request) Successful in 39s
ansible / lint (pull_request) Failing after 3m5s
2026-09-10 17:24:49 +00:00
panxiao81 422a640c0f Merge pull request '修复 Grafana 在 ZFS RWO PVC 上的升级卡死' (#48) from fix/grafana-zfs-recreate-strategy into main
yaml / yaml (push) Successful in 11s
Reviewed-on: #48
2026-09-10 16:34:55 +00:00
panxiao81 d09f62c063 fix(grafana): 使用 Recreate 挂载 ZFS PVC
yaml / yaml (pull_request) Successful in 11s
2026-09-10 15:55:28 +00:00
panxiao81 20a60b77d0 Merge pull request '激活剩余可观测性栈 Flux 接管' (#44) from feat/observability-remainder-adoption-activate into main
yaml / yaml (push) Successful in 13s
Reviewed-on: #44
2026-09-10 15:47:00 +00:00
panxiao81 15444b9eb3 feat(flux): 激活剩余可观测性栈接管
yaml / yaml (pull_request) Successful in 12s
2026-09-10 12:26:09 +00:00
panxiao81 065f96d432 Merge pull request '统一接管剩余可观测性栈' (#43) from feat/observability-remainder-adoption-stage into main
yaml / yaml (push) Successful in 11s
Reviewed-on: #43
2026-09-10 12:23:22 +00:00
panxiao81 92820fc631 feat(flux): 统一接管剩余可观测性栈
yaml / yaml (pull_request) Successful in 10s
2026-09-10 11:53:06 +00:00
panxiao81 6a54afede3 Merge pull request '激活 VictoriaMetrics Operator Flux 接管' (#42) from feat/vm-operator-adoption-activate into main
yaml / yaml (push) Successful in 11s
Reviewed-on: #42
2026-09-10 11:45:44 +00:00
panxiao81 1ea4604dfc feat(flux): 激活 VictoriaMetrics Operator 接管
yaml / yaml (pull_request) Successful in 12s
2026-09-10 11:35:25 +00:00
panxiao81 47c8133034 Merge pull request '分阶段接管 VictoriaMetrics Operator' (#41) from feat/vm-operator-adoption-stage into main
yaml / yaml (push) Successful in 10s
Reviewed-on: #41
2026-09-10 11:32:45 +00:00
panxiao81 b9645be107 feat(flux): 分阶段接管 VictoriaMetrics Operator
yaml / yaml (pull_request) Successful in 10s
2026-09-10 11:23:35 +00:00
panxiao81 ffe21da281 Merge pull request '激活 OpenEBS Flux HelmRelease 接管' (#40) from feat/openebs-adoption-activate into main
yaml / yaml (push) Successful in 10s
Reviewed-on: #40
2026-09-10 11:12:48 +00:00
panxiao81 b631bae5b3 feat(flux): 激活 OpenEBS 接管
yaml / yaml (pull_request) Successful in 11s
2026-09-10 11:06:29 +00:00
panxiao81 addd103b2c Merge pull request '分阶段接管 OpenEBS Helm release' (#39) from feat/openebs-adoption-stage into main
yaml / yaml (push) Successful in 21s
Reviewed-on: #39
2026-09-10 11:03:56 +00:00
panxiao81 f000301864 feat(flux): 分阶段接管 OpenEBS
yaml / yaml (pull_request) Successful in 30s
2026-09-10 11:01:21 +00:00
panxiao81 abf3be0eae Merge pull request #38: 激活 Envoy Gateway Flux 接管
yaml / yaml (push) Successful in 21s
2026-09-10 09:49:01 +00:00
panxiao81 4070b75436 feat(flux): 激活 Envoy Gateway 接管
yaml / yaml (pull_request) Successful in 10s
2026-09-10 09:45:25 +00:00
panxiao81 7d72e12d60 Merge pull request '分阶段接管 Envoy Gateway Helm release' (#37) from feat/envoy-gateway-adoption-stage into main
yaml / yaml (push) Successful in 10s
Reviewed-on: #37
2026-09-10 09:29:37 +00:00
panxiao81 c6e3abbfc0 feat(flux): 分阶段接管 Envoy Gateway
yaml / yaml (pull_request) Successful in 10s
2026-09-10 09:26:54 +00:00
panxiao81 392df6b6e6 Merge pull request '激活 cert-manager Flux HelmRelease 接管' (#36) from feat/cert-manager-adoption-activate into main
yaml / yaml (push) Successful in 29s
Reviewed-on: #36
2026-09-10 09:21:21 +00:00
panxiao81 d8754100f1 feat(flux): 激活 cert-manager 接管
yaml / yaml (pull_request) Successful in 37s
2026-09-10 09:20:03 +00:00
panxiao81 8c18d9daf3 Merge pull request '分阶段接管 cert-manager Helm release' (#35) from feat/cert-manager-adoption-stage into main
yaml / yaml (push) Successful in 34s
Reviewed-on: #35
2026-09-10 09:17:55 +00:00
panxiao81 2a1d1e095e feat(flux): 分阶段接管 cert-manager
yaml / yaml (pull_request) Successful in 10s
2026-09-10 09:15:26 +00:00
panxiao81 039ed46c8e Merge pull request '激活 Flux 对 External Secrets 的接管' (#33) from feat/external-secrets-adoption-activate into main
yaml / yaml (push) Successful in 11s
2026-09-10 08:44:37 +00:00
panxiao81 c47922855f feat(flux): 激活 External Secrets 接管
yaml / yaml (pull_request) Successful in 10s
2026-09-10 08:43:35 +00:00
panxiao81 9a581a2097 Merge pull request '分阶段接管 External Secrets Helm release' (#32) from feat/external-secrets-adoption-stage into main
yaml / yaml (push) Successful in 10s
Reviewed-on: #32
2026-09-10 08:40:01 +00:00
panxiao81 c4b425b079 feat(flux): 分阶段接管 External Secrets
yaml / yaml (pull_request) Successful in 9s
2026-09-10 08:39:22 +00:00
panxiao81 02fd9a932a Merge pull request '按路径拆分静态 CI 并消除重复运行' (#31) from ci/path-filtered-workflows into main
yaml / yaml (push) Successful in 23s
terraform / validate (push) Successful in 23s
ansible / lint (push) Successful in 3m57s
Reviewed-on: #31
2026-09-10 08:34:42 +00:00
panxiao81 8f6b08bd20 ci: 按路径拆分静态检查
yaml / yaml (pull_request) Successful in 20s
terraform / validate (pull_request) Successful in 22s
ansible / lint (pull_request) Successful in 4m5s
2026-09-10 08:33:26 +00:00
panxiao81 1b2e551392 Merge pull request '修复 Actions user-defined network 的 MTU' (#30) from fix/gitea-runner-user-network-mtu into main
lint / yaml (push) Successful in 11s
lint / terraform (push) Successful in 19s
lint / ansible (push) Successful in 5m28s
Reviewed-on: #30
2026-09-10 07:56:57 +00:00
panxiao81 4d6274bb8f 修复 Actions 动态网络的 MTU
lint / yaml (push) Successful in 1m6s
lint / terraform (push) Successful in 1m15s
lint / ansible (push) Successful in 5m18s
lint / yaml (pull_request) Successful in 11s
lint / terraform (pull_request) Successful in 1m15s
lint / ansible (pull_request) Successful in 8m6s
2026-09-10 07:55:39 +00:00
panxiao81 def6ba185d Merge pull request '记录 Gitea 后续精简升级策略' (#28) from docs/gitea-upgrade-policy into main
lint / yaml (push) Successful in 9s
lint / terraform (push) Successful in 29s
lint / ansible (push) Successful in 4m17s
Reviewed-on: #28
2026-09-10 07:44:04 +00:00
panxiao81 19cd0a938b docs(gitea): 记录精简升级验收策略
lint / ansible (push) Successful in 3m20s
lint / yaml (push) Successful in 9s
lint / terraform (push) Successful in 32s
lint / yaml (pull_request) Successful in 9s
lint / terraform (pull_request) Successful in 31s
lint / ansible (pull_request) Successful in 4m58s
2026-09-10 07:40:06 +00:00
panxiao81 c4f7046e1c Merge pull request '修复 Gitea runner 访问 GitHub 超时' (#29) from fix/gitea-runner-mtu into main
lint / yaml (push) Successful in 12s
lint / terraform (push) Successful in 30s
lint / ansible (push) Failing after 15m55s
Reviewed-on: #29
2026-09-10 07:34:34 +00:00
panxiao81 33627573c3 修复 Gitea runner 的 DinD MTU
lint / yaml (push) Successful in 18s
lint / terraform (pull_request) Successful in 35s
lint / terraform (push) Successful in 30s
lint / yaml (pull_request) Successful in 17s
lint / ansible (push) Failing after 19m37s
lint / ansible (pull_request) Failing after 19m55s
2026-09-10 07:30:09 +00:00
panxiao81 5539793e10 Merge pull request: 激活 Gitea 1.27.3 升级
lint / yaml (push) Successful in 14s
lint / terraform (push) Successful in 31s
lint / ansible (push) Successful in 9m3s
2026-09-10 07:14:25 +00:00
panxiao81 789872264e feat(gitea): 激活 1.27.3 升级
lint / yaml (pull_request) Successful in 12s
lint / terraform (pull_request) Successful in 33s
lint / yaml (push) Successful in 10s
lint / terraform (push) Successful in 30s
lint / ansible (pull_request) Successful in 6m52s
lint / ansible (push) Successful in 6m55s
2026-09-10 07:04:24 +00:00
panxiao81 48f6e52e61 Merge pull request: 准备 Gitea 1.27.3 suspended 升级
lint / yaml (push) Successful in 11s
lint / terraform (push) Successful in 30s
lint / ansible (push) Successful in 7m9s
2026-09-10 07:01:04 +00:00
panxiao81 7a1c83054d docs(gitea): 记录 1.26.4 OIDC 验收成功
lint / yaml (push) Successful in 13s
lint / ansible (push) Successful in 5m31s
lint / terraform (push) Successful in 30s
lint / yaml (pull_request) Successful in 8s
lint / terraform (pull_request) Successful in 33s
lint / ansible (pull_request) Successful in 6m14s
2026-09-10 06:58:49 +00:00
panxiao81 01cca07f48 feat(gitea): 准备 1.27.3 suspended 升级
lint / yaml (pull_request) Successful in 18s
lint / yaml (push) Successful in 17s
lint / terraform (pull_request) Successful in 31s
lint / terraform (push) Successful in 32s
lint / ansible (pull_request) Successful in 3m51s
lint / ansible (push) Successful in 3m34s
2026-09-10 06:56:57 +00:00
panxiao81 f9bcde53ab Merge pull request: 激活 Gitea 1.26.4 升级
lint / ansible (push) Successful in 4m16s
lint / terraform (push) Successful in 42s
lint / yaml (push) Successful in 10s
2026-09-10 06:38:18 +00:00
panxiao81 1538a8a3c8 feat(gitea): 激活 1.26.4 升级
lint / yaml (pull_request) Successful in 23s
lint / yaml (push) Successful in 23s
lint / terraform (push) Successful in 37s
lint / terraform (pull_request) Successful in 30s
lint / ansible (push) Successful in 4m14s
lint / ansible (pull_request) Successful in 4m29s
2026-09-10 06:37:28 +00:00
panxiao81 6bddb744c9 Merge pull request: 准备 Gitea 1.26.4 suspended 升级
lint / yaml (push) Successful in 12s
lint / terraform (push) Successful in 30s
lint / ansible (push) Successful in 6m35s
2026-09-10 06:01:31 +00:00
panxiao81 bac85b6335 feat(gitea): 准备 1.26.4 suspended 升级
lint / yaml (pull_request) Successful in 14s
lint / yaml (push) Successful in 15s
lint / terraform (push) Successful in 34s
lint / terraform (pull_request) Successful in 33s
lint / ansible (push) Successful in 4m21s
lint / ansible (pull_request) Successful in 5m1s
2026-09-10 05:59:23 +00:00
panxiao81 af23195b04 docs: 规划 Gitea 1.27 分阶段升级
lint / yaml (pull_request) Successful in 15s
lint / yaml (push) Successful in 15s
lint / terraform (pull_request) Successful in 32s
lint / terraform (push) Successful in 33s
lint / ansible (pull_request) Successful in 3m19s
lint / ansible (push) Successful in 3m5s
2026-09-10 05:50:40 +00:00
panxiao81 91c6defe49 Merge pull request: 启用 Flux 对 Gitea Helm release 的接管
lint / yaml (push) Successful in 13s
lint / terraform (push) Successful in 31s
lint / ansible (push) Successful in 3m52s
2026-09-10 05:39:15 +00:00
panxiao81 b4c61c5e10 feat(gitops): 启用 Gitea Helm 接管
lint / yaml (push) Successful in 17s
lint / yaml (pull_request) Successful in 20s
lint / terraform (push) Successful in 38s
lint / terraform (pull_request) Successful in 31s
lint / ansible (push) Successful in 4m16s
lint / ansible (pull_request) Successful in 4m5s
2026-09-10 05:38:43 +00:00
panxiao81 bcf4f5c0f1 Merge pull request: 分阶段接管 Gitea Helm release
lint / yaml (push) Successful in 14s
lint / terraform (push) Successful in 31s
lint / ansible (push) Has been cancelled
2026-09-10 05:35:26 +00:00
panxiao81 37abad9a04 feat(gitops): staged 接管 Gitea Helm release
lint / yaml (push) Successful in 13s
lint / terraform (pull_request) Successful in 32s
lint / terraform (push) Successful in 30s
lint / ansible (pull_request) Successful in 3m39s
lint / ansible (push) Successful in 4m9s
lint / yaml (pull_request) Successful in 13s
2026-09-10 05:32:46 +00:00
panxiao81 e83cf1932c Merge pull request: 启用 Flux 对 Gitea Actions Helm release 的接管
lint / yaml (push) Successful in 11s
lint / terraform (push) Successful in 28s
lint / ansible (push) Successful in 3m19s
2026-09-10 05:24:02 +00:00
panxiao81 5e2e28e275 feat(gitops): 启用 Gitea Actions Helm 接管
lint / yaml (push) Successful in 14s
lint / yaml (pull_request) Successful in 14s
lint / terraform (pull_request) Successful in 34s
lint / terraform (push) Successful in 29s
lint / ansible (pull_request) Successful in 3m12s
lint / ansible (push) Successful in 3m57s
2026-09-10 05:22:30 +00:00
panxiao81 a8de4d257b Merge pull request: 分阶段接管 Gitea Actions Helm release
lint / yaml (push) Successful in 13s
lint / terraform (push) Successful in 31s
lint / ansible (push) Successful in 3m26s
2026-09-10 05:18:17 +00:00
panxiao81 745d2cc6c2 feat(gitops): staged 接管 Gitea Actions Helm release
lint / yaml (push) Successful in 14s
lint / yaml (pull_request) Successful in 17s
lint / terraform (pull_request) Successful in 33s
lint / terraform (push) Successful in 29s
lint / ansible (push) Successful in 6m18s
lint / ansible (pull_request) Successful in 6m13s
2026-09-10 05:01:46 +00:00
panxiao81 a9f6663069 Merge pull request: 更新 GitOps 完成状态与下一 Helm 迁移
lint / yaml (push) Successful in 13s
lint / terraform (push) Successful in 34s
lint / ansible (push) Successful in 2m36s
2026-09-10 04:55:50 +00:00
panxiao81 a5cbe89ae2 docs: 更新 GitOps 状态与 Helm 迁移顺序
lint / yaml (pull_request) Successful in 17s
lint / yaml (push) Successful in 17s
lint / terraform (push) Successful in 37s
lint / terraform (pull_request) Successful in 29s
lint / ansible (push) Successful in 3m51s
lint / ansible (pull_request) Successful in 4m6s
2026-09-10 04:19:40 +00:00
panxiao81 f257a2aa0a Merge pull request: 最终验证 Flux canary 自动 prune
lint / yaml (push) Successful in 17s
lint / yaml (pull_request) Successful in 17s
lint / terraform (pull_request) Successful in 34s
lint / terraform (push) Successful in 32s
lint / ansible (pull_request) Successful in 3m6s
lint / ansible (push) Successful in 2m58s
2026-09-10 02:11:19 +00:00
panxiao81 b7775fce94 test(gitops): 最终删除 prune canary
lint / yaml (push) Successful in 9s
lint / ansible (push) Successful in 3m4s
lint / terraform (push) Successful in 31s
lint / yaml (pull_request) Successful in 9s
lint / ansible (pull_request) Successful in 2m56s
lint / terraform (pull_request) Successful in 35s
2026-09-09 20:20:00 +00:00
panxiao81 f1bcb9017a Merge pull request '在 prune 开启后重新纳管测试对象' (#17) from test/flux-prune-canary-readopt into main
lint / yaml (push) Successful in 9s
lint / ansible (push) Successful in 2m41s
lint / terraform (push) Successful in 33s
Reviewed-on: #17
2026-09-09 20:18:33 +00:00
panxiao81 304e18d219 test(gitops): 在 prune 开启后重新纳管测试对象
lint / terraform (push) Successful in 30s
lint / yaml (pull_request) Successful in 8s
lint / ansible (pull_request) Successful in 2m55s
lint / terraform (pull_request) Successful in 30s
lint / yaml (push) Successful in 8s
lint / ansible (push) Successful in 1m59s
2026-09-09 20:13:53 +00:00
panxiao81 342ba4f114 Merge pull request '验证 Flux canary 受控删除' (#16) from test/flux-prune-canary-remove into main
lint / yaml (push) Has been cancelled
lint / ansible (push) Has been cancelled
lint / terraform (push) Has been cancelled
Reviewed-on: #16
2026-09-09 20:10:36 +00:00
panxiao81 45e3effcbe test(gitops): 验证 canary prune
lint / yaml (pull_request) Successful in 10s
lint / ansible (pull_request) Successful in 3m42s
lint / terraform (pull_request) Successful in 30s
lint / yaml (push) Successful in 9s
lint / ansible (push) Successful in 2m32s
lint / terraform (push) Successful in 30s
2026-09-09 20:08:09 +00:00
panxiao81 f1d1a1a906 Merge pull request '添加 Flux prune 测试对象' (#15) from test/flux-prune-canary-add into main
lint / yaml (push) Has been cancelled
lint / ansible (push) Has been cancelled
lint / terraform (push) Has been cancelled
Reviewed-on: #15
2026-09-09 20:05:22 +00:00
panxiao81 531011c257 test(gitops): 添加 Flux prune canary
lint / yaml (push) Successful in 8s
lint / ansible (push) Successful in 3m19s
lint / terraform (push) Successful in 32s
lint / yaml (pull_request) Successful in 9s
lint / ansible (pull_request) Successful in 3m8s
lint / terraform (pull_request) Successful in 34s
2026-09-09 20:04:34 +00:00
panxiao81 11049f99c7 Merge pull request '添加 Flux http-echo canary' (#14) from feat/flux-http-echo-canary into main
lint / yaml (push) Successful in 16s
lint / terraform (push) Successful in 35s
lint / ansible (push) Has been cancelled
Reviewed-on: #14
2026-09-09 20:00:13 +00:00
panxiao81 5d0441144d feat(gitops): 添加 http-echo canary
lint / yaml (push) Successful in 17s
lint / terraform (push) Successful in 55s
lint / yaml (pull_request) Successful in 37s
lint / terraform (pull_request) Successful in 42s
lint / ansible (push) Successful in 3m18s
lint / ansible (pull_request) Successful in 3m57s
2026-09-09 19:56:17 +00:00
panxiao81 ee75afe2d2 Merge pull request '记录 Flux 上线与 Contour 退役完成' (#13) from docs/record-flux-live into main
lint / yaml (push) Successful in 19s
lint / terraform (push) Successful in 38s
lint / ansible (push) Successful in 4m25s
Reviewed-on: #13
2026-09-09 19:53:23 +00:00
panxiao81 d1be2b1a9f docs: 记录 Flux 上线与 Contour 退役
lint / yaml (pull_request) Successful in 19s
lint / yaml (push) Successful in 20s
lint / terraform (pull_request) Successful in 40s
lint / terraform (push) Successful in 28s
lint / ansible (push) Successful in 3m53s
lint / ansible (pull_request) Successful in 3m36s
2026-09-09 19:49:50 +00:00
panxiao81 aaa54a1289 Merge pull request '引入 Flux 2.9.5 并彻底退役 Contour' (#12) from feat/flux-bootstrap into main
lint / yaml (push) Successful in 14s
lint / terraform (push) Successful in 32s
lint / ansible (push) Successful in 4m47s
Reviewed-on: #12
2026-09-09 19:45:24 +00:00
panxiao81 b343cb95f0 docs: 修正 Contour 清理文档格式
lint / yaml (push) Successful in 10s
lint / yaml (pull_request) Successful in 10s
lint / terraform (push) Successful in 34s
lint / ansible (push) Successful in 4m29s
lint / ansible (pull_request) Successful in 4m57s
lint / terraform (pull_request) Successful in 30s
2026-09-09 19:33:33 +00:00
panxiao81 96f44c7e6a feat(gitops): 引入 Flux 2.9.5 并退役 Contour
lint / yaml (push) Successful in 15s
lint / terraform (push) Has been cancelled
lint / ansible (push) Has been cancelled
2026-09-09 19:32:57 +00:00
panxiao81 4bfeef464d Merge pull request '文档:默认使用中文协作' (#10) from docs/prefer-chinese into main
lint / yaml (push) Successful in 9s
lint / terraform (push) Successful in 29s
lint / ansible (push) Successful in 3m57s
Reviewed-on: #10
2026-09-09 18:35:34 +00:00
panxiao81 6d5f507faa 文档:默认使用中文协作
lint / yaml (push) Successful in 13s
lint / yaml (pull_request) Successful in 13s
lint / terraform (pull_request) Successful in 33s
lint / terraform (push) Successful in 29s
lint / ansible (push) Successful in 3m56s
lint / ansible (pull_request) Successful in 4m17s
2026-09-09 18:35:09 +00:00
panxiao81 102fbf2ae6 Merge pull request 'docs: record live CI bootstrap' (#8) from docs/runner-bootstrap-status into main
lint / yaml (push) Successful in 10s
lint / terraform (push) Successful in 29s
lint / ansible (push) Successful in 4m40s
Reviewed-on: #8
2026-09-09 18:29:38 +00:00
panxiao81 0984a2d0ef docs: record live CI bootstrap
lint / yaml (push) Successful in 15s
lint / yaml (pull_request) Successful in 15s
lint / terraform (pull_request) Successful in 34s
lint / terraform (push) Successful in 31s
lint / ansible (push) Successful in 3m26s
lint / ansible (pull_request) Successful in 4m7s
2026-09-09 18:28:45 +00:00
panxiao81 10a21df890 Merge pull request 'fix(ci): share Ansible Galaxy collections' (#6) from fix/ansible-lint-collections into main
lint / yaml (push) Successful in 12s
lint / terraform (push) Successful in 33s
lint / ansible (push) Successful in 5m0s
Reviewed-on: #6
2026-09-09 18:21:33 +00:00
panxiao81 eeea0dc092 fix(ci): share Ansible Galaxy collections
lint / yaml (pull_request) Successful in 15s
lint / terraform (pull_request) Successful in 38s
lint / terraform (push) Successful in 34s
lint / yaml (push) Successful in 14s
lint / ansible (pull_request) Successful in 5m5s
lint / ansible (push) Successful in 5m10s
2026-09-09 18:15:35 +00:00
panxiao81 4ee123965c Merge pull request 'fix(ci): install uv from official PyPI' (#5) from fix/ci-uv-mirror into main
lint / yaml (push) Successful in 11s
lint / terraform (push) Successful in 38s
lint / ansible (push) Failing after 1m6s
Reviewed-on: #5
2026-09-09 18:03:55 +00:00
panxiao81 d5d529dcd9 fix(ci): install uv from official PyPI
lint / yaml (push) Successful in 13s
lint / yaml (pull_request) Successful in 13s
lint / terraform (push) Successful in 32s
lint / terraform (pull_request) Successful in 31s
lint / ansible (push) Failing after 1m9s
lint / ansible (pull_request) Failing after 1m9s
2026-09-09 18:03:22 +00:00
panxiao81 661f841837 Merge pull request 'fix(ci): bootstrap lint toolchain' (#4) from fix/ci-tool-bootstrap into main
lint / ansible (push) Failing after 13s
lint / terraform (push) Successful in 28s
lint / yaml (push) Failing after 13s
Reviewed-on: #4
2026-09-09 17:56:53 +00:00
panxiao81 3de2a078b9 fix(ci): bootstrap lint toolchain
lint / yaml (push) Failing after 14s
lint / ansible (push) Failing after 13s
lint / yaml (pull_request) Failing after 13s
lint / terraform (push) Has been cancelled
lint / ansible (pull_request) Failing after 12s
lint / terraform (pull_request) Successful in 31s
2026-09-09 17:56:18 +00:00
panxiao81 0644c25567 Merge pull request 'fix(ci): use compatible DinD runner' (#3) from fix/gitea-runner-dind into main
lint / terraform (push) Has been cancelled
lint / yaml (push) Has been cancelled
lint / ansible (push) Has been cancelled
Reviewed-on: #3
2026-09-09 17:49:29 +00:00
panxiao81 e513739ba0 fix(ci): use compatible DinD runner
lint / terraform (push) Failing after 2s
lint / yaml (push) Failing after 52s
lint / terraform (pull_request) Failing after 2s
lint / yaml (pull_request) Failing after 52s
lint / ansible (push) Failing after 1m42s
lint / ansible (pull_request) Failing after 1m43s
2026-09-09 17:48:40 +00:00
110 changed files with 9192 additions and 204 deletions
+79
View File
@@ -0,0 +1,79 @@
---
name: ansible
on:
push:
branches: [main]
paths:
- 'infrastructure/**/ansible/**'
- 'infrastructure/dns/**'
- '.ansible-lint'
- '.gitea/workflows/ansible.yml'
pull_request:
paths:
- 'infrastructure/**/ansible/**'
- 'infrastructure/dns/**'
- '.ansible-lint'
- '.gitea/workflows/ansible.yml'
env:
ANSIBLE_COLLECTIONS_PATH: /root/.ansible/collections
jobs:
lint:
runs-on: self-hosted
steps:
- uses: actions/checkout@v4
- name: Bootstrap uv
run: |
python3 -m pip install --user --break-system-packages \
--index-url https://pypi.org/simple --quiet uv==0.11.7
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
- name: Install ansible-lint and collections
run: |
for i in 1 2 3 4 5; do
uv tool install ansible-core --with paramiko --with pywinrm --quiet && break
echo "attempt $i failed"; sleep 10
done
for i in 1 2 3 4 5; do
uv tool install ansible-lint --quiet && break
echo "attempt $i failed"; sleep 10
done
export PATH="$HOME/.local/bin:$PATH"
for p in infrastructure/proxmox infrastructure/samba-ad infrastructure/openbao; do
ansible-galaxy collection install \
-r "$p/ansible/requirements.yml" -p "$ANSIBLE_COLLECTIONS_PATH"
done
- name: ansible-lint
run: |
export PATH="$HOME/.local/bin:$PATH"
export ANSIBLE_COLLECTIONS_PATH="$PWD/infrastructure/samba-ad/ansible/collections:/root/.ansible/collections"
rc=0
for p in infrastructure/openbao infrastructure/samba-ad infrastructure/proxmox; do
echo "::group::$p"
(cd "$p/ansible" && ansible-lint -c ../../../.ansible-lint --nocolor -f pep8 .) || rc=1
echo "::endgroup::"
done
exit $rc
collection-test:
runs-on: self-hosted
steps:
- uses: actions/checkout@v4
- name: Install ansible-core
run: |
python3 -m pip install --user --break-system-packages \
--index-url https://pypi.org/simple --quiet uv==0.11.7
export PATH="$HOME/.local/bin:$PATH"
uv tool install ansible-core --quiet
- name: Run ansible-test
working-directory: infrastructure/samba-ad/ansible/collections/ansible_collections/ddupan/homelab
run: |
export PATH="$HOME/.local/bin:$PATH"
ansible-test sanity --venv --requirements --python 3.12 --color no
ansible-test units --venv --requirements --python 3.12 --color no
+63
View File
@@ -0,0 +1,63 @@
---
name: kind-on-kata-smoke
on:
push:
branches:
- poc/kind-on-kata
paths:
- .gitea/workflows/kind-on-kata-smoke.yml
workflow_dispatch:
jobs:
smoke:
runs-on: kata-poc
steps:
- name: Create nested kind cluster
shell: sh
env:
KIND_VERSION: v0.27.0
KIND_NODE_IMAGE: kindest/node:v1.32.2@sha256:f226345927d7e348497136874b6d207e0b32cc52154ad8323129352923a3142f
run: |
set -eu
apk add --no-cache ca-certificates curl docker-cli
curl --retry 5 --retry-all-errors --connect-timeout 15 -fsSLo /tmp/kind \
"https://kind.sigs.k8s.io/dl/${KIND_VERSION}/kind-linux-amd64"
curl --retry 5 --retry-all-errors --connect-timeout 15 -fsSLo /tmp/kind.sha256sum \
"https://kind.sigs.k8s.io/dl/${KIND_VERSION}/kind-linux-amd64.sha256sum"
expected="$(awk '{print $1}' /tmp/kind.sha256sum)"
printf '%s %s\n' "$expected" /tmp/kind | sha256sum -c -
install -m 0755 /tmp/kind /usr/local/bin/kind
docker info --format 'kernel={{.KernelVersion}} driver={{.Driver}}'
test "$(docker info --format '{{.Driver}}')" = overlay2
cleanup() {
kind delete cluster --name nested >/dev/null 2>&1 || true
}
trap cleanup EXIT
cat >/tmp/kind-config.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
name: nested
nodes:
- role: control-plane
extraMounts:
- hostPath: /dev/kmsg
containerPath: /dev/kmsg
EOF
kind create cluster -v 9 --retain --config /tmp/kind-config.yaml --image "$KIND_NODE_IMAGE" --wait 5m
docker exec nested-control-plane kubectl \
--kubeconfig=/etc/kubernetes/admin.conf wait \
--for=condition=Ready node/nested-control-plane --timeout=2m
docker exec nested-control-plane kubectl \
--kubeconfig=/etc/kubernetes/admin.conf run smoke \
--image=docker.io/library/busybox:1.37 --restart=Never \
--command -- sh -c 'echo kind-on-kata-ok'
docker exec nested-control-plane kubectl \
--kubeconfig=/etc/kubernetes/admin.conf wait \
--for=jsonpath='{.status.phase}'=Succeeded pod/smoke --timeout=2m
test "$(docker exec nested-control-plane kubectl \
--kubeconfig=/etc/kubernetes/admin.conf logs smoke)" = kind-on-kata-ok
kind delete cluster --name nested
trap - EXIT
+20 -73
View File
@@ -1,28 +1,26 @@
---
# Stage 1 of the infra pipeline: static checks only. No cluster access, no
# credentials, no mutation — so this is safe to run on every push from day one.
# credentials or mutation. It runs only when YAML-related paths change.
#
# Stages 2 (kubectl --dry-run=server) and 3 (k3d / molecule) come later and DO
# need cluster access; keep them in separate workflows so a credential problem
# there can never block this one.
name: lint
name: yaml
on:
push:
branches: [main]
paths:
- '**/*.yaml'
- '**/*.yml'
- '.yamllint.yml'
- '.gitea/workflows/lint.yml'
pull_request:
env:
# pypi.org is NOT reachable from this network — it resolves fine but TCP/443 to
# Fastly (151.101.x) times out, while github.com and cloudflare.com are fine.
# This is not the usual flaky-WAN symptom and a plain `uv tool install` will
# hang until timeout. Use a mirror; verified reachable 2026-07-28.
UV_DEFAULT_INDEX: https://pypi.tuna.tsinghua.edu.cn/simple
# ansible-lint and ansible-core install as SEPARATE uv tools, each with its own
# venv. Collections installed under the ansible-core tool are invisible to
# ansible-lint, which then reports every module as `syntax-check[unknown-module]`
# — a false failure that looks exactly like a real one. Pin both to a shared path.
ANSIBLE_COLLECTIONS_PATH: /root/.ansible/collections
paths:
- '**/*.yaml'
- '**/*.yml'
- '.yamllint.yml'
- '.gitea/workflows/lint.yml'
jobs:
yaml:
@@ -30,6 +28,13 @@ jobs:
steps:
- uses: actions/checkout@v4
- name: Bootstrap uv
# Pin the tool for reproducibility; PyPI also avoids another setup action.
run: |
python3 -m pip install --user --break-system-packages \
--index-url https://pypi.org/simple --quiet uv==0.11.7
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
- name: Install yamllint
# The WAN drops at random (see CLAUDE.md); retry rather than fail a run.
run: |
@@ -47,61 +52,3 @@ jobs:
export PATH="$HOME/.local/bin:$PATH"
files=$(git ls-files '*.yaml' '*.yml' | grep -vE '^apps/netboot/')
yamllint -c .yamllint.yml --no-warnings -f parsable $files
ansible:
runs-on: self-hosted
steps:
- uses: actions/checkout@v4
- name: Install ansible-lint and collections
# pywinrm is not optional — without it every ansible.windows.* task dies
# with "No module named 'winrm'" (CLAUDE.md documents this trap).
run: |
for i in 1 2 3 4 5; do
uv tool install ansible-core --with ansible --with paramiko --with pywinrm --quiet && break
echo "attempt $i failed"; sleep 10
done
for i in 1 2 3 4 5; do
uv tool install ansible-lint --quiet && break
echo "attempt $i failed"; sleep 10
done
export PATH="$HOME/.local/bin:$PATH"
for p in infrastructure/proxmox infrastructure/samba-ad infrastructure/openbao; do
ansible-galaxy collection install \
-r "$p/ansible/requirements.yml" -p "$ANSIBLE_COLLECTIONS_PATH"
done
- name: ansible-lint
# Each project has its own ansible.cfg and relative roles_path, so lint
# must run from inside each one — a single run at the repo root resolves
# roles_path incorrectly and reports spurious missing-role errors.
run: |
export PATH="$HOME/.local/bin:$PATH"
rc=0
for p in infrastructure/openbao infrastructure/samba-ad infrastructure/proxmox; do
echo "::group::$p"
(cd "$p/ansible" && ansible-lint -c ../../../.ansible-lint --nocolor -f pep8 .) || rc=1
echo "::endgroup::"
done
exit $rc
terraform:
runs-on: self-hosted
steps:
- uses: actions/checkout@v4
- name: fmt and validate
# -backend=false so validate never touches real state or needs credentials.
# These roots deliberately use different providers AND different interactive
# auth (bao login -method=oidc, az login), which is exactly why they are not
# merged — so validate is as far as static checking can go here.
run: |
rc=0
for d in $(git ls-files '*.tf' | xargs -n1 dirname | sort -u); do
echo "::group::$d"
terraform -chdir="$d" fmt -check -diff || rc=1
terraform -chdir="$d" init -backend=false -input=false || rc=1
terraform -chdir="$d" validate || rc=1
echo "::endgroup::"
done
exit $rc
+35
View File
@@ -0,0 +1,35 @@
---
name: terraform
on:
push:
branches: [main]
paths:
- '**/*.tf'
- '**/.terraform.lock.hcl'
- '.gitea/workflows/terraform.yml'
pull_request:
paths:
- '**/*.tf'
- '**/.terraform.lock.hcl'
- '.gitea/workflows/terraform.yml'
jobs:
validate:
runs-on: self-hosted
steps:
- uses: actions/checkout@v4
- uses: hashicorp/setup-terraform@v3
- name: fmt and validate
# -backend=false so validate never touches real state or needs credentials.
run: |
rc=0
for d in $(git ls-files '*.tf' | xargs -n1 dirname | sort -u); do
echo "::group::$d"
terraform -chdir="$d" fmt -check -diff || rc=1
terraform -chdir="$d" init -backend=false -input=false || rc=1
terraform -chdir="$d" validate || rc=1
echo "::endgroup::"
done
exit $rc
+5
View File
@@ -4,3 +4,8 @@
- `apps/http-echo/` and `archive/traefik/` contain Kubernetes/Gateway API manifests; inspect their parent Gateway references before applying archived or brownfield resources.
- `apps/tailscale/helm.sh` contains live Tailscale OAuth values; do not copy, print, or commit those values anywhere else.
- Preserve the existing README intent in `apps/http-echo/` and `archive/traefik/` when updating manifests.
- `CHANGELOG.md` 是冻结的历史快照,不再更新。持久的服务状态与运维知识写入对应
README/runbook;单次变化由 commit 和 PR 记录,agent 陷阱写入 `CLAUDE.md`。
- 在 homelab 工作中,所有提交到 `git.ddupan.top` 的 commit message、PR、issue
和项目文档默认优先使用中文。代码标识符、命令、配置键、上游专有名称,以及
使用英文能避免歧义的技术字段可保留英文。
+57 -4
View File
@@ -15,6 +15,42 @@ What changed in this homelab, when, and why. Newest first.
---
## 2026-09-10
**Flux 的 deployment、drift repair 和 scoped prune 闭环验证完成。**
| area | change |
|---|---|
| GitOps | PR #18 合并后,Flux 自行发现 revision `f257a2a` 并删除已在 `prune: true` 下重新进入 inventory 的测试 ConfigMap;未发送 reconcile annotation,`http-echo` Deployment/Service 保持 Ready,root 继续 `prune: false` |
| Helm migration | 选定 `gitea-actions` 作为第一个 Flux HelmRelease adoption:它不承载 Git、入口、DNS、证书、数据库或 secrets controller。live StatefulSet 与 Git 都使用 regular DinD,但 Helm 保存的 release values/manifest 仍是失败的 rootless 配置;接管先固定 chart `0.1.1` 并验证 live Pod spec 不变,升级另开 PR |
| Helm adoption stage | 为 `gitea-actions` 加入固定 chart `0.1.1` 的 HelmRepository、values ConfigMap 和 `suspend: true` HelmRelease;第一阶段只让 Flux 登记对象,确认 source 与固定 chart render 后再解除 suspend,失败策略使用 `RetryOnFailure` 以避免回滚到 stored rootless manifest |
| Helm adoption activate | 第一阶段合并后 Flux source 与子 Kustomization 均 Ready,Helm release 仍为 revision 1,runner Pod 未 rollout;再次确认固定 chart 的完整 render 对 live 集群为零差异后,第二阶段移除 `suspend`,允许 Flux 修正 Helm 存储状态并开始 drift detection |
| Gitea adoption stage | 开始用 Flux 接管关键 `gitea` release:固定现有 chart `12.5.3`,以 `suspend: true` 登记 HelmRelease、source、values ConfigMap 和现有 HTTPRoute,子 Kustomization 保持 `prune: false`;现有 OIDC Secret 继续只引用不覆盖,其尚未进入 OpenBao/ESO 的缺口独立跟踪 |
| Gitea adoption activate | 第一阶段合并后 source、子 Kustomization 和 HTTPRoute 均 Ready,Helm release 仍为 revision 14,Gitea Pod 未重建或重启;再次确认固定 chart 对 live 业务资源零差异后移除 `suspend`,允许 Flux 修正 Helm 存储状态并启用 drift detection |
| Gitea upgrade plan | 规划两跳升级:chart `12.6.0` + 显式 Gitea `1.26.4`,再到 chart `12.7.0` + 显式 Gitea `1.27.3`;每个 minor 都先以 suspended desired state 合并、停机建立 CNPG/PVC 一致回滚点,再用独立 PR 激活。当前 CNPG 无连续备份、local-path PVC 无 snapshot class,因此禁止无备份直接触发数据库 migration |
| Gitea 1.26 preparation | 将第一跳目标写入 Git:chart 固定为 `12.6.0`、rootless 镜像显式固定为 `1.26.4`,同时重新设置 HelmRelease `suspend: true`;该准备 revision 合并后只更新 desired state,不触发 Pod replacement 或数据库 migration |
| Gitea 1.26 backup | 预拉取 `1.26.4-rootless` 后,在 HelmRelease suspended 状态将 Gitea scale 到 0;生成并校验 539355-byte CNPG custom dump、2609188-byte PVC tar 和 2848667-byte GPG encrypted bundle,随后恢复旧版 `1.25.5` 并验证内外 API 与 Flux source。按明确决定不上传 OCI,本阶段接受只有节点本地回滚点的风险 |
| Gitea 1.26 activation | 停机一致备份门槛完成后,激活变更只移除 HelmRelease 的 `suspend`;chart `12.6.0`、显式 `1.26.4-rootless` image、values、数据库与 PVC 均保持已 review 的准备状态 |
| Gitea 1.26 result | Flux 以 Helm revision 16 成功完成 chart `12.6.0` / Gitea `1.26.4-rootless` 的 Recreate upgrade 和 migration 323–330;Pod 内/统一域名 API、临时 branch push/delete、Flux source 及 main/smoke 的全部 CI jobs 均通过,Pod 在约 15 分钟采样中保持零重启,Authelia OIDC init 同步与浏览器交互式管理员登录也已确认成功 |
| Gitea 1.27 preparation | 预拉取 `1.27.3-rootless` 并将第二跳 desired state 原子设置为 chart `12.7.0`、显式 image `1.27.3` 和 `suspend: true`;合并只暂停并登记目标,不执行 migration,激活前必须从当前 1.26.4 数据建立新的配套回滚点 |
| Gitea 1.27 activation | 按明确决定跳过新的 1.26.4 数据库/PVC 备份,激活变更只移除 HelmRelease 的 `suspend`;接受 migration 失败后不能无损回退到 1.26.4 的风险,现有 1.25.5 本地备份仅能作为会丢失第一跳后状态的灾难恢复点 |
| CI runner network | 修复 Actions job 容器访问 GitHub 超时:k3s Pod MTU 为 1450,而 DinD 动态 bridge 默认为 1500;为 Docker daemon 固定 `--mtu=1450`。隔离测试证明相同 curl 镜像在默认 bridge 超时、在 MTU 1450 bridge 下访问 GitHub 与 API 均约 0.1 秒成功 |
### Incident: Gitea 备份后的恢复命令被 stdin 校验阻塞
最初把本地 custom-format dump 通过 `kubectl exec -i` 输送给 CNPG Pod 内的
`pg_restore --list`;远端 stdin 没有正常结束,组合脚本因此停在校验步骤,尚未执行
后面的 scale-up。Deployment 保持预期的 0,没有失败 Pod 或数据写入。发现后终止
会话、先恢复 Gitea,再把 dump 临时复制到 CNPG 可写数据卷完成校验并立即删除。
旧版 Gitea 恢复后内外 API 和 Flux source 均正常;后续 runbook 不再把 stdin 管道与
恢复命令放进同一个 shell transaction。
`Carried forward`: complete the two-stage zero-change `gitea` HelmRelease
adoption, migrate its remaining manual OIDC Secret to OpenBao/ESO, then upgrade
Gitea and add credential-free PR plan output before ordering the remaining Helm
migrations by dependency and blast radius. Root Flux prune remains disabled
until brownfield ownership is audited.
## 2026-09-09
**Recorded the brownfield GitOps/IaC redesign before changing live infrastructure.**
@@ -26,14 +62,31 @@ What changed in this homelab, when, and why. Newest first.
| secrets | Recorded that ESO 2.8.0, five ExternalSecrets and the scoped OpenBao Kubernetes-auth path already exist; the next gate is live recovery testing and migration of any remaining manual Secrets |
| Terraform | Recorded Gitea 1.27 State Registry as the preferred candidate for local roots after version and recovery testing; the OCI recovery root remains in OCI Object Storage to avoid a home-control-plane dependency loop |
| cleanup | Removed the retired NapCat tree, the Contour and Kanidm archive trees, and seven generated Terraform plan files before establishing the clean Git baseline; plans may embed complete state and remain globally ignored |
| CI | Added a review-first Gitea Actions runner bootstrap: official actions chart 0.1.1, pinned runner 2.3.0, one persistent Kubernetes runner with capacity four and rootless DinD, plus an ESO reference to a repository-scoped registration token in OpenBao. It is not deployed until the PR is merged |
| CI | Added a review-first Gitea Actions runner bootstrap: official actions chart 0.1.1, pinned runner 2.3.0, one persistent instance-scoped Kubernetes runner with capacity four, plus an ESO reference to its registration token in OpenBao. The first deployment proved that rootlesskit is blocked by the node's AppArmor unprivileged-userns policy; because the chart requires privileged DinD in either mode, the reviewed fix uses regular DinD instead of weakening the host-wide policy. The runner image intentionally carries neither `uv` nor Terraform: Terraform uses its versioned setup action, while `uv` is pinned and installed from official PyPI because the nested job network reaches PyPI but times out against the GitHub API queried by `setup-uv`. Ansible installs only `ansible-core` in its tool venv and puts declared Galaxy collections in a shared path visible to ansible-lint; installing the `ansible` meta-package had made Galaxy falsely skip that shared installation |
| identity | Declared the Samba AD `gitea-admins` group with `panxiao81` as its initial member. Gitea already maps this OIDC group to site administrators; the local `gitea_admin` account remains as break-glass access |
| docs | Reconciled the redesign and CI status with reality: the Gitea remote, instance-scoped runner, OpenBao-projected registration token, green Stage 1 and Flux bootstrap are live; off-site mirroring and recovery verification remain pending |
| 协作规范 | 在 `AGENTS.md` 中明确:homelab 向 `git.ddupan.top` 提交的 commit message、PR、issue 与项目文档默认优先使用中文,同时保留必要的英文技术标识符 |
| k3s | 在本机逐级从 `v1.33.6+k3s1` 升级到 `v1.33.13+k3s2`、`v1.34.11+k3s1`、`v1.35.8+k3s1`,最终到 `v1.36.4+k3s1`;每一级均建立 SQLite/server 冷备份并验证节点、工作负载、PVC、Gateway、DNS 与 Gitea。k3s 每次重启都会覆盖 CoreDNS 的手工 `serve_stale`,已按 `platform/k3s/Corefile.desired` 恢复 |
| GitOps | Flux `v2.9.5` 的四个核心 controller 已上线;集群内只读 Gitea source 与 `prune: false` 的 root Kustomization 均在合并 revision `aaa54a1` 上 Ready,完成了首个 pull reconciliation 闭环 |
| cleanup | 已把 `bao-acme` HTTP-01 solver 改到 Envoy Gateway 的明文 listener,并删除不再承载流量的 Contour namespace、provisioner、RBAC、GatewayClass 和全部 `projectcontour.io` CRD;Envoy Gateway、证书、DNS 与 Gitea 复查正常 |
| GitOps canary | 加入由 Flux 部署到独立 `gitops-canary` namespace 的 `http-echo` Deployment 和 Service;历史 Contour HTTPRoute 明确排除在 Kustomization 之外,初始保持 `prune: false` |
| prune 验证 | 为 `http-echo` 加入无业务依赖的 `flux-prune-canary` ConfigMap;先在 `prune: false` 下确认 Flux inventory,后续通过独立 PR 删除并仅为 canary 开启 prune |
| prune 验证第二阶段 | 第一阶段已确认 `flux-prune-canary` 带 Flux ownership 标签并进入 `http-echo` inventory;从 Git 删除该测试对象,同时仅为 `http-echo` 开启 `prune: true`,root 继续保持 `prune: false` |
| prune 验证修正 | 第二阶段证明“同一 revision 开启 prune 并删除旧对象”不会回收该对象:Flux 已从 inventory 移除它,但 live ConfigMap 保留。将 ConfigMap 在已经生效的 `prune: true` 下重新纳管,下一 revision 只做删除 |
| prune 验证最终阶段 | 已确认 `prune: true` 生效且测试 ConfigMap 重新进入 Flux inventory;本次只从 Git 删除该对象,不修改 canary 或 root 的 prune 设置,用于完成精确垃圾回收验证 |
### Incident: prune 启用与对象删除放在同一 revision
测试把 `http-echo` 从 `prune: false` 改为 `true` 的同时从 Git 删除测试
ConfigMap。Flux 按新 revision 更新了 inventory,但没有删除按旧设置管理的 live
对象,导致 ConfigMap 成为 inventory 之外的残留。没有业务影响。修正方式是在
`prune: true` 已经生效后先重新纳管对象,再用下一 revision 单独删除。
`Carried forward`: re-verify OpenBao/ESO recovery and remaining Secret inventory;
configure a Git remote and off-site mirror; confirm the running Gitea version;
configure an off-site Git mirror; plan the Gitea upgrade beyond 1.25.5;
take an encrypted independent OCI state copy before enabling bucket versioning;
reconstruct the missing root to a zero-change plan; bootstrap Gitea Actions and Flux on a low-risk
service; then move Tunnel origins to Envoy one hostname at a time.
reconstruct the missing root to a zero-change plan; add a low-risk Flux canary workload;
then move Tunnel origins to Envoy one hostname at a time.
## 2026-08-15
+9 -3
View File
@@ -132,9 +132,10 @@ recovered, so `.vault_pass.gpg` is the authoritative recovery path.
## Working rules
- **Record changes in `CHANGELOG.md`.** One dated section per day, newest first; incidents
get their own subsection. Traps and procedures belong *here* in CLAUDE.md, not there —
the changelog is for humans reading what changed.
- **Do not update `CHANGELOG.md`.** It is a frozen historical snapshot; requiring every PR
to append to one shared text file caused needless conflicts and duplicated Git/PR history.
Put durable service state and operational knowledge in the component README or runbook,
agent-facing traps here, and let commits/PRs record individual changes.
- **Verify, don't assert.** Check the end state (`pvesm status`, `linstor node list`,
`kubectl get pod`, `show ip route`) rather than trusting that a command "should have" worked.
Several confident diagnoses in this repo's history were wrong until measured.
@@ -149,6 +150,11 @@ recovered, so `.vault_pass.gpg` is the authoritative recovery path.
all-clear. Use `git check-ignore --no-index` and `git rm --cached` to actually remove it.
- **A `.tfplan` is a zip containing a full `tfstate`.** It walks straight past `*.tfstate`
ignore rules. Ignore `*.tfplan` everywhere.
- **SPIRE CLI JSON can be an array of response blocks.** `spire-agent api fetch jwt
-output json` in 1.15.3 returns blocks containing `svids` and `bundles`. Capture stdout
privately and type-check before extracting fields; `list(response)` prints full tokens
when the response is already an array. Never inspect credential payloads by printing
their containers, and never put fetched JWTs in command arguments or Pod logs.
- **Quoting does not survive two ssh hops.** `ssh pve1 "ssh pve3 'cmd | qm monitor 103'"`
loses the inner quotes — ssh re-joins argv with spaces, so the pipeline splits and the
tail runs on the **jump host**. It fails silently if you discard stderr: a `screendump`
+46
View File
@@ -0,0 +1,46 @@
# Gitea
Gitea 使用外部 CloudNativePG 数据库和现有 `gitea-shared-storage` RWO PVC,入口由
Envoy Gateway HTTPRoute 提供。Helm chart 自带的无 class Ingress 暂时保留以确保
首次接管零变化;清理该 Ingress 与升级 chart 必须使用后续独立 PR。
## Flux 接管
初始接管 release 是 `gitea-12.5.3`(Gitea `1.25.5`)。接管分为两个 PR:第一阶段创建
固定版本且 `suspend: true` 的 HelmRelease,只让 Flux 登记对象;确认 source Ready
并重新验证完整 chart render 后,第二阶段解除 suspend。第一阶段已经确认 source、
子 Kustomization 和 HTTPRoute 均 Ready,Helm revision 与 Gitea Pod 未变化;第二次
live diff 仍只有不常驻的 test hook Pod。子 Kustomization 和集群 root 均保持
`prune: false`。
数据库密码已经由 External Secrets Operator 从 OpenBao 投射到 `gitea-db`。OIDC
client secret 仍是历史手工 Secret `gitea-oidc-secret`,本次接管只引用、不覆盖它;
将剩余 Secret 迁移到 OpenBao 是独立的后续工作。
Gitea 是 Flux GitRepository 的上游。升级或重启期间 Git source 暂时不可用不会删除
已经应用的资源;Gitea 恢复后 Flux 会继续同步。任何会改变 Pod template、数据库迁移
或 PVC identity 的变更都不得与首次接管合并。
跨 minor 的执行顺序、停机一致备份和失败恢复步骤见
[`../../docs/gitea-upgrade-plan.md`](../../docs/gitea-upgrade-plan.md)。
当前 release 是 chart `12.7.0` / Gitea `1.27.3-rootless`。两次跨 minor migration、
API、OIDC、Git/Flux 和 runner 均已验证。第二跳按明确决定跳过了新的 1.26.4 停机
一致回滚点;现有 1.25.5 本地备份没有 OCI 或异机副本,因此只作为会丢失后续状态的
灾难恢复点。
## 后续升级默认策略
已连续验证两次 Flux 驱动的跨 minor Recreate upgrade,后续常规 patch/minor 升级
不再默认执行长时间观察、临时 branch push/delete、重复 CI、逐条 migration 日志审计
或每个 minor 的停机备份。默认只需要:
1. 固定 chart 和实际 `image.tag`,阅读与本配置相关的 breaking/security notes;
2. Helm/Kustomize render 通过 review;
3. 合并后确认 HelmRelease `UpgradeSucceeded`、Pod Ready 且没有 CrashLoop;
4. API 返回目标版本,并简单确认 OIDC 登录和一次正常 Git 操作。
只有变更数据库后端或存储布局、rootless 模式、PVC identity、部署策略、重大 chart
结构,或者 release notes 指出相关 breaking migration 时,才恢复停机一致备份、分阶段
suspend、详细日志审计和扩展验收。出现启动失败或 migration error 时也立即升级为完整
故障流程。
+39
View File
@@ -0,0 +1,39 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: gitea
namespace: gitea
spec:
chart:
spec:
chart: gitea
interval: 1h
sourceRef:
kind: HelmRepository
name: gitea-charts
version: 12.7.0
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: gitea
targetNamespace: gitea
timeout: 15m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
# Keep the patch release explicit because chart 12.7.0 defaults to 1.27.0.
# This override is in the suspended HelmRelease itself so chart, image and
# suspension are applied atomically; changing the watched values ConfigMap in
# the same revision could otherwise trigger reconciliation first.
values:
image:
rootless: true
tag: "1.27.3"
valuesFrom:
- kind: ConfigMap
name: gitea-values
+8
View File
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: gitea-charts
namespace: gitea
spec:
interval: 1h
url: https://dl.gitea.com/charts/
+18
View File
@@ -0,0 +1,18 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
annotations:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: gitea-values
namespace: gitea
files:
- values.yaml=gitea-values.yaml
resources:
- helmrepository.yaml
- helmrelease.yaml
- httproute.yaml
+21 -16
View File
@@ -1,22 +1,27 @@
# HTTP Echo (Gateway API workload)
# HTTP Echo(Flux GitOps canary)
**Purpose**
- Preserve a tiny HTTP echo `Deployment + Service` as a future GitOps canary.
- The current `HTTPRoute` still references the retired `contour-gateway`; do not
apply this folder until it is migrated and reviewed against Envoy Gateway.
这是 Flux 首个低风险工作负载,用于验证 PR 合并后的自动部署、健康检查和漂移修复。
Flux 将它部署到独立的 `gitops-canary` namespace。
**Resources**
| File | Description |
| 文件 | 说明 |
| --- | --- |
| `deployment.yaml` | Two replicas of `hashicorp/http-echo` returning `hello from contour gateway`. |
| `service.yaml` | ClusterIP service on port 80 targeted by the route. |
| `httproute.yaml` | Gateway API `HTTPRoute` that targets `contour-gateway` and the `http-echo` service. |
| `deployment.yaml` | 两个 `hashicorp/http-echo` 副本。 |
| `service.yaml` | 只在集群内可达的 ClusterIP Service。 |
| `kustomization.yaml` | Flux 实际构建入口;明确排除历史 HTTPRoute。 |
| `httproute.yaml` | 保留的历史 Contour 示例,**不在 Kustomization 中,不会部署**。 |
**How to verify**
## 验证
Migration and verification are intentionally deferred. Update `parentRefs` to the
reviewed Envoy Gateway and choose a hostname covered by its listener before apply.
合并 canary PR 后不手动 apply。等待 Flux 自动创建资源:
**Notes**
- This README records the old test intent; the retired Contour manifests are
available only in legacy Git history.
```bash
sudo k3s kubectl -n flux-system get kustomization http-echo
sudo k3s kubectl -n gitops-canary get deployment,service,pod
```
Deployment 漂移修复已经验证:手动把 replicas 改成 1 后,Flux 能按 Git 恢复为
2。删除验证使用无业务依赖的 `flux-prune-canary` ConfigMap。首次尝试在同一个
revision 中同时开启 prune 并删除对象,Flux 更新了 inventory 但保留了 live 对象。
因此先在已经生效的 `prune: true` 下重新纳管 ConfigMap,再用下一 revision 只删除
对象。最终删除阶段不再修改 prune,确保可以准确验证垃圾回收。root Kustomization
始终保持 `prune: false`,brownfield 资源不会进入此次删除范围。
+1 -1
View File
@@ -19,6 +19,6 @@ spec:
- name: http-echo
image: hashicorp/http-echo:0.2.3
args:
- '-text=hello from contour gateway'
- '-text=hello from flux gitops canary'
ports:
- containerPort: 5678
+5
View File
@@ -0,0 +1,5 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- deployment.yaml
- service.yaml
+22 -2
View File
@@ -11,7 +11,9 @@
| `helm.sh` | Installs or upgrades the SeaweedFS release. |
**Install**
1. Set real S3 access and secret keys in `values.yaml`.
1. 在 OpenBao `kv/k8s/seaweedfs-s3` 维护基础 S3 配置;zot 凭据单独以
`kv/k8s/zot-s3` 为唯一来源。ESO 合成为 `seaweedfs-s3-config`,详见下文。
不要把真实 AK/SK 放进 `values.yaml`。
2. Apply the manifests:
```bash
bash ~/services/apps/seaweedfs/helm.sh
@@ -31,4 +33,22 @@
**Notes**
- The chart manages master, volume, filer, S3, and admin components.
- The chart-managed S3 secret uses the current AK/SK pair for the admin user.
- The filer uses the ESO-managed `seaweedfs-s3-config` Secret for static S3 identities.
## zot 制品存储
`zot` bucket 专用于 [zot Registry](../zot/README.md),OCI 数据位于 `registry/`
前缀。静态身份 `zot` 只有该 bucket 的 Read/Write/List/Tagging 权限,凭据唯一来源为
Bao `kv/k8s/zot-s3` 的 `access_key` / `secret_key`,同时供 zot consumer 和
SeaweedFS 服务端使用。
[ExternalSecret 模板](../../platform/external-secrets/externalsecrets.yaml) 保留
`kv/k8s/seaweedfs-s3` 的原有身份及其他配置,再追加 zot 身份与限定 bucket 的权限。
基础配置当前版本不保存 zot AK/SK;旧 KV 版本历史仍保留。新增其他身份时使用
KV compare-and-set 保留已有内容,不覆盖 Terraform 或其他应用的 AK/SK。
不要直接编辑生成的 Kubernetes Secret。该 ExternalSecret 已单独应用到集群,
目前仍未加入 ESO 的 Flux Kustomization,遵循该组件现有 ownership 边界。
运行版本 `4.22` 可在 Secret volume 更新后向 filer/内嵌 S3 的 `weed` 进程发送
SIGHUP,重新加载静态配置,无需重启共享 S3 服务。本次接入没有启用 SeaweedFS
OIDC/STS;SPIRE 认证发生在 zot 的客户端入口。
+123
View File
@@ -0,0 +1,123 @@
# zot OCI Registry
内网入口为 `https://zot.ad.ddupan.top`。使用官方 Helm chart `0.1.124`,运行
zot `v2.1.21`,镜像固定到官方 linux/amd64 digest。
## 存储与凭据
制品、manifest 和 OCI layout 保存在现有 SeaweedFS 的 `zot` bucket,前缀为
`registry/`,S3 endpoint 为 `https://s3.ad.ddupan.top`。**不创建 PVC**;chart 的
`/var/lib/registry` 是 `emptyDir`,仅用于运行时本地工作数据。
首期单副本,关闭跨仓库 dedupe,不额外部署 Redis/DynamoDB 缓存。保留 zot GC,
暂不配置自动删除已发布版本的 retention policy。增加副本、启用 dedupe 或搜索等
扩展前,需要重新检查共享元数据与缓存的持久化要求。
凭据链路:
```text
OpenBao kv/k8s/seaweedfs-s3
→ 原有 S3 身份及基础配置 ─┐
├→ ESO 模板 → seaweedfs/seaweedfs-s3-config
OpenBao kv/k8s/zot-s3 ────┘ → 完整 s3.config
└→ ESO → zot/zot-s3 → zot 的 AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY
```
`kv/k8s/zot-s3` 是 zot AK/SK 的唯一维护来源。基础配置保留原有身份及其他字段,
不再保存 zot 凭据副本;[SeaweedFS ExternalSecret](../../platform/external-secrets/externalsecrets.yaml)
使用 ESO v2 模板追加 zot 身份。两个 Kubernetes Secret 都是自动生成的消费副本,
不手工编辑。Bao 的旧版本历史保留,回滚基础配置时模板也会替换其中的旧 zot 身份。
专用 S3 身份只有 `Read:zot`、`Write:zot`、`List:zot`、`Tagging:zot`,不能读取
Terraform 的 `tfstate` bucket。`kv/k8s/zot-s3` 的字段是 `access_key` 和
`secret_key`。AK/SK 不进入 Git、Helm values 或 CI;这里仍是静态 S3 凭据,尚未
接入 SPIRE/STS。
本次归一没有轮换密钥,生成的完整配置与归一前语义一致。当前模板只有一组 zot
凭据,尚未实现新旧密钥重叠轮换。后续轮换只修改 `kv/k8s/zot-s3`,但仍需协调
两个 ExternalSecret 同步:确认 SeaweedFS Secret volume 更新后向 filer 的
`weed` 进程发送 SIGHUP,再确认 zot Secret 更新并重启 zot(环境变量不会热更新)。
两端异步更新期间可能短暂认证失败;需要无中断轮换时先扩展模板支持新旧凭据重叠。
## SPIRE 认证和授权
| 参数 | 值 |
|---|---|
| issuer | `https://spire-oidc.ad.ddupan.top` |
| JWT audience | `zot` |
| subject | `spiffe://ddupan.top/` 下的 workload SPIFFE ID |
| token endpoint | `https://zot.ad.ddupan.top/zot/auth/token` |
| 当前权限 | 受信身份可以读取所有仓库;没有常驻写入或删除授权 |
zot 通过已配置的 issuer discovery/JWKS 验证 JWT-SVID,再以 `sub` 作为授权身份。
不接受任意 issuer,不关闭 TLS/issuer 验证。新的 Kata CI 负责取得并更新自己的
JWT-SVID;确认其身份命名后,再添加针对具体 repository 的 `create`/`update`
授权,不能把整个 trust domain 都授予写权限。
现阶段拉取也需要 JWT-SVID。原定内网匿名拉取尚未启用:zot `v2.1.21` 的
OIDC Bearer middleware 会在授权阶段之前拒绝无 token 请求,单独增加
`anonymousPolicy` 无法解决。匿名读取与 SPIRE 写入共存需后续单独验证方案。
已有 SPIRE 身份的进程可以通过 Workload API 获取 `aud=zot` 的 JWT-SVID,然后
通过 `docker login` 或 `crane auth login` 的 `--password-stdin` 交给 Registry。
用户名可以使用 `zot`,实际权限取自已验证 JWT 的身份。使用独立、权限为 `0700`
的临时 `DOCKER_CONFIG`,结束后删除;不要开启 shell tracing,不要打印 token,
不要把 token 放进命令参数。token 接口不会延长 SVID 有效期。
## 部署与网络
- 官方 chart 管理 Deployment、Service、ConfigMap 和 HTTPRoute。
- `persistence: false`,Service 为 ClusterIP,TLS 由已有 Envoy Gateway 的
`https` listener 与内网通配符证书终止。
- 仅配置 Samba AD 内网 DNS;不创建公网 DNS 或 Cloudflare Tunnel route。
- NetworkPolicy 只允许现有 Envoy Gateway 数据面访问 zot 的 5000 端口。
- HTTPRoute 只暴露 `/v2/` 和 `/zot/auth/token`,不暴露内部健康检查或管理端点。
- namespace 使用 restricted PodSecurity,容器非 root、只读根文件系统。
首次已按用户授权从本地执行 `kubectl apply -k apps/zot`,由集群 Helm controller
安装。`clusters/homelab/apps/zot.yaml` 是 GitOps composition;对应文件合并进入
Flux 跟踪分支后,才由根 Kustomization 持续管理,不能把未提交的本地部署写成
已完成 Git 接管。
检查与渲染:
```bash
helm template zot --repo https://zotregistry.dev/helm-charts \
--version 0.1.124 --namespace zot -f apps/zot/values.yaml --skip-tests
sudo k3s kubectl -n zot get helmrelease,pods,externalsecret,httproute
sudo k3s kubectl -n zot get pvc
```
上游 chart 的 Helm test Pod 不满足本 namespace 的 restricted 策略,也没有
SPIRE 凭据,因此不运行默认 `helm test`;使用下述真实身份验收。
## 验收与恢复
验收使用独立临时 Pod,通过 SPIFFE CSI socket 和真实 Workload API 取得 JWT-SVID,
没有修改现有 runner。仅在初始化 `verification/smoke:spire-s3` 测试镜像时临时
授予该测试身份针对该仓库的写权限;完成后必须撤回 HelmRelease override,并删除
临时 Pod、ServiceAccount 与 ClusterSPIFFEID。
验收项目:有效 SVID + crane pull、manifest digest 一致、错误 audience、错误
signature、过期 token、无凭据写入、跨仓库写入、只读身份写入和删除拒绝;另外检查
Pod 重建后镜像仍可拉取,以及 S3 身份不能访问 `tfstate`。
2026-09-14 已完成上述验收:HelmRelease Ready、HTTPRoute Accepted/ResolvedRefs,
DNS 第二次 Ansible check 为 `changed=0`;一分钟真实 JWT-SVID 到期后返回 401。
临时写权限已移除。测试镜像可供后续 CI 验证拉取:
```text
zot.ad.ddupan.top/verification/smoke:spire-s3
sha256:b8d3b977a1235022759470903dab4a46b7cf8107958624f1f76a323eabe37c5e
```
它是仅含验证文本的 OCI 测试镜像,没有可执行入口,不用于运行服务。
Registry 恢复需要完整的 SeaweedFS bucket 数据、Bao 专用凭据和此目录配置。
zot 的临时目录不是制品备份。独立异机/离线备份尚未在本次部署中建立;不能把同一
SeaweedFS 内的数据副本当作独立灾备。重装 zot 不得删除 `zot` bucket。
参考:[官方 Kubernetes 安装](https://zotregistry.dev/v2.1.21/install-guides/install-guide-k8s/)、
[S3 存储](https://zotregistry.dev/v2.1.21/articles/storage/)、
[OIDC workload identity](https://github.com/project-zot/zot/blob/v2.1.21/examples/README-OIDC-WORKLOAD-IDENTITY.md)。
+22
View File
@@ -0,0 +1,22 @@
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: zot-s3
namespace: zot
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: openbao
target:
name: zot-s3
creationPolicy: Owner
data:
- secretKey: access_key
remoteRef:
key: k8s/zot-s3
property: access_key
- secretKey: secret_key
remoteRef:
key: k8s/zot-s3
property: secret_key
+30
View File
@@ -0,0 +1,30 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: zot
namespace: zot
spec:
chart:
spec:
chart: zot
version: 0.1.124
interval: 1h
sourceRef:
kind: HelmRepository
name: zot
releaseName: zot
interval: 30m
timeout: 5m
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: zot-values
+8
View File
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: zot
namespace: zot
spec:
interval: 1h
url: https://zotregistry.dev/helm-charts
+18
View File
@@ -0,0 +1,18 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- serviceaccount.yaml
- external-secret.yaml
- helmrepository.yaml
- helmrelease.yaml
- networkpolicy.yaml
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: zot-values
namespace: zot
files:
- values.yaml=values.yaml
+7
View File
@@ -0,0 +1,7 @@
apiVersion: v1
kind: Namespace
metadata:
name: zot
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/enforce-version: v1.36
+22
View File
@@ -0,0 +1,22 @@
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: zot-ingress
namespace: zot
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: zot
policyTypes: [Ingress]
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: envoy-gateway-system
podSelector:
matchLabels:
gateway.envoyproxy.io/owning-gateway-name: eg
gateway.envoyproxy.io/owning-gateway-namespace: envoy-gateway-system
ports:
- protocol: TCP
port: 5000
+6
View File
@@ -0,0 +1,6 @@
apiVersion: v1
kind: ServiceAccount
metadata:
name: zot
namespace: zot
automountServiceAccountToken: false
+147
View File
@@ -0,0 +1,147 @@
# 官方 chart 0.1.124 / zot v2.1.21;制品与 manifests 保存在 SeaweedFS S3。
# persistence=false 仅保留 chart 的 emptyDir,不创建 PVC。
# 首期关闭跨仓库 dedupe,不额外引入 Redis/DynamoDB 持久缓存。
replicaCount: 1
image:
repository: ghcr.io/project-zot/zot
tag: v2.1.21@sha256:8258443838e95989c13c891f78a02bc1c391b5a00591ffef24cb8c17cde28038
persistence: false
strategy:
type: Recreate
serviceAccount:
create: false
name: zot
service:
type: ClusterIP
port: 5000
mountConfig: true
mountSecret: false
secretFiles: {}
configFiles:
config.json: |
{
"distSpecVersion": "1.1.1",
"storage": {
"rootDirectory": "/var/lib/registry",
"dedupe": false,
"gc": true,
"gcDelay": "24h",
"gcInterval": "24h",
"storageDriver": {
"name": "s3",
"region": "us-east-1",
"regionendpoint": "https://s3.ad.ddupan.top",
"bucket": "zot",
"rootdirectory": "/registry",
"secure": true,
"skipverify": false,
"forcepathstyle": true
}
},
"http": {
"address": "0.0.0.0",
"port": "5000",
"externalUrl": "https://zot.ad.ddupan.top",
"compat": [
"docker2s2"
],
"auth": {
"bearer": {
"realm": "https://zot.ad.ddupan.top/zot/auth/token",
"service": "zot.ad.ddupan.top",
"oidc": [
{
"issuer": "https://spire-oidc.ad.ddupan.top",
"audiences": [
"zot"
],
"claimMapping": {
"username": "claims.sub",
"validations": [
{
"expression": "claims.sub.startsWith('spiffe://ddupan.top/')",
"message": "SPIFFE trust domain mismatch"
}
]
}
}
]
}
},
"accessControl": {
"repositories": {
"**": {
"defaultPolicy": [
"read"
]
}
}
}
},
"log": {
"level": "info"
}
}
env:
- name: AWS_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: zot-s3
key: access_key
- name: AWS_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: zot-s3
key: secret_key
- name: AWS_EC2_METADATA_DISABLED
value: 'true'
podSecurityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: '1'
memory: 512Mi
extraVolumes:
- name: tmp
emptyDir:
sizeLimit: 128Mi
extraVolumeMounts:
- name: tmp
mountPath: /tmp
startupProbe:
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 60
httproute:
enabled: true
parentRefs:
- name: eg
namespace: envoy-gateway-system
sectionName: https
hostnames:
- zot.ad.ddupan.top
rules:
- matches:
- path:
type: PathPrefix
value: /v2/
- path:
type: Exact
value: /zot/auth/token
timeouts:
request: 900s
backendRequest: 900s
+51 -11
View File
@@ -1,15 +1,55 @@
# Homelab cluster
# Homelab 集群
This directory will become the Flux reconciliation entrypoint for the homelab
k3s cluster. It is intentionally documentation-only until a Git remote, CI
checks and a low-risk bootstrap workload have been verified.
这里是单节点 k3s 集群的 Flux reconciliation 入口。集群当前运行 Kubernetes
`v1.36.4+k3s1`,Flux 固定为 `v2.9.5`。
Planned reconciliation order:
## 首次 bootstrap
1. namespaces and CRDs;
2. shared platform controllers;
3. secret references and storage;
4. applications.
仓库经过 PR 审查并合并后,在本机从合并后的 `main` 执行:
Do not enable pruning for a path until its live resources and field ownership
have been audited.
```bash
sudo k3s kubectl apply -f clusters/homelab/flux-system/gotk-components.yaml
sudo k3s kubectl apply -f clusters/homelab/flux-system/gotk-sync.yaml
```
`GitRepository/flux-system` 通过集群内 Gitea Service 读取公开仓库,不需要长期
管理员 token,也不依赖 Cloudflare、公网 DNS 或 Envoy Gateway。Gitea 暂时不可用
时,已经应用的资源继续运行,Flux 在 Gitea 恢复后重新同步。
root Kustomization 从 `./clusters/homelab` 开始 reconciliation。初始设置
`prune: false`;在逐项审计现有资源和 field ownership 之前不得开启全局 prune。
计划中的 reconciliation 顺序:
1. namespaces 和 CRD;
2. platform controllers;
3. secret references 和 storage;
4. applications。
首次部署后的最低验证:
```bash
sudo k3s kubectl -n flux-system get pods
sudo k3s kubectl -n flux-system get gitrepositories,kustomizations
```
四个 controller、GitRepository 和 root Kustomization 都必须为 Ready,随后才能
通过单独 PR 引入低风险 canary workload。
## 当前状态
- Flux `v2.9.5`、GitRepository 和 root Kustomization 均为 Ready;
- `http-echo` canary 已验证 merge 后自动部署和 replicas 漂移修复;
- `http-echo` 的专用测试 ConfigMap 已在 `prune: true` 生效后重新纳管,并由下一
revision 自动删除;
- `http-echo` 保持 `prune: true`,root 保持 `prune: false`;
- `gitea-actions`、`gitea` 与 External Secrets 已由 Flux HelmRelease 接管,Gitea 已升级到 `1.27.3`;
- cert-manager 已固定现有 `v1.21.0` 并完成分阶段 Flux HelmRelease 接管;
- Envoy Gateway 已固定现有 `v1.5.6` 并完成分阶段 Flux HelmRelease 接管;
- OpenEBS 已固定现有 `4.4.0` 并完成分阶段 Flux HelmRelease 接管;
- VictoriaMetrics Operator 已固定现有 chart `0.66.2` 并完成分阶段 Flux HelmRelease
接管;Metrics、Logs、Traces 与 Grafana 也已统一完成 Flux 接管;
- External Secrets Operator 已固定 chart `2.8.0` 并完成分阶段接管;
- SPIRE 已按官方 hardened chart `0.30.2`(SPIRE `1.15.3`)声明,使用共享
PostgreSQL 与独立 signing-key PVC;首次上线和 OpenBao JWT-SVID PoC 尚待合并后验证;
- root Kustomization 与所有 brownfield 子 Kustomization 继续保持 `prune: false`。
+14
View File
@@ -0,0 +1,14 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: cert-manager
namespace: flux-system
spec:
interval: 10m
path: ./platform/cert-manager
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
wait: false
+14
View File
@@ -0,0 +1,14 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: envoy-gateway
namespace: flux-system
spec:
interval: 10m
path: ./platform/envoy-gateway
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
wait: false
@@ -0,0 +1,14 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: external-secrets
namespace: flux-system
spec:
interval: 10m
path: ./platform/external-secrets
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
wait: false
+14
View File
@@ -0,0 +1,14 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: gitea-actions
namespace: flux-system
spec:
interval: 10m
path: ./platform/gitea-runner
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
wait: false
+13
View File
@@ -0,0 +1,13 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: gitea
namespace: flux-system
spec:
interval: 10m
path: ./apps/gitea
prune: false
sourceRef:
kind: GitRepository
name: flux-system
wait: false
+20
View File
@@ -0,0 +1,20 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: http-echo
namespace: flux-system
spec:
healthChecks:
- apiVersion: apps/v1
kind: Deployment
name: http-echo
namespace: gitops-canary
interval: 10m
path: ./apps/http-echo
prune: true
sourceRef:
kind: GitRepository
name: flux-system
targetNamespace: gitops-canary
timeout: 3m
wait: true
+14
View File
@@ -0,0 +1,14 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: observability
namespace: flux-system
spec:
interval: 10m
path: ./platform/observability
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
wait: false
+14
View File
@@ -0,0 +1,14 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: openebs
namespace: flux-system
spec:
interval: 10m
path: ./platform/openebs
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
wait: false
+14
View File
@@ -0,0 +1,14 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: spire
namespace: flux-system
spec:
interval: 10m
path: ./platform/spire
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 15m
wait: false
+18
View File
@@ -0,0 +1,18 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: zot
namespace: flux-system
spec:
dependsOn:
- name: envoy-gateway
- name: external-secrets
- name: spire
interval: 10m
path: ./apps/zot
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 5m
wait: true
+19
View File
@@ -0,0 +1,19 @@
# Contour 退役清理
Envoy Gateway 已经承载全部现用 `Gateway` 和 `HTTPRoute`。Contour 在集群中仅剩
gateway provisioner、RBAC、旧 `GatewayClass` 和没有实例的 CRD,不再承载流量。
清理必须在本变更合并后进行,顺序如下:
1. 应用更新后的 `platform/cert-manager/clusterissuer-bao-acme.yaml`,把 OpenBao
HTTP-01 solver 改到 `envoy-gateway-system/eg` 的 `http` listener;
2. 确认 `Gateway/eg` 为 `Programmed=True`,所有现用 `HTTPRoute` 保持正常;
3. 删除 `GatewayClass/contour`;
4. 删除 `projectcontour` namespace;
5. 删除名称包含 `contour` 的遗留 ClusterRole/ClusterRoleBinding;
6. 在确认所有 Contour 自定义资源均为空后,删除 `projectcontour.io` 的五个 CRD;
7. 复查 Envoy Gateway、证书、DNS 和现用入口。
这些对象是 Flux 启用前留下的孤立资源,首次清理由本机 `kubectl` 完成。Flux 的
root Kustomization 初始保持 `prune: false`,不会借 bootstrap 顺带删除 brownfield
资源。
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,27 @@
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: flux-system
namespace: flux-system
spec:
interval: 1m
ref:
branch: main
timeout: 60s
url: http://gitea-http.gitea.svc.cluster.local:3000/panxiao81/homelab-infra.git
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: flux-system
namespace: flux-system
spec:
interval: 10m
path: ./clusters/homelab
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
wait: true
@@ -0,0 +1,5 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- gotk-components.yaml
- gotk-sync.yaml
+15
View File
@@ -0,0 +1,15 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- flux-system
- namespaces/gitops-canary.yaml
- apps/cert-manager.yaml
- apps/envoy-gateway.yaml
- apps/external-secrets.yaml
- apps/gitea.yaml
- apps/gitea-actions.yaml
- apps/http-echo.yaml
- apps/openebs.yaml
- apps/spire.yaml
- apps/observability.yaml
- apps/zot.yaml
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: gitops-canary
+19 -14
View File
@@ -1,6 +1,7 @@
# CI/CD — what we are building
Status: **design, partly built.** Stage 1 is live. The substrate is undecided.
Status: **partly built.** Stage 1、Kubernetes runner 与 Flux 已上线;credentialed
stages 尚未实现。
Started 2026-07-28.
## Goal
@@ -22,19 +23,21 @@ to make drift between this repo and reality visible when it happens.
| | |
|---|---|
| git | History since 2026-07-28. Four commits, **no remote yet** |
| stage 1 | Live and green — `yamllint`, `ansible-lint`, `terraform fmt`/`validate`. Configs tuned against a real run |
| gitea | 1.25.5, Actions **enabled**, `DEFAULT_ACTIONS_URL=github`. **No runner deployed**, so nothing executes |
| git | `homelab-infra` is hosted on the local Gitea; an independent off-site mirror is still missing |
| stage 1 | Gitea runner 上的 `yamllint`、`ansible-lint`、Terraform fmt/validate 已上线;三个 workflow 按路径触发,feature push 不再与 PR 事件重复运行 |
| gitea | 1.27.3,Actions 已启用;一个 instance-scoped Kubernetes runner 以 capacity 4 运行 |
| ansible | 33 roles across `infrastructure/proxmox/`, `infrastructure/samba-ad/`, `infrastructure/openbao/` |
| terraform | 4 roots, **local state**, each with **different interactive auth** (`bao login -method=oidc`, `az login`) |
| k8s | ~13 Helm releases, all deployed by hand |
| k8s | Flux 已接管 Gitea、Gitea Actions 与 http-echo canary;其余 brownfield release 逐项迁移 |
| secrets | 4 config files are gitignored because they embed live secrets, so their contents are **not** version controlled |
## The four stages
**Stage 1 — static. Built.**
`yamllint`, `ansible-lint`, `terraform fmt -check` / `validate -backend=false`.
No cluster, no credentials, no mutation, so it is safe on every push. It already
No cluster, no credentials, no mutation. YAML、Ansible 和 Terraform 各自按相关路径
触发;feature branch 只由 `pull_request` 检查,合并后再由 `main` push 检查,避免
同一 revision 因 branch push 和 PR 各跑一遍。它已经
found a real defect: `infrastructure/proxmox/ansible/` had no `requirements.yml` at all, so a
fresh checkout could not reproduce its collections.
@@ -90,11 +93,13 @@ Gitea Actions is the CI control plane. It integrates directly with repository
permissions and status checks and preserves GitHub Actions workflow syntax.
The bootstrap worker is the official Gitea Runner chart in Kubernetes: one
persistent StatefulSet Pod, rootless Docker-in-Docker and capacity four. Job
persistent StatefulSet Pod, Docker-in-Docker and capacity four. Job
containers are dynamic, while the runner and its Docker daemon remain resident.
Rootless DinD still needs a privileged Pod to establish its user namespace, so
the runner is repository-scoped and restricted to trusted workflows. See
`platform/gitea-runner/`.
The chart's DinD container is privileged in both modes. Rootless mode is blocked
by the node's AppArmor unprivileged-userns policy, so regular DinD avoids weakening
that host-wide policy without pretending the Pod has a stronger isolation boundary.
The runner is instance-scoped and restricted to trusted repositories and workflows.
See `platform/gitea-runner/`.
This is not native pod-per-job execution. If stronger isolation becomes useful,
the runner's ephemeral registration and Gitea `workflow_job` webhook can later
@@ -113,7 +118,7 @@ Proxmox provider exists, so ephemeral Proxmox VMs would mean writing one.
Bootstrap dependencies and current status:
1. **Secret delivery exists.** OpenBao and External Secrets Operator already
synchronize five Secrets. Add the repository-scoped runner registration token
synchronize five Secrets. Add the instance-scoped runner registration token
at `kv/k8s/gitea-runner`; Git contains only its `ExternalSecret` reference.
2. **Git remote exists.** `homelab-infra` is hosted in Gitea. An off-cluster
read-only mirror remains required for disaster recovery.
@@ -169,10 +174,10 @@ the CI system itself.
## Sequencing
1. Externalise secrets — unblocks everything, valuable on its own
1. Externalise secrets — **partly complete**; ESO delivery works, recovery and the remaining inventory are pending
2. Capture `e5renew` and `rustfs` into the repo (`helm get values`)
3. Git remote
4. Pick the substrate; stand it up in an isolated namespace
3. Git remote — **complete locally**; off-site mirror pending
4. Pick the substrate; stand it up in an isolated namespace — **complete**
5. Stage 2, then stage 3 on **one** container-friendly role first
6. Flux on one low-stakes namespace (`http-echo` or `marker`)
7. Drift detection for Terraform and Ansible — scoped machine identities
+174
View File
@@ -0,0 +1,174 @@
# Gitea 1.25.5 → 1.27.3 升级计划
## 决策
现有 release 为 chart `12.5.3` / Gitea `1.25.5`,已经由 Flux HelmRelease 接管。
升级分两个独立阶段执行,不跨过中间 minor:
1. chart `12.6.0`,显式设置 `image.tag: 1.26.4`;
2. chart `12.7.0`,显式设置 `image.tag: 1.27.3`。
chart 当前默认 appVersion 分别只是 `1.26.1` 和 `1.27.0`,因此不能依赖默认镜像。
官方 `1.26.4` 修复了 1.26.3 的代码页回归并包含安全修复;`1.27.3` 又修复了
Actions fork PR 审批绕过等安全问题。两个 rootless 镜像标签都已经在官方 registry
验证存在。Gitea Runner `2.3.0` 已高于 1.27 推荐的 `2.0.0`,无需与服务端同时升级。
参考:
- [Gitea 升级说明](https://docs.gitea.com/installation/upgrade-from-gitea/)
- [Gitea 备份一致性说明](https://docs.gitea.com/1.26/administration/backup-and-restore/)
- [Gitea 1.26.4 发布说明](https://blog.gitea.com/release-of-1.26.3-and-1.26.4/)
- [Gitea 1.27.0 与 Runner 兼容说明](https://blog.gitea.com/release-of-1.27.0/)
- [Gitea 1.27.3 发布说明](https://blog.gitea.com/release-of-1.27.3/)
## 当前风险边界
- Gitea 使用外部单实例 CNPG;`gitea` 数据库约 19 MB,CNPG 没有配置连续备份,
`firstRecoverabilityPoint` 为空。
- `/data` 使用 `local-path` RWO PVC,实际数据约 6.2 MB;集群没有
VolumeSnapshotClass,因此不能把 CSI snapshot 当成回滚点。
- Gitea 是 Flux GitRepository 的上游。Gitea 停机时已应用的资源继续运行,但不能
依靠新的 Git commit 解救失败的 Gitea。
- 跨 minor 启动会执行数据库 migration。官方明确说明升级后的数据库不能由旧 minor
安全使用;回退必须同时恢复数据库和 `/data`。
- `gitea-oidc-secret` 仍是手工 Secret。本次只继续引用它,不与版本升级一起迁移。
## 已完成的静态检查
- 使用现有 `gitea-values.yaml` 渲染三个 chart,并从比较中排除 Secret 与 Helm test
hook;`12.5.3 → 12.6.0` 的业务资源变化只有 chart/app 标签和四处 Gitea 镜像。
- `12.6.0 → 12.7.0` 除同类版本变化外,只包含 Service template 文件重命名、字段
排版变化和 test hook 参数排版,没有新增或删除 live 业务对象。
- `docker.gitea.com/gitea:1.26.4-rootless` 与
`docker.gitea.com/gitea:1.27.3-rootless` 的 multi-arch manifests 均存在。
- 升级前没有 deprecation 或 error;启动会报告若干数据库 default 比较及内网明文 SMTP
warning,均为已有状态。HTTPRoute、内部 API、统一域名 API 与 Flux GitRepository
均 Ready。
## 每个 minor 的执行单元
两个阶段各自重复以下流程。1.26 验收完成前不得创建 1.27 的执行 PR。
### 1. 预拉取与 review
1. 在 HelmRelease 中显式设置目标 `image.tag`,同时把 chart 固定到该阶段目标版本,
并重新执行 Helm/Kustomize render。目标 chart、镜像与 `suspend: true` 必须在同一
HelmRelease 对象内原子提交;不得同时修改会触发 watch 的 values ConfigMap。
2. 在节点预拉取目标 rootless 镜像,避免维护窗口受 WAN 波动影响。
3. 创建 **保持 `spec.suspend: true`** 的准备 PR;合并并等 Flux 同步。此时 Git 只记录
目标版本,不执行 Helm action。
4. 从该准备 commit 创建只删除 `suspend` 的激活 PR,提前 push 到 Gitea,但暂不合并。
### 2. 建立一致回滚点
Gitea 官方要求停机备份才能保证数据库、仓库、LFS、附件和 metadata 一致。本环境
数据量很小,因此接受一个短维护窗口,不使用在线 `gitea dump` 冒充一致备份。
1. 确认 HelmRelease 已 suspended,记录 Deployment、ReplicaSet、Pod UID、Helm
revision、数据库大小和 PVC 路径。
2. 手工将 Gitea Deployment scale 到 0 并等待 Pod 消失。此时 runner 会暂时无法上报,
Flux source 会暂时 NotReady,但现有工作负载不受影响。
3. 对 CNPG 的 `gitea` 数据库执行 custom-format `pg_dump`;对 PVC host path 创建保留
owner/xattr 的 tar archive。两个文件写入节点上新建的、权限为 0700 的时间戳目录。
4. 生成 SHA-256,执行 `pg_restore --list` 和 `tar --list`,确认两份文件可读。
5. 用个人 GPG 公钥加密后复制一份到 OCI Object Storage(或另一台物理机);未验证
第二份副本前不得升级。
6. 将 Deployment scale 回 1,确认仍以旧版本启动,并验证 API、OIDC 登录、clone、
push 和 Actions runner online。
手工 scale 只用于取得一致备份;HelmRelease 此时 suspended,不会发生 drift repair。
备份结束后服务恢复,后续激活 PR 仍通过正常 Gitea review/merge,不需要在 Gitea 停机
时修改 Git 或 live HelmRelease。
### 3. 由 Flux 升级
1. review 并合并已经 push 的激活 PR;它唯一的行为变化应是删除 `suspend: true`。
2. 不手工 reconcile,观察 Flux 正常发现 revision、Helm upgrade 和 Gitea migration。
3. `Recreate` strategy 会先终止旧 Pod 再启动新 Pod,避免 RWO PVC 与 leveldb queue
lock 导致双 Pod deadlock。
4. 若 15 分钟内 HelmRelease 未 Ready,停止自动重试并进入回滚判断,不连续修改
values 猜测修复。
### 4. 每阶段验收
- HelmRelease `Ready=True`、`UpgradeSucceeded`,Helm revision 只增加一次;
- Deployment 和 Pod 使用目标 rootless 镜像,PVC UID/volume name 不变;
- Gitea `/api/v1/version` 从 Pod 内和 `https://git.ddupan.top` 返回目标版本;
- Authelia OIDC 管理员登录与 break-glass 本地管理员登录均有效;
- 对现有仓库完成 HTTPS clone、创建临时 branch、push、删除临时 branch;
- 当前仓库 Actions workflow 能入队并由 runner `2.3.0` 完成;
- Flux GitRepository 恢复 Ready 并能拉取升级 revision;
- package registry、Terraform package/state API(启用后)与附件/LFS 按实际使用情况抽查;
- 日志中没有 migration failure、panic、持续数据库错误或配置弃用警告;
- 稳定观察至少 15 分钟,再把该阶段记录为完成。
## 回滚
patch 版本原则上保持数据库结构兼容;本计划的两个步骤都是跨 minor,不能仅回退镜像。
若 migration 后目标版本无法健康启动:
1. 暂停 `Kustomization/gitea` 和 `HelmRelease/gitea`,停止 Gitea Deployment;
2. 保存失败现场的 Pod logs、Helm status 与 migration error;
3. 删除并从同一阶段备份恢复 `gitea` 数据库;
4. 清空并从配套 tar 恢复原 PVC 内容,保持 owner、mode 和 xattr;
5. 将 live HelmRelease 临时恢复到上一阶段 chart/image,并启动验证;
6. Gitea 恢复后,通过 PR 把 Git desired state 恢复到上一阶段,再恢复 Flux
Kustomization。不得让 Flux 在数据库尚未恢复时自动拉回新版本。
数据库 dump 与 PVC archive 是一个不可拆分的回滚点;禁止混用不同时间或不同阶段的
两份备份。
## 1.26.4 阶段执行记录
- 目标 `1.26.4-rootless` 镜像已经预拉取到 k3s 节点;
- Gitea 停写后生成了 custom-format PostgreSQL dump 与保留 owner/ACL/xattr 的 PVC
tar,两者分别通过 `pg_restore --list`、`tar --list` 和 SHA-256 校验;
- 本地回滚点位于
`/var/backups/homelab/gitea/20260910T060400Z-1.25.5/`,权限为 0700;其中另有使用
个人 GPG 公钥加密并通过 packet 检查的 bundle;
- 操作者明确选择不上传 OCI,因此本阶段接受只有节点本地副本的风险例外;
- 备份后旧版 `1.25.5` 已恢复,Pod 内与统一域名 API、Gitea API 和 Flux Git source
均验证正常。
激活后 Helm revision 16 以 chart `12.6.0` 成功部署
`docker.gitea.com/gitea:1.26.4-rootless`。Migration 323–330 完成;Pod 内和统一域名
API、临时 branch push/delete、Flux source,以及 main/smoke 的 YAML、Ansible、
Terraform CI 均通过。Pod 在约 15 分钟观察期内保持 Running、零重启;Authelia OIDC
配置 init 同步及浏览器交互式管理员登录也已确认成功。
## 1.27.3 阶段准备状态
- 目标 `1.27.3-rootless` 镜像已经预拉取到 k3s 节点;
- Git desired state 固定为 chart `12.7.0` 与显式 image `1.27.3`,并重新设置
`suspend: true`;
- 合并准备状态后必须从已经迁移的 1.26.4 数据重新建立一组数据库/PVC 回滚点,禁止
复用 1.25.5 备份作为 1.27 阶段的直接回滚点。
执行时操作者明确选择跳过上述 1.26.4 备份门槛并继续激活。该风险例外意味着 1.27
migration 后若失败,不能无损回到 1.26.4;现有本地 1.25.5 备份只可用于接受丢失
第一跳之后状态的灾难恢复。目标 1.27.3 镜像已预拉取,其余准备检查均已通过。
激活后 Helm revision 17 以 chart `12.7.0` 成功部署实际镜像 `1.27.3-rootless`,
migration 331–342 和全部 init containers 成功;Pod 内/统一域名 API、Git pull、临时
branch push/delete 与 Flux source 均通过,Pod Ready 且零重启。Chart metadata 显示
appVersion `1.27.0`,实际版本以固定 image 和 API 返回的 `1.27.3` 为准。
## 后续升级策略
本次连续两次跨 minor 升级证明现有 chart、外部 CNPG、rootless PVC 和 `Recreate`
组合工作稳定。后续常规 patch/minor 升级采用精简验收:固定 chart/image、审阅相关
release notes、render、确认 `UpgradeSucceeded`/Pod Ready/API 版本,再人工抽查 OIDC
与 Git。长时间观察、临时 branch、重复 CI、逐条 migration 日志和逐 minor 停机备份
不再是默认步骤。
以下任一条件出现时恢复本文的完整流程:数据库或存储变更、rootless/权限模型变化、
PVC identity 或 deployment strategy 变化、重大 chart 结构变化、相关 breaking migration,
以及任何启动失败、CrashLoop 或 migration error。
## 后续但不并入升级
- 将 `gitea-oidc-secret` 等剩余手工 Secret 迁入 OpenBao/ESO;
- Gitea 1.27 稳定后验证最小权限 Actions job token,再实现无集群凭据的 PR render
diff/summary;live `kubectl diff` 仍使用独立、受信任且限权的执行路径;
- 单独评估关闭 chart 生成但无人处理的 Ingress;
- 为 CNPG 和 Gitea `/data` 建立周期性、异机可恢复备份,替代升级前一次性备份。
+34 -20
View File
@@ -1,6 +1,6 @@
# Homelab GitOps and IaC redesign
Status: **design; no infrastructure changes have been applied.**
Status: **implementation in progress; CI、Flux bootstrap、漂移修复与受控 prune 已验证。**
Started 2026-09-09. This is the durable record of the redesign discussion. It
separates observations, decisions and open work so an assumption cannot silently
@@ -35,13 +35,19 @@ or reconcile later CPU, memory, NIC or boot drift.
### Kubernetes and delivery
- Kubernetes is a single-node k3s cluster.
- Helm releases and manifests have historically been applied by hand.
- No Flux or Argo CD installation was found during the initial audit.
- Kubernetes is a single-node k3s `v1.36.4+k3s1` cluster.
- Helm releases and manifests have historically been applied by hand and are now
being adopted by Flux one release at a time.
- Flux `v2.9.5` is live. Its internal Gitea source and root Kustomization are
Ready; `http-echo` proved automatic deployment, drift repair and scoped prune.
- Repository history records External Secrets Operator 2.8.0 as deployed. Five
`ExternalSecret` resources cover Authelia, Gitea, Cloudflared and SeaweedFS.
All five reported `SecretSynced=True` during a live check on 2026-09-09.
- The repository has no configured Git remote yet.
- Gitea Actions now has one instance-scoped runner in namespace `gitea-actions`.
Its Helm release is deployed, its Pod is `2/2 Running`, its identity PVC is
bound, and a sixth `ExternalSecret` delivers the registration token from OpenBao.
- The repository is hosted at `panxiao81/homelab-infra` on the local Gitea
instance. A one-way off-site mirror is still missing.
- Terraform roots remain per-service and must not be merged.
- The k3s node was `Ready` on 2026-09-09. The Snap-packaged `kubectl` could
not start because the user systemd session was degraded; `k3s kubectl` with the
@@ -84,10 +90,12 @@ without becoming another deployment controller.
### CI execution
Most CI uses Gitea Actions for its GitHub Actions compatibility. A persistent
Gitea Runner StatefulSet runs in Kubernetes with rootless Docker-in-Docker and
capacity four; individual job containers are created dynamically. Rootless DinD
still requires a privileged Pod, so the runner is repository-scoped and accepts
trusted workflows only.
Gitea Runner StatefulSet runs in Kubernetes with Docker-in-Docker and capacity
four; individual job containers are created dynamically. The chart requires a
privileged DinD container in both modes, and rootlesskit is blocked by the node's
AppArmor unprivileged-userns policy, so regular DinD is used instead of weakening
that host-wide policy. The runner is instance-scoped and accepts trusted
repositories and workflows only.
Only explicitly labelled jobs needing privilege, nested virtualization,
amd64-only software or isolation from k3s use an ephemeral Proxmox VM. IaC owns
@@ -196,17 +204,23 @@ offline break-glass path. ESO-generated Secrets are projections, not backups.
## Implementation phases
1. Preserve OCI state and sanitized libvirt/Helm evidence.
2. Re-verify the existing OpenBao/ESO delivery and recovery path, inventory
1. **In progress:** sanitized libvirt/Helm evidence is recorded; the independent
encrypted OCI state backup is still missing.
2. **In progress:** live OpenBao/ESO delivery is verified; verify recovery and inventory
remaining manually managed Secrets, then migrate them incrementally.
3. Configure Gitea remote plus a one-way off-site mirror. Confirm whether the
running Gitea supports the 1.27 Terraform State Registry.
4. Manually deploy the reviewed Gitea Runner bootstrap, then bootstrap Flux on
`http-echo` or `marker` without enabling prune until live ownership is audited.
5. Move Tunnel origins to Envoy and consolidate split DNS through Blocky.
6. Deploy Backstage read-only with Catalog, Kubernetes, Flux and TechDocs.
7. Reconstruct the OCI root to a zero-change plan and add libvirt drift reports.
8. Add dynamic PVE VM workers only when a real job requires one, with TTL cleanup
3. **In progress:** the Gitea remote exists; add a one-way off-site mirror and
revisit the Terraform State Registry after upgrading beyond Gitea 1.25.5.
4. **Complete:** the reviewed Gitea Runner and Stage 1 CI are live. Flux deploys
`http-echo`; automatic deployment, replica drift repair and scoped deletion
were verified. Root prune remains disabled for brownfield safety.
5. **In progress:** Gitea Actions and Gitea are managed by Flux. Adopt External
Secrets Operator next with its existing chart `2.8.0` and repository values;
register the suspended release first, then activate it in a separate PR after
proving the fixed render matches Helm's stored manifest.
6. Move Tunnel origins to Envoy and consolidate split DNS through Blocky.
7. Deploy Backstage read-only with Catalog, Kubernetes, Flux and TechDocs.
8. Reconstruct the OCI root to a zero-change plan and add libvirt drift reports.
9. Add dynamic PVE VM workers only when a real job requires one, with TTL cleanup
and hard concurrency/resource limits first.
## Open decisions
@@ -216,7 +230,7 @@ offline break-glass path. ESO-generated Secrets are projections, not backups.
- Whether OCI Object Storage passes the concurrent lockfile test.
- Whether the OCI VM should later move from its current public subnet.
- Schema/generator for the Git-owned service declaration.
- Which live Helm releases are absent from or differ from Git.
- Migration order for Helm releases after the `gitea-actions` adoption.
- Whether each local libvirt VM should autostart.
- Which first job genuinely requires a dynamic Proxmox VM.
+1 -1
View File
@@ -20,7 +20,7 @@ SUB-SKILL", checkbox task lists) that only ever suited one migration. Design
documents now live directly in `docs/` — see `../cicd.md` — and the split that
matters is:
- `CHANGELOG.md` — what changed, for humans
- `CHANGELOG.md` — frozen historical snapshot; Git commits and PRs now record changes
- `CLAUDE.md` — traps and procedures, for agents
- `docs/*.md` — design docs for work not yet built
- `<service>/README.md` — how a service actually works
+29
View File
@@ -0,0 +1,29 @@
# DNS 声明与权威边界
`records.yml` 是 homelab DNS 的唯一声明清单,但不是 DNS 服务本身。不同视图仍由最适合
它们的后端提供:
| 视图 | 权威或递归服务 | 配置方式 |
|---|---|---|
| 公网 `ddupan.top` | Cloudflare | Terraform;尚待完整导入已有记录 |
| AD `ad.ddupan.top` | Samba internal DNS | `samba_dns_record` Ansible module |
| LAN split horizon | Blocky | 尚待从 inventory 渲染或校验 |
| Kubernetes Pod split horizon | CoreDNS | 尚待从 inventory 渲染或校验 |
## 安全边界
- Samba module 只管理 `homelab_dns.samba.records` 明确列出的 RRset,不遍历或清理 zone。
- `_ldap`、`_kerberos`、域控制器 locator 等由 Samba 自动维护的记录不进入 inventory。
- 一个受管 RRset 默认使用 `exact: true`:同名同类型的额外值会被删除,但其他名称和类型
不受影响。
- DHCP 切换不属于本阶段。Blocky 仍未成为 LAN 客户端的正式 resolver。
## 分阶段接管
1. 用 Samba module 接管现有静态 A RRset,首次 check mode 应为零变更。
2. 将 Cloudflare 已有 tunnel DNS 记录导入 Terraform state。
3. 让 Blocky 与 CoreDNS 从 `split_horizon.records` 生成配置或执行 CI 一致性检查。
4. 验证公网、LAN、Pod、AD 四个视图后,再单独修改 DHCP。
当前 inventory 已明确暴露一个既有差异:`obj.ddupan.top` 在 Blocky 中存在,但 CoreDNS
尚无对应覆盖。本阶段不会偷偷修复它;后续在两个 resolver 同时接管时统一修复。
+46
View File
@@ -0,0 +1,46 @@
---
# Homelab DNS desired state. This file is the canonical inventory; individual
# backends consume only the views they own.
homelab_dns:
samba:
# Samba remains authoritative for the AD zone. Only these explicitly listed
# RRsets are reconciled; Samba-generated AD/Kerberos records are untouched.
records:
- { zone: ad.ddupan.top, name: bao, type: A, values: [192.168.10.8] }
- { zone: ad.ddupan.top, name: pve1, type: A, values: [192.168.10.4] }
- { zone: ad.ddupan.top, name: pve2, type: A, values: [192.168.10.7] }
- { zone: ad.ddupan.top, name: pve3, type: A, values: [192.168.10.9] }
- { zone: ad.ddupan.top, name: retrolab, type: A, values: [10.60.0.10] }
- { zone: ad.ddupan.top, name: netbox, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: s3, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-oidc, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: zot, type: A, values: [192.168.10.127] }
split_horizon:
# LAN and pod resolvers should eventually render the same set from here.
# Adoption of Blocky/CoreDNS is deliberately a separate change.
records:
- { name: git.ddupan.top, type: A, values: [192.168.10.127] }
- { name: auth.ddupan.top, type: A, values: [192.168.10.127] }
- { name: obj.ddupan.top, type: A, values: [192.168.10.127] }
public:
# Names expected at Cloudflare. Terraform adoption is a separate change;
# complete RRsets here make the current ownership gap explicit.
records:
- name: auth.ddupan.top
type: CNAME
values: [ff392451-b0b1-45bb-964e-6d9372c3a9e3.cfargotunnel.com]
proxied: true
- name: git.ddupan.top
type: CNAME
values: [ff392451-b0b1-45bb-964e-6d9372c3a9e3.cfargotunnel.com]
proxied: true
- name: obj.ddupan.top
type: CNAME
values: [ff392451-b0b1-45bb-964e-6d9372c3a9e3.cfargotunnel.com]
proxied: true
- name: e5renew.ddupan.top
type: CNAME
values: [ff392451-b0b1-45bb-964e-6d9372c3a9e3.cfargotunnel.com]
proxied: true
+6
View File
@@ -36,6 +36,7 @@ configure any secrets engines/auth methods — that is a separate bootstrap play
```
terraform/ # OpenBao's API-level CONFIGURATION (see below)
mounts.tf pki.tf ssh.tf auth.tf policies.tf
auth-spire.tf # SPIFFE JWT-SVID -> short-lived Bao tokens
imports.tf # adopts the already-running instance into state
policies/*.hcl # policy bodies, kept diffable
```
@@ -205,6 +206,11 @@ after this, use `BAO_ADDR=https://bao.ad.ddupan.top:8200` (no skip-verify), not
## Using it
Kubernetes workload 不接收长期 `BAO_TOKEN`:它通过 SPIRE Workload API 获取
JWT-SVID,再经 `auth/jwt-spire/login` 换取短期、最小权限 token。完整接入流程、
manifest、exchange 脚本、安全要求和排障方法见
[`../../platform/spire/RUNBOOK.md`](../../platform/spire/RUNBOOK.md)。
```bash
# human: log in via Authelia (2FA)
bao login -method=oidc # browser → auth.ddupan.top
@@ -0,0 +1,27 @@
# Workload authentication via SPIFFE JWT-SVIDs. The discovery document and
# JWKS contain public verification material, so this backend carries no secret.
resource "vault_jwt_auth_backend" "spire" {
path = "jwt-spire"
description = "SPIFFE JWT-SVID workload authentication"
oidc_discovery_url = "https://spire-oidc.ad.ddupan.top"
bound_issuer = "https://spire-oidc.ad.ddupan.top"
}
# First end-to-end identity. Keep the subject exact: this role is deliberately
# not a wildcard escape hatch for every workload in the trust domain.
resource "vault_jwt_auth_backend_role" "spire_poc" {
backend = vault_jwt_auth_backend.spire.path
role_name = "spire-poc"
role_type = "jwt"
user_claim = "sub"
bound_audiences = ["openbao"]
bound_claims = {
sub = "spiffe://ddupan.top/ns/spire-poc/sa/spire-jwt-poc"
}
token_policies = [vault_policy.spire_poc.name]
token_no_default_policy = true
token_ttl = 300
token_max_ttl = 900
}
@@ -14,6 +14,11 @@ resource "vault_policy" "ai_agent_ssh" {
policy = file("${path.module}/policies/ai-agent-ssh.hcl")
}
resource "vault_policy" "spire_poc" {
name = "spire-poc"
policy = file("${path.module}/policies/spire-poc.hcl")
}
resource "vault_policy" "snapshot" {
name = "snapshot"
policy = file("${path.module}/policies/snapshot.hcl")
@@ -0,0 +1,10 @@
# Intentionally grants no secret access. This policy proves that an exact
# SPIFFE ID can exchange a JWT-SVID for a bounded OpenBao token and inspect or
# revoke only that token.
path "auth/token/lookup-self" {
capabilities = ["read"]
}
path "auth/token/revoke-self" {
capabilities = ["update"]
}
@@ -56,7 +56,7 @@
# --check-connection makes PVE actually bind before saving, so a wrong DN,
# password, or an untrusted certificate fails HERE instead of silently
# producing a realm nobody can log in to.
ansible.builtin.shell:
ansible.builtin.shell: # noqa command-instead-of-shell
cmd: >-
pveum realm add {{ pve_auth_realm }} --type ad {{ _realm_opts }}
--password '{{ pve_auth_bind_password }}'
@@ -95,7 +95,7 @@
run_once: true
- name: Update the realm
ansible.builtin.shell:
ansible.builtin.shell: # noqa command-instead-of-shell
cmd: >-
pveum realm modify {{ pve_auth_realm }} {{ _realm_opts }}
--password '{{ pve_auth_bind_password }}'
@@ -37,7 +37,7 @@
- name: Refuse to proceed unless that really is a whole disk
# Last line of defence: sgdisk against a partition is destructive, so verify
# the derived device is TYPE=disk and not a partition before touching it.
ansible.builtin.shell:
ansible.builtin.command:
cmd: "lsblk -dno TYPE {{ _ssd_disk.stdout | trim }}"
register: _ssd_type
changed_when: false
@@ -101,7 +101,7 @@
# Same guard as the SSD path. wipefs/vgcreate against a partition by mistake is
# how the pve VG on pve2/pve3 got destroyed on 2026-07-25; assert the device
# type rather than trusting the variable.
ansible.builtin.shell:
ansible.builtin.command:
cmd: "lsblk -dno TYPE {{ pve_linstor_hdd_disk }}"
register: _hdd_type
changed_when: false
+4
View File
@@ -54,6 +54,7 @@ second run is a no-op there and only reconciles config drift.
|---|---|
| `roles/dc_vm/` | creates the DC VM locally via libvirt + cloud-init |
| `roles/samba_ad_dc/` | provisions the DC (§1) — idempotent |
| `ansible/collections/ansible_collections/ddupan/homelab/` | repository-local collection containing the Samba DNS module and its `ansible-test` tests |
| `roles/samba_ad_dc/tasks/verify.yml` | smoke tests (`--tags verify`) |
| `roles/samba_ad_dc/tasks/legacy.yml` | opt-in retro-client protocols (§ Retro clients) |
| `roles/windows_vm/` | unattended-installs the Windows Server 2025 admin box (§2.0) |
@@ -62,6 +63,9 @@ second run is a no-op there and only reconciles config drift.
| `group_vars/all/vault.yml` | admin passwords (gitignored; encrypt with ansible-vault) |
| `inventory/hosts.yml` | DC + Windows hosts |
Static DNS records are declared once in `../../dns/records.yml`. The role manages
only those RRsets and deliberately leaves Samba-generated AD locator records alone.
---
## 0. Decisions — fill these in before touching anything
@@ -1,6 +1,7 @@
[defaults]
inventory = inventory/hosts.yml
roles_path = roles
collections_paths = collections
host_key_checking = False
callback_result_format = yaml
nocows = True
@@ -0,0 +1,4 @@
# ddupan.homelab
Repository-local Ansible collection for homelab infrastructure modules. It is consumed
directly through the `collections_paths` setting and is not published.
@@ -0,0 +1,16 @@
---
namespace: ddupan
name: homelab
version: 0.1.0
readme: README.md
authors:
- panxiao81
description: Ansible plugins used to manage the ddupan.top homelab.
license:
- GPL-3.0-or-later
tags:
- infrastructure
- dns
repository: https://git.ddupan.top/panxiao81/homelab-infra
build_ignore:
- .git
@@ -0,0 +1,159 @@
#!/usr/bin/python
# Copyright: (c) 2026, homelab-infra contributors
# GNU General Public License v3.0+ (see COPYING or https://www.gnu.org/licenses/gpl-3.0.txt)
DOCUMENTATION = r"""
---
module: samba_dns_record
short_description: Reconcile one Samba internal DNS RRset
author:
- Pan Xiao (@panxiao81)
description:
- Reconciles only the named RRset and never deletes an undeclared zone or name.
- Runs C(samba-tool dns) on a Samba AD domain controller using its machine account.
options:
server:
description: DNS server accepted by C(samba-tool dns).
type: str
required: true
zone:
description: Authoritative DNS zone.
type: str
required: true
name:
description: Record owner relative to the zone.
type: str
required: true
type:
description: DNS record type.
type: str
choices: [A, AAAA, CNAME, PTR, TXT]
required: true
values:
description: Complete desired value set for this owner and type.
type: list
elements: str
default: []
state:
description: Whether the desired RRset is present or absent.
type: str
choices: [present, absent]
default: present
exact:
description: Remove live values not present in C(values).
type: bool
default: true
"""
EXAMPLES = r"""
- name: Reconcile an A RRset
samba_dns_record:
server: dc1.ad.ddupan.top
zone: ad.ddupan.top
name: pve1
type: A
values: [192.168.10.4]
exact: true
"""
RETURN = r"""
before:
description: Values observed before reconciliation.
type: list
returned: always
after:
description: Values expected after reconciliation.
type: list
returned: always
"""
import re
from ansible.module_utils.basic import AnsibleModule
ABSENT_ERRORS = (
"WERR_DNS_ERROR_NAME_DOES_NOT_EXIST",
"WERR_DNS_ERROR_RECORD_DOES_NOT_EXIST",
"WERR_DNS_ERROR_NXDOMAIN",
)
def normalize_value(record_type, value):
value = value.strip()
if record_type in ("CNAME", "PTR"):
return value.rstrip(".").lower()
if record_type == "TXT" and len(value) >= 2 and value[0] == value[-1] == '"':
return value[1:-1]
return value.lower() if record_type == "AAAA" else value
def parse_query(stdout, record_type):
values = []
prefix = re.compile(r"^\s*%s:\s*(.*?)\s*(?:\([^)]*\))?\s*$" % re.escape(record_type))
for line in stdout.splitlines():
match = prefix.match(line)
if match:
values.append(normalize_value(record_type, match.group(1)))
return sorted(set(values))
def run(module, args, check_rc=True):
command = [module.get_bin_path("samba-tool", required=True), "dns"] + args + ["-P"]
return module.run_command(command, check_rc=check_rc)
def main():
module = AnsibleModule(
argument_spec=dict(
server=dict(type="str", required=True),
zone=dict(type="str", required=True),
name=dict(type="str", required=True),
type=dict(type="str", required=True, choices=["A", "AAAA", "CNAME", "PTR", "TXT"]),
values=dict(type="list", elements="str", default=[]),
state=dict(type="str", choices=["present", "absent"], default="present"),
exact=dict(type="bool", default=True),
),
supports_check_mode=True,
)
p = module.params
desired = sorted(set(normalize_value(p["type"], value) for value in p["values"]))
if p["state"] == "present" and not desired:
module.fail_json(msg="values must not be empty when state=present")
rc, stdout, stderr = run(
module,
["query", p["server"], p["zone"], p["name"], p["type"]],
check_rc=False,
)
if rc == 0:
current = parse_query(stdout, p["type"])
elif any(marker in stdout + stderr for marker in ABSENT_ERRORS):
current = []
else:
module.fail_json(msg="samba-tool dns query failed", rc=rc, stdout=stdout, stderr=stderr)
target = desired if p["state"] == "present" else []
additions = sorted(set(target) - set(current))
removals = sorted(set(current) - set(target)) if p["exact"] or p["state"] == "absent" else []
changed = bool(additions or removals)
if changed and not module.check_mode:
for value in removals:
run(module, ["delete", p["server"], p["zone"], p["name"], p["type"], value])
for value in additions:
run(module, ["add", p["server"], p["zone"], p["name"], p["type"], value])
after = sorted((set(current) - set(removals)) | set(additions))
module.exit_json(
changed=changed,
before=current,
after=after,
diff={"before": {p["type"]: current}, "after": {p["type"]: after}},
)
if __name__ == "__main__":
main()
@@ -0,0 +1,23 @@
import unittest
from ansible_collections.ddupan.homelab.plugins.modules import samba_dns_record
class ParseQueryTests(unittest.TestCase):
def test_parses_a_records_and_ignores_other_types(self):
output = """ Name=host, Records=3, Children=0
A: 192.168.10.4 (flags=f0, serial=1, ttl=900)
A: 192.168.10.7 (flags=f0, serial=2, ttl=900)
TXT: ignored (flags=f0, serial=3, ttl=900)
"""
self.assertEqual(
samba_dns_record.parse_query(output, "A"),
["192.168.10.4", "192.168.10.7"],
)
def test_normalizes_dns_targets(self):
output = " CNAME: Gateway.AD.DDUPAN.TOP. (flags=f0, serial=1, ttl=900)\n"
self.assertEqual(
samba_dns_record.parse_query(output, "CNAME"),
["gateway.ad.ddupan.top"],
)
@@ -10,29 +10,8 @@ samba_ad_dc_ip: "192.168.10.5"
samba_ad_dns_forwarder: "192.168.10.1"
samba_ad_reverse_zone: "10.168.192.in-addr.arpa" # reverse of 192.168.10.0/24
# Extra A records for non-domain hosts published in the AD DNS zone.
samba_ad_extra_a_records:
- { name: "bao", ip: "192.168.10.8" } # OpenBao (../openbao), not domain-joined
# Proxmox cluster nodes (../proxmox). Not domain-joined; they authenticate
# USERS against this DC rather than being members themselves.
- { name: "pve1", ip: "192.168.10.4" }
- { name: "pve2", ip: "192.168.10.7" }
- { name: "pve3", ip: "192.168.10.9" }
# Lab VMs on the SDN VNets (routed via the VyOS router, see ../../proxmox).
# These are NOT on 192.168.10.0/24, so they have no PTR in the existing
# reverse zone — forward resolution only unless a 0.60.10.in-addr.arpa zone
# is added later.
- { name: "retrolab", ip: "10.60.0.10" }
# k3s services exposed on the LAN through the Envoy gateway (../../../platform/envoy-gateway;
# Contour was retired 2026-07-25). They all point at the k3s node, which is where
# Envoy's LoadBalancer lands; the gateway routes by Host header and serves the
# *.ad.ddupan.top wildcard cert.
# Adding another such service = one more line here + an HTTPRoute, nothing else.
- { name: "netbox", ip: "192.168.10.127" } # NetBox (../../../apps/netbox)
# SeaweedFS S3. Exists so Terraform state does NOT ride the Cloudflare tunnel:
# obj.ddupan.top works, but it hairpins through the WAN, and on 2026-07-28 that
# path was blackholed for hours by a dead VPN tunnel. State must stay on the LAN.
- { name: "s3", ip: "192.168.10.127" } # SeaweedFS S3 (../../../apps/seaweedfs)
# Static records now live in ../../dns/records.yml and are reconciled as complete
# RRsets by the local samba_dns_record module.
# Support legacy clients (Win9x/NT4/2000/XP)? INSECURE — see README "Retro clients".
samba_ad_legacy_clients: false
@@ -6,6 +6,8 @@
hosts: samba_dc
become: true
gather_facts: true
vars_files:
- ../../dns/records.yml
roles:
- role: samba_ad_dc
post_tasks:
@@ -0,0 +1,14 @@
---
- name: Reconcile explicitly managed AD DNS RRsets
ddupan.homelab.samba_dns_record:
server: "{{ samba_ad_dc_ip }}"
zone: "{{ item['zone'] }}"
name: "{{ item['name'] }}"
type: "{{ item['type'] }}"
values: "{{ item['values'] }}"
state: present
exact: true
loop: "{{ homelab_dns.samba.records }}"
loop_control:
label: "{{ item['name'] }}.{{ item['zone'] }} {{ item['type'] }}"
tags: [dns]
@@ -6,7 +6,7 @@
- name: Inject legacy protocol settings into smb.conf [global]
ansible.builtin.blockinfile:
path: /etc/samba/smb.conf
marker: "\t# {mark} ANSIBLE MANAGED — legacy clients (INSECURE)"
marker: "\t# {mark} ANSIBLE MANAGED — legacy clients (INSECURE)" # noqa no-tabs
insertafter: '^\[global\]'
block: |2
server min protocol = NT1
@@ -198,20 +198,11 @@
- "'already exists' not in (kms_srv.stderr | default('')) + (kms_srv.stdout | default(''))"
no_log: true
# --- Extra A records for non-domain hosts (e.g. OpenBao) -----------------------
# --- Directory objects and explicitly managed DNS records ----------------------
- name: Service accounts and RBAC groups
ansible.builtin.import_tasks: directory_objects.yml
tags: [directory, accounts]
- name: Register extra A records in the AD DNS zone
ansible.builtin.command:
cmd: >-
samba-tool dns add {{ samba_ad_dc_ip }} {{ samba_ad_realm | lower }}
{{ item.name }} A {{ item.ip }} -P
loop: "{{ samba_ad_extra_a_records }}"
register: extra_a
changed_when: "'Record added successfully' in (extra_a.stdout | default(''))"
failed_when:
- extra_a.rc != 0
- "'already exists' not in (extra_a.stderr | default('')) + (extra_a.stdout | default(''))"
- name: Reconcile static AD DNS records
ansible.builtin.import_tasks: dns_records.yml
tags: [dns]
+12
View File
@@ -51,6 +51,18 @@ the only challenge that can issue a wildcard.
## Deploy
### Flux 接管状态
现有 release 为 chart/app `v1.21.0`,Git 中固定同一版本并已由 Flux HelmRelease
完成接管。第一阶段暂停登记后确认 chart artifact Ready、Helm revision 保持为 1,且
三个 workload Pod 均未被替换;随后通过独立 PR 解除暂停。
Issuer 与 Certificate 清单随本目录的 Kustomization 由 Flux 管理;现有 Cloudflare
token Secret 只被引用,本次接管不改变其所有权。删除保护期间保持 `prune: false`,且
CRD 同时启用 chart 的 `crds.keep` 与 Flux Helm action 的 `CreateReplace`。
以下命令保留为 break-glass 手工恢复流程;正常变更应提交 Git:
```bash
helm repo add jetstack https://charts.jetstack.io && helm repo update jetstack
helm upgrade --install cert-manager jetstack/cert-manager --version v1.21.0 \
@@ -32,12 +32,12 @@ spec:
# http-01, not dns01: bao resolves ad.ddupan.top and can reach LAN hosts
# directly (noted as verified in openbao/terraform/pki.tf), so it can fetch
# the challenge over the LAN with no public exposure. cert-manager creates a
# temporary HTTPRoute on the shared Contour gateway to answer it.
# temporary HTTPRoute on the shared Envoy Gateway to answer it.
- http01:
gatewayHTTPRoute:
parentRefs:
- name: contour-gateway
namespace: projectcontour
- name: eg
namespace: envoy-gateway-system
kind: Gateway
group: gateway.networking.k8s.io
sectionName: http # the plaintext :80 listener
+33
View File
@@ -0,0 +1,33 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: cert-manager
namespace: cert-manager
spec:
chart:
spec:
chart: cert-manager
interval: 1h
sourceRef:
kind: HelmRepository
name: jetstack
version: v1.21.0
driftDetection:
mode: enabled
install:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: cert-manager
targetNamespace: cert-manager
timeout: 10m
upgrade:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: cert-manager-values
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: jetstack
namespace: cert-manager
spec:
interval: 1h
url: https://charts.jetstack.io
+19
View File
@@ -0,0 +1,19 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: cert-manager-values
namespace: cert-manager
files:
- values.yaml=values.yaml
resources:
- helmrepository.yaml
- helmrelease.yaml
- clusterissuer-letsencrypt.yaml
- clusterissuer-bao-acme.yaml
- certificate-wildcard-ad.yaml
- certificate-auth-ddupan.yaml
- certificate-git-ddupan.yaml
+15
View File
@@ -79,6 +79,21 @@ unauthenticated requests to an app whose auth model is "trust the header".
## Install
### Flux 接管状态
现有 release 为 OCI chart/app `v1.5.6`,且没有 user-supplied values。Git 中固定
同一版本并已完成分阶段 Flux HelmRelease 接管。第一阶段确认 Kustomization Ready、
Helm revision 保持为 1,controller/data-plane Pod UID 与入口响应均未变化后,再通过
独立 PR 解除暂停。OCI 类型 HelmRepository 是按需 chart generator,本身不产生
artifact 或 Ready condition;解除暂停后由 HelmRelease 请求并验证 chart。
`GatewayClass eg` 与 `Gateway eg` 随本目录 Kustomization 由 Flux 管理。接管期间保持
`prune: false`;Gateway API 与 Envoy Gateway CRD 使用 `CreateReplace`,延续现有 Helm
所有权且绝不通过删除 CRD 迁移。Gateway 的三个 `certificateRefs` 最初属于
`kubectl-client-side-apply`;首次 reconcile 已由 Flux 默认 SSA `Override` 接管这些在
Git 中声明且值相同的字段,没有改变 listener spec。不得用 `force: true`,它用于
不可变字段失败时删除重建资源。以下命令保留为 break-glass 手工恢复流程。
```bash
# The Gateway API CRDs may already be owned by another tool's field manager (Contour's
# quickstart used client-side apply), which makes Helm fail with
+30
View File
@@ -0,0 +1,30 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: envoy-gateway
namespace: envoy-gateway-system
spec:
chart:
spec:
chart: gateway-helm
interval: 1h
sourceRef:
kind: HelmRepository
name: envoy-gateway
version: v1.5.6
driftDetection:
mode: enabled
install:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: envoy-gateway
targetNamespace: envoy-gateway-system
timeout: 10m
upgrade:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
@@ -0,0 +1,9 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: envoy-gateway
namespace: envoy-gateway-system
spec:
interval: 1h
type: oci
url: oci://docker.io/envoyproxy
@@ -0,0 +1,6 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- helmrepository.yaml
- helmrelease.yaml
- gateway.yaml
+20
View File
@@ -0,0 +1,20 @@
# External Secrets Operator
External Secrets Operator(ESO)把 OpenBao `kv/k8s/*` 下的值投影为 Kubernetes
Secret。`ClusterSecretStore/openbao` 使用 `external-secrets` ServiceAccount 的短期
JWT 登录 OpenBao,不在 Git 中保存长期凭据。
## Flux 接管
现有 release 是 2026-07-28 手工安装的 chart `external-secrets` `2.8.0`,Helm
revision 1。接管前审计确认:本目录 `values.yaml` 与 Helm stored user values 一致;
用固定 chart 生成的 33,528 行 manifest 与 stored manifest 只有末尾空行差异。
接管分两阶段。第一阶段以 `suspend: true` 登记 HelmRepository、values ConfigMap 与
HelmRelease;合并后已确认 source Ready、release 仍为 revision 1,三个 Pod UID 保持
不变且零重启。第二阶段只移除 `suspend`,允许 Flux 修正 Helm stored ownership 并
启用 drift detection;chart、values 和子 Kustomization 的 `prune: false` 均不变。
`clustersecretstore.yaml` 和 `externalsecrets.yaml` 是现有 secret delivery intent,
本阶段故意不把它们加入该 Kustomization,避免在 Helm release 接管时同时扩大 Flux
ownership。Helm 接管稳定后再单独审计、纳管这些对象。
+35 -2
View File
@@ -101,6 +101,7 @@ spec:
# The chart normally GENERATES seaweedfs-s3-secret from s3.credentials. We point
# filer.s3.existingConfigSecret at this one instead, so the chart stops rendering
# credentials from values entirely.
# zot 的 AK/SK 只保存在 k8s/zot-s3;在此组装服务端配置,不在基础配置中维护副本。
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
@@ -114,6 +115,38 @@ spec:
target:
name: seaweedfs-s3-config
creationPolicy: Owner
dataFrom:
- extract:
template:
engineVersion: v2
mergePolicy: Replace
data:
seaweedfs_s3_config: |-
{{- $config := mustFromJson .baseConfig -}}
{{- if not (kindIs "slice" $config.identities) -}}
{{- fail "base S3 configuration must contain an identities array" -}}
{{- end -}}
{{- if or (eq .zotAccessKey "") (eq .zotSecretKey "") -}}
{{- fail "zot S3 credentials must not be empty" -}}
{{- end -}}
{{- $identities := list -}}
{{- range $config.identities -}}
{{- if ne .name "zot" -}}
{{- $identities = append $identities . -}}
{{- end -}}
{{- end -}}
{{- $credential := dict "accessKey" .zotAccessKey "secretKey" .zotSecretKey -}}
{{- $zot := dict "name" "zot" "credentials" (list $credential) "actions" (list "Read:zot" "Write:zot" "List:zot" "Tagging:zot") -}}
{{- $_ := set $config "identities" (append $identities $zot) -}}
{{- mustToJson $config -}}
data:
- secretKey: baseConfig
remoteRef:
key: k8s/seaweedfs-s3
property: seaweedfs_s3_config
- secretKey: zotAccessKey
remoteRef:
key: k8s/zot-s3
property: access_key
- secretKey: zotSecretKey
remoteRef:
key: k8s/zot-s3
property: secret_key
@@ -0,0 +1,31 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: external-secrets
namespace: external-secrets
spec:
chart:
spec:
chart: external-secrets
interval: 1h
sourceRef:
kind: HelmRepository
name: external-secrets
version: 2.8.0
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: external-secrets
targetNamespace: external-secrets
timeout: 10m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: external-secrets-values
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: external-secrets
namespace: external-secrets
spec:
interval: 1h
url: https://charts.external-secrets.io
@@ -0,0 +1,14 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: external-secrets-values
namespace: external-secrets
files:
- values.yaml=values.yaml
resources:
- helmrepository.yaml
- helmrelease.yaml
+38 -6
View File
@@ -2,21 +2,53 @@
This is the bootstrap runner for Gitea Actions. One persistent runner Pod accepts
up to four jobs; each job runs in a dynamically created container inside a
rootless Docker-in-Docker daemon. Rootless DinD still requires a privileged Pod
to create its user namespace, so this runner is restricted to this repository
and trusted workflows.
Docker-in-Docker daemon. The official chart runs DinD privileged. Rootless DinD
would still be privileged and is blocked by the node's AppArmor user-namespace
policy, so this deployment uses regular DinD instead of weakening that host-wide
policy. Only trusted workflows may target this runner.
DinD 同时使用 `--mtu=1450` 和
`--default-network-opt=bridge=com.docker.network.driver.mtu=1450`,与 k3s Pod 的
`eth0` 一致。前者只覆盖 Docker 默认 bridge;act 为每个 job 创建 user-defined
bridge,必须由后者设置默认 MTU。不要在未验证节点 Pod MTU 的情况下删除或修改这
两个参数:MTU 1500 的 job 容器虽然能够解析 GitHub、甚至建立 TCP 连接,但较大的
TLS 数据包会在嵌套网络路径中丢失,表现为 `github.com` / `api.github.com` 超时或
`setup-go` 每次请求卡满 6 分钟后重试。Pod 网络和默认 Docker bridge 正常不代表
Actions job bridge 正常。
The runner is registered at instance scope so it is available to every repository
on this Gitea instance. Repository permissions and protected-branch review are
therefore the security boundary; do not enable Actions for untrusted repositories.
The runner registration token is authoritative in OpenBao at
`kv/k8s/gitea-runner`. External Secrets Operator projects its `token` property to
the `gitea-runner-token` Secret. Never put the token in this directory or a Helm
command line.
## Review-first bootstrap
## Flux 接管状态
This is a one-time manual deployment because Flux is not installed yet:
该 release 最初通过下述 review-first 流程手动 bootstrap。下一个 GitOps 阶段将
使用 Flux `HelmRelease` 接管它,并首先固定现有 chart `0.1.1`,不在接管 PR 中升级。
迁移前审计发现:Helm 保存的 user-supplied values 和 release manifest 仍描述失败的
rootless DinD 尝试,但 live StatefulSet 与本目录 `values.yaml` 都已经使用 regular
DinD。首次 reconcile 的验收条件是修正 Helm 存储状态,同时 live Pod spec、PVC
identity、runner capacity 和在线状态保持不变。接管稳定后再用独立 PR 升级 chart。
接管分两阶段:第一阶段提交 `suspend: true` 的 HelmRelease、HelmRepository 和由
`values.yaml` 生成的 ConfigMap。Flux 只登记这些对象,不执行 Helm action。合并后
检查 HelmRepository Ready,并用固定 chart 重复比较期望清单与 live StatefulSet;
第二阶段解除 suspend。第一阶段已经确认 source Ready、完整 chart render 与 live
资源零差异,且登记过程中现有 runner 没有 rollout。失败重试使用
`RetryOnFailure`,不会用 stored rootless release 做 rollback。
## 历史 review-first bootstrap
这是 Flux 安装前执行过的一次性手动部署流程,保留用于恢复和审计:
1. Merge the reviewed PR.
2. Create a repository-scoped runner registration token in Gitea.
2. As a Gitea site administrator, create an instance-scoped runner registration
token under **Site Administration → Actions → Runners**.
3. Store it as the `token` property at `kv/k8s/gitea-runner` without exposing it
in shell history:
+31
View File
@@ -0,0 +1,31 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: gitea-actions
namespace: gitea-actions
spec:
chart:
spec:
chart: actions
interval: 1h
sourceRef:
kind: HelmRepository
name: gitea-charts
version: 0.1.1
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: gitea-actions
targetNamespace: gitea-actions
timeout: 10m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: gitea-actions-values
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: gitea-charts
namespace: gitea-actions
spec:
interval: 1h
url: https://dl.gitea.com/charts/
+16
View File
@@ -0,0 +1,16 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: gitea-actions-values
namespace: gitea-actions
files:
- values.yaml=values.yaml
resources:
- namespace.yaml
- external-secret.yaml
- helmrepository.yaml
- helmrelease.yaml
+12 -8
View File
@@ -44,14 +44,18 @@ statefulset:
docker_timeout: 300s
dind:
rootless: true
uid: 1000
# The node enforces AppArmor's unprivileged-userns restriction, which blocks
# rootlesskit even though this chart must run DinD privileged either way.
rootless: false
registry: docker.io
repository: docker
tag: 29.7.1-dind-rootless
tag: 29.7.1-dind
pullPolicy: IfNotPresent
extraEnvs:
- name: DOCKERD_ROOTLESS_ROOTLESSKIT_NET
value: slirp4netns
- name: DOCKERD_ROOTLESS_ROOTLESSKIT_MTU
value: "65520"
# k3s uses a 1450-byte pod MTU. Without matching it here, nested Actions
# networks advertise 1500 and GitHub TLS packets disappear on the outer
# overlay path while direct pod traffic remains healthy.
extraArgs:
- --mtu=1450
# --mtu only changes Docker's default bridge. act creates a user-defined
# bridge per job, so give every new bridge the same explicit default.
- --default-network-opt=bridge=com.docker.network.driver.mtu=1450
+28
View File
@@ -13,6 +13,34 @@ Jaeger API), Grafana (3 datasources healthy, VM dashboard loaded, Tailscale ingr
`localpv-zfs-ceph`. Note the vmagent hot-reload RBAC fix (`metrics/reload-rbac.yaml`,
see `metrics/README-reload.md`) and that VL/VT CRDs are served under `.../v1`.
### Flux 接管状态
可观测性栈使用一个 `observability` Flux Kustomization,按依赖顺序逐步加入 Operator、
Metrics、Logs、Traces 和 Grafana,避免多个 reconciler 争夺共享 namespace、Helm source
或 VictoriaMetrics CRD。brownfield 迁移期间保持 `prune: false`。
VictoriaMetrics Operator 已固定 chart `0.66.2` / app `v0.73.1` 完成分阶段 Flux
HelmRelease 接管,保存 values 与 `operator/values.yaml` 一致。阶段一确认 Helm revision
保持为 2、Operator Pod UID 未变、24 个 CRD 及 6 个核心 VM/VL/VT CR 未变化;随后
通过独立 PR 解除暂停。
其余已部署资源也已统一完成接管:Metrics CR、rules、scrapes、node-exporter,Logs 的
VLSingle 与固定 `0.3.6` 的 Collector,以及 Traces 的 VTSingle、OTel Collector,最后是
固定 `10.5.15` 的 Grafana 与 dashboard ConfigMap。阶段一确认 11 个 Pod UID、4 个 PVC、
6 个核心 CR、17 个 rule/scrape 对象及 dashboard 哈希均未变化,Collector 和 Grafana
Helm revision 均保持为 1;随后在同一独立 PR 中解除两个 HelmRelease 的暂停。
首次 Grafana Helm reconcile 暴露了本地 ZFS RWO 卷限制:chart 默认 `RollingUpdate`
会先创建新 Pod,但 CSI 拒绝在旧 Pod 仍挂载 PVC 时再次 mount,使 upgrade 卡在
`pending-upgrade`。`grafana/values.yaml` 因此显式使用 `deploymentStrategy.type: Recreate`。
Grafana 升级会有一次短暂停机,但旧 Pod 会先退出,新 Pod 才挂载同一 PVC;不要改回
RollingUpdate,除非存储改为真正支持并发挂载的 RWX。
以下内容明确不属于本批接管:尚未部署的 kube-state-metrics、`vlogs-ingress.yaml`,以及
仅作 standalone chart 参考的 `traces-values.yaml`。`grafana/oidc-secret.yaml` 是不含真实
值的占位模板;live `grafana-oidc` Secret 继续只被 HelmRelease 引用,尚未由 OpenBao/ESO
纳管。
Still host-side (not yet done): deploy `docker-hosts/compose.yaml` on the Docker
hosts and `logs/vlogs-ingress.yaml` to push their logs.
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: grafana
namespace: monitoring
spec:
interval: 1h
url: https://grafana.github.io/helm-charts
@@ -0,0 +1,31 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: grafana
namespace: monitoring
spec:
chart:
spec:
chart: grafana
interval: 1h
sourceRef:
kind: HelmRepository
name: grafana
version: 10.5.15
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: grafana
targetNamespace: monitoring
timeout: 10m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: grafana-values
@@ -0,0 +1,21 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
configMapGenerator:
- name: grafana-values
namespace: monitoring
files:
- values.yaml=values.yaml
options:
labels:
reconcile.fluxcd.io/watch: Enabled
- name: grafana-dashboard-victoriametrics
namespace: monitoring
files:
- victoriametrics.json=dashboards/victoriametrics.json
options:
labels:
grafana_dashboard: "1"
resources:
- helmrelease.yaml
@@ -48,6 +48,12 @@ persistence:
storageClassName: localpv-zfs-ceph
size: 5Gi
# localpv-zfs-ceph cannot mount the same RWO volume into the old and new Grafana
# Pods concurrently. RollingUpdate leaves the old Pod serving while the new Pod
# blocks forever in verifyMount, so upgrades must stop the old Pod first.
deploymentStrategy:
type: Recreate
# --- Private exposure via the Tailscale ingress (like seaweedfs-admin) ---
# The tailscale operator provisions grafana.<tailnet>.ts.net and a TLS cert.
ingress:
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: vm
namespace: monitoring
spec:
interval: 1h
url: https://victoriametrics.github.io/helm-charts/
+11
View File
@@ -0,0 +1,11 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- helmrepository.yaml
- grafana-helmrepository.yaml
- operator
- metrics
- logs
- traces
- grafana
@@ -0,0 +1,31 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: victoria-logs-collector
namespace: monitoring
spec:
chart:
spec:
chart: victoria-logs-collector
interval: 1h
sourceRef:
kind: HelmRepository
name: vm
version: 0.3.6
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: victoria-logs-collector
targetNamespace: monitoring
timeout: 10m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: victoria-logs-collector-values
@@ -0,0 +1,14 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: victoria-logs-collector-values
namespace: monitoring
files:
- values.yaml=collector-values.yaml
resources:
- vlsingle.yaml
- helmrelease.yaml
@@ -0,0 +1,16 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- vmsingle.yaml
- vmagent.yaml
- vmalert.yaml
- vmalertmanager.yaml
- reload-rbac.yaml
- rules/vm-health.yaml
- rules/vmagent.yaml
- rules/vmalert.yaml
- rules/vmsingle.yaml
- scrapes/blocky.yaml
- scrapes/docker-hosts.yaml
- scrapes/kubelet.yaml
- exporters/node-exporter.yaml
@@ -0,0 +1,33 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: vm-operator
namespace: monitoring
spec:
chart:
spec:
chart: victoria-metrics-operator
interval: 1h
sourceRef:
kind: HelmRepository
name: vm
version: 0.66.2
driftDetection:
mode: enabled
install:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: vm-operator
targetNamespace: monitoring
timeout: 10m
upgrade:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: vm-operator-values
@@ -0,0 +1,13 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: vm-operator-values
namespace: monitoring
files:
- values.yaml=values.yaml
resources:
- helmrelease.yaml
@@ -0,0 +1,5 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- vtsingle.yaml
- otel-collector.yaml
+53
View File
@@ -0,0 +1,53 @@
# OpenEBS — 本地 ZFS 存储
OpenEBS umbrella chart 为单节点 k3s 提供 `zfs-localpv` CSI provisioner,并保留
chart 自带的 local PV、Loki、Alloy 与 MinIO 组件。业务 PVC 使用独立声明的
`localpv-zfs-ceph` StorageClass,在宿主机 ZFS pool `data/ceph` 上创建 dataset/zvol。
## 当前版本与引擎
- release:`openebs`,namespace:`openebs`
- chart/app:`4.4.0`
- ZFS LocalPV:启用
- LVM LocalPV、RawFile LocalPV、Mayastor:显式禁用
- bundled Loki SingleBinary:单副本
不要在 chart 升级时顺带启用其他存储引擎。`localpv-zfs-ceph` 当前承载 SeaweedFS、
shared PostgreSQL、NetBox、VictoriaMetrics、VictoriaLogs、VictoriaTraces、Grafana 与
RustFS 的持久卷;删除 StorageClass 不会删除已有 PV,但误删/重建 chart 资源可能中断
provisioner,因此 brownfield 接管期间保持 `prune: false`。
## Flux 接管状态
现有 Helm release 保存的 user-supplied values 与本目录 `values.yaml` 一致,并已固定
`4.4.0` 完成分阶段 Flux HelmRelease 接管。阶段一确认 source Ready、Helm revision
保持为 3、8 个 OpenEBS Pod UID 未变、10/10 ZFS PV 和集群 20/20 PVC Bound,且
`data/ceph` 下 10 个 volume 未变;随后通过独立 PR 解除暂停。
CRD 在 install/upgrade 时使用 `CreateReplace`,但绝不通过删除 CRD 迁移。接管和升级
是两个独立动作;首次激活不得改变 chart、values、StorageClass 或存储引擎。
## Break-glass 手工恢复
正常变更应提交 Git。Flux 不可用时可暂时执行:
```bash
helm upgrade --install openebs openebs/openebs --version 4.4.0 \
-n openebs --create-namespace -f platform/openebs/values.yaml
kubectl apply -f platform/openebs/storageclasses.yaml
```
恢复 Flux 后,应以 Git 为准并确认 HelmRelease 没有 drift。
## 验证
```bash
kubectl -n openebs get pods
kubectl get storageclass localpv-zfs-ceph
kubectl get pv | grep localpv-zfs-ceph
kubectl get pvc -A
zfs list -r data/ceph
```
所有既有 PV 必须保持 `Bound`。验证新 provisioner 时只创建单独的临时 PVC,不要修改
或删除业务 PVC,也不要用 `helm uninstall` 验证接管。
+33
View File
@@ -0,0 +1,33 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: openebs
namespace: openebs
spec:
chart:
spec:
chart: openebs
interval: 1h
sourceRef:
kind: HelmRepository
name: openebs
version: 4.4.0
driftDetection:
mode: enabled
install:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: openebs
targetNamespace: openebs
timeout: 15m
upgrade:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: openebs-values
+8
View File
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: openebs
namespace: openebs
spec:
interval: 1h
url: https://openebs.github.io/openebs

Some files were not shown because too many files have changed in this diff Show More