Author SHA1 Message Date
panxiao81 57b13387cb ci: 输出 kind 节点准备调试日志
kind-on-kata-smoke / smoke (push) Failing after 23m50s
2026-09-14 16:16:36 +00:00
panxiao81 1deb024621 ci: 移除错误的容器内磁盘检查
kind-on-kata-smoke / smoke (push) Failing after 13m13s
2026-09-14 16:02:13 +00:00
panxiao81 311a506982 ci: 记录 kind guest 临时盘容量
kind-on-kata-smoke / smoke (push) Failing after 8s
2026-09-14 16:00:51 +00:00
panxiao81 a0d38fc9ab ci: 验证 kind 使用 guest-local overlay2
kind-on-kata-smoke / smoke (push) Failing after 10m59s
2026-09-14 15:54:28 +00:00
panxiao81 fa293ffc80 ci: 重试 kind 工具下载
kind-on-kata-smoke / smoke (push) Failing after 16m43s
2026-09-14 14:53:46 +00:00
panxiao81 77df686981 ci: 向 kind node 提供 guest kmsg
kind-on-kata-smoke / smoke (push) Failing after 28s
2026-09-14 14:51:37 +00:00
panxiao81 9e6515406f ci: 将 guest Docker 暴露给 kind job
kind-on-kata-smoke / smoke (push) Failing after 5m45s
2026-09-14 14:42:10 +00:00
panxiao81 f6c216ce93 ci: 验证 Kata 内运行 kind 集群
kind-on-kata-smoke / smoke (push) Failing after 15s
2026-09-14 14:38:14 +00:00
panxiao81 e0e8213b8f Merge pull request '部署 zot:接入 SeaweedFS、SPIRE 并统一 S3 凭据来源' (#55) from feat/zot-spire-s3 into main
yaml / yaml (push) Successful in 11s
ansible / collection-test (push) Successful in 50s
ansible / lint (push) Successful in 11m30s
Reviewed-on: #55
2026-09-14 13:38:17 +00:00
panxiao81 5c2b575a4f feat(zot): 接入 SeaweedFS 与 SPIRE 并统一 S3 凭据来源
yaml / yaml (pull_request) Successful in 20s
ansible / collection-test (pull_request) Successful in 59s
ansible / lint (pull_request) Successful in 10m32s
2026-09-14 13:26:54 +00:00
panxiao81 67dfe1ddcb Merge pull request 54: 记录 SPIRE 与 OpenBao workload identity 用法 2026-09-14 10:33:14 +00:00
panxiao81 ed86738c4b 文档:记录 SPIRE 与 OpenBao workload identity 用法 2026-09-14 10:31:36 +00:00
panxiao81 8c27a7286b Merge pull request 53: 配置 SPIRE JWT-SVID 登录 OpenBao
terraform / validate (push) Successful in 1m5s
2026-09-13 16:01:52 +00:00
panxiao81 c9987122e6 配置 SPIRE JWT-SVID 登录 OpenBao
terraform / validate (pull_request) Successful in 1m12s
2026-09-13 15:56:08 +00:00
22 changed files with 1023 additions and 6 deletions
+63
View File
@@ -0,0 +1,63 @@
---
name: kind-on-kata-smoke
on:
push:
branches:
- poc/kind-on-kata
paths:
- .gitea/workflows/kind-on-kata-smoke.yml
workflow_dispatch:
jobs:
smoke:
runs-on: kata-poc
steps:
- name: Create nested kind cluster
shell: sh
env:
KIND_VERSION: v0.27.0
KIND_NODE_IMAGE: kindest/node:v1.32.2@sha256:f226345927d7e348497136874b6d207e0b32cc52154ad8323129352923a3142f
run: |
set -eu
apk add --no-cache ca-certificates curl docker-cli
curl --retry 5 --retry-all-errors --connect-timeout 15 -fsSLo /tmp/kind \
"https://kind.sigs.k8s.io/dl/${KIND_VERSION}/kind-linux-amd64"
curl --retry 5 --retry-all-errors --connect-timeout 15 -fsSLo /tmp/kind.sha256sum \
"https://kind.sigs.k8s.io/dl/${KIND_VERSION}/kind-linux-amd64.sha256sum"
expected="$(awk '{print $1}' /tmp/kind.sha256sum)"
printf '%s %s\n' "$expected" /tmp/kind | sha256sum -c -
install -m 0755 /tmp/kind /usr/local/bin/kind
docker info --format 'kernel={{.KernelVersion}} driver={{.Driver}}'
test "$(docker info --format '{{.Driver}}')" = overlay2
cleanup() {
kind delete cluster --name nested >/dev/null 2>&1 || true
}
trap cleanup EXIT
cat >/tmp/kind-config.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
name: nested
nodes:
- role: control-plane
extraMounts:
- hostPath: /dev/kmsg
containerPath: /dev/kmsg
EOF
kind create cluster -v 9 --retain --config /tmp/kind-config.yaml --image "$KIND_NODE_IMAGE" --wait 5m
docker exec nested-control-plane kubectl \
--kubeconfig=/etc/kubernetes/admin.conf wait \
--for=condition=Ready node/nested-control-plane --timeout=2m
docker exec nested-control-plane kubectl \
--kubeconfig=/etc/kubernetes/admin.conf run smoke \
--image=docker.io/library/busybox:1.37 --restart=Never \
--command -- sh -c 'echo kind-on-kata-ok'
docker exec nested-control-plane kubectl \
--kubeconfig=/etc/kubernetes/admin.conf wait \
--for=jsonpath='{.status.phase}'=Succeeded pod/smoke --timeout=2m
test "$(docker exec nested-control-plane kubectl \
--kubeconfig=/etc/kubernetes/admin.conf logs smoke)" = kind-on-kata-ok
kind delete cluster --name nested
trap - EXIT
+5
View File
@@ -150,6 +150,11 @@ recovered, so `.vault_pass.gpg` is the authoritative recovery path.
all-clear. Use `git check-ignore --no-index` and `git rm --cached` to actually remove it.
- **A `.tfplan` is a zip containing a full `tfstate`.** It walks straight past `*.tfstate`
ignore rules. Ignore `*.tfplan` everywhere.
- **SPIRE CLI JSON can be an array of response blocks.** `spire-agent api fetch jwt
-output json` in 1.15.3 returns blocks containing `svids` and `bundles`. Capture stdout
privately and type-check before extracting fields; `list(response)` prints full tokens
when the response is already an array. Never inspect credential payloads by printing
their containers, and never put fetched JWTs in command arguments or Pod logs.
- **Quoting does not survive two ssh hops.** `ssh pve1 "ssh pve3 'cmd | qm monitor 103'"`
loses the inner quotes — ssh re-joins argv with spaces, so the pipeline splits and the
tail runs on the **jump host**. It fails silently if you discard stderr: a `screendump`
+22 -2
View File
@@ -11,7 +11,9 @@
| `helm.sh` | Installs or upgrades the SeaweedFS release. |
**Install**
1. Set real S3 access and secret keys in `values.yaml`.
1. 在 OpenBao `kv/k8s/seaweedfs-s3` 维护基础 S3 配置;zot 凭据单独以
`kv/k8s/zot-s3` 为唯一来源。ESO 合成为 `seaweedfs-s3-config`,详见下文。
不要把真实 AK/SK 放进 `values.yaml`。
2. Apply the manifests:
```bash
bash ~/services/apps/seaweedfs/helm.sh
@@ -31,4 +33,22 @@
**Notes**
- The chart manages master, volume, filer, S3, and admin components.
- The chart-managed S3 secret uses the current AK/SK pair for the admin user.
- The filer uses the ESO-managed `seaweedfs-s3-config` Secret for static S3 identities.
## zot 制品存储
`zot` bucket 专用于 [zot Registry](../zot/README.md),OCI 数据位于 `registry/`
前缀。静态身份 `zot` 只有该 bucket 的 Read/Write/List/Tagging 权限,凭据唯一来源为
Bao `kv/k8s/zot-s3` 的 `access_key` / `secret_key`,同时供 zot consumer 和
SeaweedFS 服务端使用。
[ExternalSecret 模板](../../platform/external-secrets/externalsecrets.yaml) 保留
`kv/k8s/seaweedfs-s3` 的原有身份及其他配置,再追加 zot 身份与限定 bucket 的权限。
基础配置当前版本不保存 zot AK/SK;旧 KV 版本历史仍保留。新增其他身份时使用
KV compare-and-set 保留已有内容,不覆盖 Terraform 或其他应用的 AK/SK。
不要直接编辑生成的 Kubernetes Secret。该 ExternalSecret 已单独应用到集群,
目前仍未加入 ESO 的 Flux Kustomization,遵循该组件现有 ownership 边界。
运行版本 `4.22` 可在 Secret volume 更新后向 filer/内嵌 S3 的 `weed` 进程发送
SIGHUP,重新加载静态配置,无需重启共享 S3 服务。本次接入没有启用 SeaweedFS
OIDC/STS;SPIRE 认证发生在 zot 的客户端入口。
+123
View File
@@ -0,0 +1,123 @@
# zot OCI Registry
内网入口为 `https://zot.ad.ddupan.top`。使用官方 Helm chart `0.1.124`,运行
zot `v2.1.21`,镜像固定到官方 linux/amd64 digest。
## 存储与凭据
制品、manifest 和 OCI layout 保存在现有 SeaweedFS 的 `zot` bucket,前缀为
`registry/`,S3 endpoint 为 `https://s3.ad.ddupan.top`。**不创建 PVC**;chart 的
`/var/lib/registry` 是 `emptyDir`,仅用于运行时本地工作数据。
首期单副本,关闭跨仓库 dedupe,不额外部署 Redis/DynamoDB 缓存。保留 zot GC,
暂不配置自动删除已发布版本的 retention policy。增加副本、启用 dedupe 或搜索等
扩展前,需要重新检查共享元数据与缓存的持久化要求。
凭据链路:
```text
OpenBao kv/k8s/seaweedfs-s3
→ 原有 S3 身份及基础配置 ─┐
├→ ESO 模板 → seaweedfs/seaweedfs-s3-config
OpenBao kv/k8s/zot-s3 ────┘ → 完整 s3.config
└→ ESO → zot/zot-s3 → zot 的 AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY
```
`kv/k8s/zot-s3` 是 zot AK/SK 的唯一维护来源。基础配置保留原有身份及其他字段,
不再保存 zot 凭据副本;[SeaweedFS ExternalSecret](../../platform/external-secrets/externalsecrets.yaml)
使用 ESO v2 模板追加 zot 身份。两个 Kubernetes Secret 都是自动生成的消费副本,
不手工编辑。Bao 的旧版本历史保留,回滚基础配置时模板也会替换其中的旧 zot 身份。
专用 S3 身份只有 `Read:zot`、`Write:zot`、`List:zot`、`Tagging:zot`,不能读取
Terraform 的 `tfstate` bucket。`kv/k8s/zot-s3` 的字段是 `access_key` 和
`secret_key`。AK/SK 不进入 Git、Helm values 或 CI;这里仍是静态 S3 凭据,尚未
接入 SPIRE/STS。
本次归一没有轮换密钥,生成的完整配置与归一前语义一致。当前模板只有一组 zot
凭据,尚未实现新旧密钥重叠轮换。后续轮换只修改 `kv/k8s/zot-s3`,但仍需协调
两个 ExternalSecret 同步:确认 SeaweedFS Secret volume 更新后向 filer 的
`weed` 进程发送 SIGHUP,再确认 zot Secret 更新并重启 zot(环境变量不会热更新)。
两端异步更新期间可能短暂认证失败;需要无中断轮换时先扩展模板支持新旧凭据重叠。
## SPIRE 认证和授权
| 参数 | 值 |
|---|---|
| issuer | `https://spire-oidc.ad.ddupan.top` |
| JWT audience | `zot` |
| subject | `spiffe://ddupan.top/` 下的 workload SPIFFE ID |
| token endpoint | `https://zot.ad.ddupan.top/zot/auth/token` |
| 当前权限 | 受信身份可以读取所有仓库;没有常驻写入或删除授权 |
zot 通过已配置的 issuer discovery/JWKS 验证 JWT-SVID,再以 `sub` 作为授权身份。
不接受任意 issuer,不关闭 TLS/issuer 验证。新的 Kata CI 负责取得并更新自己的
JWT-SVID;确认其身份命名后,再添加针对具体 repository 的 `create`/`update`
授权,不能把整个 trust domain 都授予写权限。
现阶段拉取也需要 JWT-SVID。原定内网匿名拉取尚未启用:zot `v2.1.21` 的
OIDC Bearer middleware 会在授权阶段之前拒绝无 token 请求,单独增加
`anonymousPolicy` 无法解决。匿名读取与 SPIRE 写入共存需后续单独验证方案。
已有 SPIRE 身份的进程可以通过 Workload API 获取 `aud=zot` 的 JWT-SVID,然后
通过 `docker login` 或 `crane auth login` 的 `--password-stdin` 交给 Registry。
用户名可以使用 `zot`,实际权限取自已验证 JWT 的身份。使用独立、权限为 `0700`
的临时 `DOCKER_CONFIG`,结束后删除;不要开启 shell tracing,不要打印 token,
不要把 token 放进命令参数。token 接口不会延长 SVID 有效期。
## 部署与网络
- 官方 chart 管理 Deployment、Service、ConfigMap 和 HTTPRoute。
- `persistence: false`,Service 为 ClusterIP,TLS 由已有 Envoy Gateway 的
`https` listener 与内网通配符证书终止。
- 仅配置 Samba AD 内网 DNS;不创建公网 DNS 或 Cloudflare Tunnel route。
- NetworkPolicy 只允许现有 Envoy Gateway 数据面访问 zot 的 5000 端口。
- HTTPRoute 只暴露 `/v2/` 和 `/zot/auth/token`,不暴露内部健康检查或管理端点。
- namespace 使用 restricted PodSecurity,容器非 root、只读根文件系统。
首次已按用户授权从本地执行 `kubectl apply -k apps/zot`,由集群 Helm controller
安装。`clusters/homelab/apps/zot.yaml` 是 GitOps composition;对应文件合并进入
Flux 跟踪分支后,才由根 Kustomization 持续管理,不能把未提交的本地部署写成
已完成 Git 接管。
检查与渲染:
```bash
helm template zot --repo https://zotregistry.dev/helm-charts \
--version 0.1.124 --namespace zot -f apps/zot/values.yaml --skip-tests
sudo k3s kubectl -n zot get helmrelease,pods,externalsecret,httproute
sudo k3s kubectl -n zot get pvc
```
上游 chart 的 Helm test Pod 不满足本 namespace 的 restricted 策略,也没有
SPIRE 凭据,因此不运行默认 `helm test`;使用下述真实身份验收。
## 验收与恢复
验收使用独立临时 Pod,通过 SPIFFE CSI socket 和真实 Workload API 取得 JWT-SVID,
没有修改现有 runner。仅在初始化 `verification/smoke:spire-s3` 测试镜像时临时
授予该测试身份针对该仓库的写权限;完成后必须撤回 HelmRelease override,并删除
临时 Pod、ServiceAccount 与 ClusterSPIFFEID。
验收项目:有效 SVID + crane pull、manifest digest 一致、错误 audience、错误
signature、过期 token、无凭据写入、跨仓库写入、只读身份写入和删除拒绝;另外检查
Pod 重建后镜像仍可拉取,以及 S3 身份不能访问 `tfstate`。
2026-09-14 已完成上述验收:HelmRelease Ready、HTTPRoute Accepted/ResolvedRefs,
DNS 第二次 Ansible check 为 `changed=0`;一分钟真实 JWT-SVID 到期后返回 401。
临时写权限已移除。测试镜像可供后续 CI 验证拉取:
```text
zot.ad.ddupan.top/verification/smoke:spire-s3
sha256:b8d3b977a1235022759470903dab4a46b7cf8107958624f1f76a323eabe37c5e
```
它是仅含验证文本的 OCI 测试镜像,没有可执行入口,不用于运行服务。
Registry 恢复需要完整的 SeaweedFS bucket 数据、Bao 专用凭据和此目录配置。
zot 的临时目录不是制品备份。独立异机/离线备份尚未在本次部署中建立;不能把同一
SeaweedFS 内的数据副本当作独立灾备。重装 zot 不得删除 `zot` bucket。
参考:[官方 Kubernetes 安装](https://zotregistry.dev/v2.1.21/install-guides/install-guide-k8s/)、
[S3 存储](https://zotregistry.dev/v2.1.21/articles/storage/)、
[OIDC workload identity](https://github.com/project-zot/zot/blob/v2.1.21/examples/README-OIDC-WORKLOAD-IDENTITY.md)。
+22
View File
@@ -0,0 +1,22 @@
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: zot-s3
namespace: zot
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: openbao
target:
name: zot-s3
creationPolicy: Owner
data:
- secretKey: access_key
remoteRef:
key: k8s/zot-s3
property: access_key
- secretKey: secret_key
remoteRef:
key: k8s/zot-s3
property: secret_key
+30
View File
@@ -0,0 +1,30 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: zot
namespace: zot
spec:
chart:
spec:
chart: zot
version: 0.1.124
interval: 1h
sourceRef:
kind: HelmRepository
name: zot
releaseName: zot
interval: 30m
timeout: 5m
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: zot-values
+8
View File
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: zot
namespace: zot
spec:
interval: 1h
url: https://zotregistry.dev/helm-charts
+18
View File
@@ -0,0 +1,18 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- serviceaccount.yaml
- external-secret.yaml
- helmrepository.yaml
- helmrelease.yaml
- networkpolicy.yaml
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: zot-values
namespace: zot
files:
- values.yaml=values.yaml
+7
View File
@@ -0,0 +1,7 @@
apiVersion: v1
kind: Namespace
metadata:
name: zot
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/enforce-version: v1.36
+22
View File
@@ -0,0 +1,22 @@
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: zot-ingress
namespace: zot
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: zot
policyTypes: [Ingress]
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: envoy-gateway-system
podSelector:
matchLabels:
gateway.envoyproxy.io/owning-gateway-name: eg
gateway.envoyproxy.io/owning-gateway-namespace: envoy-gateway-system
ports:
- protocol: TCP
port: 5000
+6
View File
@@ -0,0 +1,6 @@
apiVersion: v1
kind: ServiceAccount
metadata:
name: zot
namespace: zot
automountServiceAccountToken: false
+147
View File
@@ -0,0 +1,147 @@
# 官方 chart 0.1.124 / zot v2.1.21;制品与 manifests 保存在 SeaweedFS S3。
# persistence=false 仅保留 chart 的 emptyDir,不创建 PVC。
# 首期关闭跨仓库 dedupe,不额外引入 Redis/DynamoDB 持久缓存。
replicaCount: 1
image:
repository: ghcr.io/project-zot/zot
tag: v2.1.21@sha256:8258443838e95989c13c891f78a02bc1c391b5a00591ffef24cb8c17cde28038
persistence: false
strategy:
type: Recreate
serviceAccount:
create: false
name: zot
service:
type: ClusterIP
port: 5000
mountConfig: true
mountSecret: false
secretFiles: {}
configFiles:
config.json: |
{
"distSpecVersion": "1.1.1",
"storage": {
"rootDirectory": "/var/lib/registry",
"dedupe": false,
"gc": true,
"gcDelay": "24h",
"gcInterval": "24h",
"storageDriver": {
"name": "s3",
"region": "us-east-1",
"regionendpoint": "https://s3.ad.ddupan.top",
"bucket": "zot",
"rootdirectory": "/registry",
"secure": true,
"skipverify": false,
"forcepathstyle": true
}
},
"http": {
"address": "0.0.0.0",
"port": "5000",
"externalUrl": "https://zot.ad.ddupan.top",
"compat": [
"docker2s2"
],
"auth": {
"bearer": {
"realm": "https://zot.ad.ddupan.top/zot/auth/token",
"service": "zot.ad.ddupan.top",
"oidc": [
{
"issuer": "https://spire-oidc.ad.ddupan.top",
"audiences": [
"zot"
],
"claimMapping": {
"username": "claims.sub",
"validations": [
{
"expression": "claims.sub.startsWith('spiffe://ddupan.top/')",
"message": "SPIFFE trust domain mismatch"
}
]
}
}
]
}
},
"accessControl": {
"repositories": {
"**": {
"defaultPolicy": [
"read"
]
}
}
}
},
"log": {
"level": "info"
}
}
env:
- name: AWS_ACCESS_KEY_ID
valueFrom:
secretKeyRef:
name: zot-s3
key: access_key
- name: AWS_SECRET_ACCESS_KEY
valueFrom:
secretKeyRef:
name: zot-s3
key: secret_key
- name: AWS_EC2_METADATA_DISABLED
value: 'true'
podSecurityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: '1'
memory: 512Mi
extraVolumes:
- name: tmp
emptyDir:
sizeLimit: 128Mi
extraVolumeMounts:
- name: tmp
mountPath: /tmp
startupProbe:
initialDelaySeconds: 5
periodSeconds: 5
failureThreshold: 60
httproute:
enabled: true
parentRefs:
- name: eg
namespace: envoy-gateway-system
sectionName: https
hostnames:
- zot.ad.ddupan.top
rules:
- matches:
- path:
type: PathPrefix
value: /v2/
- path:
type: Exact
value: /zot/auth/token
timeouts:
request: 900s
backendRequest: 900s
+18
View File
@@ -0,0 +1,18 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: zot
namespace: flux-system
spec:
dependsOn:
- name: envoy-gateway
- name: external-secrets
- name: spire
interval: 10m
path: ./apps/zot
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 5m
wait: true
+1
View File
@@ -12,3 +12,4 @@ resources:
- apps/openebs.yaml
- apps/spire.yaml
- apps/observability.yaml
- apps/zot.yaml
+1
View File
@@ -14,6 +14,7 @@ homelab_dns:
- { zone: ad.ddupan.top, name: netbox, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: s3, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-oidc, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: zot, type: A, values: [192.168.10.127] }
split_horizon:
# LAN and pod resolvers should eventually render the same set from here.
+6
View File
@@ -36,6 +36,7 @@ configure any secrets engines/auth methods — that is a separate bootstrap play
```
terraform/ # OpenBao's API-level CONFIGURATION (see below)
mounts.tf pki.tf ssh.tf auth.tf policies.tf
auth-spire.tf # SPIFFE JWT-SVID -> short-lived Bao tokens
imports.tf # adopts the already-running instance into state
policies/*.hcl # policy bodies, kept diffable
```
@@ -205,6 +206,11 @@ after this, use `BAO_ADDR=https://bao.ad.ddupan.top:8200` (no skip-verify), not
## Using it
Kubernetes workload 不接收长期 `BAO_TOKEN`:它通过 SPIRE Workload API 获取
JWT-SVID,再经 `auth/jwt-spire/login` 换取短期、最小权限 token。完整接入流程、
manifest、exchange 脚本、安全要求和排障方法见
[`../../platform/spire/RUNBOOK.md`](../../platform/spire/RUNBOOK.md)。
```bash
# human: log in via Authelia (2FA)
bao login -method=oidc # browser → auth.ddupan.top
@@ -0,0 +1,27 @@
# Workload authentication via SPIFFE JWT-SVIDs. The discovery document and
# JWKS contain public verification material, so this backend carries no secret.
resource "vault_jwt_auth_backend" "spire" {
path = "jwt-spire"
description = "SPIFFE JWT-SVID workload authentication"
oidc_discovery_url = "https://spire-oidc.ad.ddupan.top"
bound_issuer = "https://spire-oidc.ad.ddupan.top"
}
# First end-to-end identity. Keep the subject exact: this role is deliberately
# not a wildcard escape hatch for every workload in the trust domain.
resource "vault_jwt_auth_backend_role" "spire_poc" {
backend = vault_jwt_auth_backend.spire.path
role_name = "spire-poc"
role_type = "jwt"
user_claim = "sub"
bound_audiences = ["openbao"]
bound_claims = {
sub = "spiffe://ddupan.top/ns/spire-poc/sa/spire-jwt-poc"
}
token_policies = [vault_policy.spire_poc.name]
token_no_default_policy = true
token_ttl = 300
token_max_ttl = 900
}
@@ -14,6 +14,11 @@ resource "vault_policy" "ai_agent_ssh" {
policy = file("${path.module}/policies/ai-agent-ssh.hcl")
}
resource "vault_policy" "spire_poc" {
name = "spire-poc"
policy = file("${path.module}/policies/spire-poc.hcl")
}
resource "vault_policy" "snapshot" {
name = "snapshot"
policy = file("${path.module}/policies/snapshot.hcl")
@@ -0,0 +1,10 @@
# Intentionally grants no secret access. This policy proves that an exact
# SPIFFE ID can exchange a JWT-SVID for a bounded OpenBao token and inspect or
# revoke only that token.
path "auth/token/lookup-self" {
capabilities = ["read"]
}
path "auth/token/revoke-self" {
capabilities = ["update"]
}
+35 -2
View File
@@ -101,6 +101,7 @@ spec:
# The chart normally GENERATES seaweedfs-s3-secret from s3.credentials. We point
# filer.s3.existingConfigSecret at this one instead, so the chart stops rendering
# credentials from values entirely.
# zot 的 AK/SK 只保存在 k8s/zot-s3;在此组装服务端配置,不在基础配置中维护副本。
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
@@ -114,6 +115,38 @@ spec:
target:
name: seaweedfs-s3-config
creationPolicy: Owner
dataFrom:
- extract:
template:
engineVersion: v2
mergePolicy: Replace
data:
seaweedfs_s3_config: |-
{{- $config := mustFromJson .baseConfig -}}
{{- if not (kindIs "slice" $config.identities) -}}
{{- fail "base S3 configuration must contain an identities array" -}}
{{- end -}}
{{- if or (eq .zotAccessKey "") (eq .zotSecretKey "") -}}
{{- fail "zot S3 credentials must not be empty" -}}
{{- end -}}
{{- $identities := list -}}
{{- range $config.identities -}}
{{- if ne .name "zot" -}}
{{- $identities = append $identities . -}}
{{- end -}}
{{- end -}}
{{- $credential := dict "accessKey" .zotAccessKey "secretKey" .zotSecretKey -}}
{{- $zot := dict "name" "zot" "credentials" (list $credential) "actions" (list "Read:zot" "Write:zot" "List:zot" "Tagging:zot") -}}
{{- $_ := set $config "identities" (append $identities $zot) -}}
{{- mustToJson $config -}}
data:
- secretKey: baseConfig
remoteRef:
key: k8s/seaweedfs-s3
property: seaweedfs_s3_config
- secretKey: zotAccessKey
remoteRef:
key: k8s/zot-s3
property: access_key
- secretKey: zotSecretKey
remoteRef:
key: k8s/zot-s3
property: secret_key
+11 -2
View File
@@ -3,6 +3,9 @@
SPIRE 是 homelab 的机器与 workload identity 根。人类身份继续由 Samba AD 与
Authelia 提供;SPIRE 不替代人类 OIDC,也不承担目标服务的资源授权。
部署状态、workload 接入、JWT-SVID → OpenBao exchange、安全规则和故障恢复详见
[RUNBOOK.md](RUNBOOK.md)。本文只保留部署声明与关键恢复边界。
## 部署范围
Flux 安装 SPIFFE hardened charts:
@@ -73,8 +76,14 @@ sudo k3s kubectl -n spire-system get daemonset,pods
必须先确认 `spire-crds` Ready,随后 `spire` Ready。SPIRE Server 应连接 PostgreSQL,
Agent 应通过 PSAT attestation 注册,CSI Driver 应在节点 Ready。
首个业务验收另行增加一个专用测试 Pod 与 `ClusterSPIFFEID`,验证取得
`aud=openbao` 的 JWT-SVID 后登录 OpenBao。PoC 完成前不修改生产认证方式。
2026-09-14 已使用临时测试 Pod 与 `ClusterSPIFFEID` 完成
`aud=openbao` JWT-SVID → OpenBao 登录、token 自省和主动吊销的端到端验收;临时
Kubernetes 资源与 registration entries 已清理。
OpenBao 中对应的 Terraform 资源位于
`../../infrastructure/openbao/terraform/auth-spire.tf`。PoC role 只接受精确 subject
`spiffe://ddupan.top/ns/spire-poc/sa/spire-jwt-poc`,token 不包含 default policy,
且 `spire-poc` policy 不允许读取任何业务 secret。
## 恢复边界
+436
View File
@@ -0,0 +1,436 @@
# SPIRE 与 OpenBao workload identity runbook
本文记录 homelab 中 workload 如何取得 SPIFFE 身份、如何把 JWT-SVID 交换成
OpenBao 短期 token,以及相关的部署、接入、验证和恢复操作。这里的命令默认在
`laptop` 上执行。
## 1. 当前架构
```text
Kubernetes Pod
│ Pod ServiceAccount + Pod UID
▼
SPIRE Agent(每节点 DaemonSet,k8s workload attestor)
│ Unix socket: /spiffe-workload-api/spire-agent.sock
│ Agent 自身通过 k8s_psat 向 Server 证明节点身份
▼
SPIRE Server(trust domain: ddupan.top)
├─ registration state → shared PostgreSQL
├─ signing keys → localpv-zfs-ceph PVC
└─ JWT public keys → OIDC Discovery Provider
│
▼
https://spire-oidc.ad.ddupan.top
│ discovery + JWKS
▼
OpenBao auth/jwt-spire/login
│ exact sub + aud + role
▼
短期、最小权限 Bao token
```
各层职责必须保持分离:
- Kubernetes ServiceAccount 是 Pod 的初始身份证明,不是跨平台 IAM token;
- SPIRE 负责 workload 身份、证明和 SVID 签发,不保存业务 secret;
- OpenBao 验证 JWT-SVID,并把身份映射为本地 policy;
- 目标服务最终仍负责自己的资源授权,SPIFFE ID 本身不等于权限;
- 人类身份继续使用 Samba AD + Authelia OIDC,不经过 `jwt-spire`。
## 2. 线上对象与稳定标识
| 项目 | 当前值 |
|---|---|
| SPIRE chart | `0.30.2` |
| SPIRE | `1.15.3` |
| SPIRE CRDs chart | `0.6.1` |
| trust domain | `ddupan.top` |
| Kubernetes cluster name | `homelab` |
| Controller Manager class | `spire-mgmt-spire` |
| JWT issuer | `https://spire-oidc.ad.ddupan.top` |
| OpenBao auth mount | `jwt-spire` |
| OpenBao login endpoint | `auth/jwt-spire/login` |
| SPIRE Server namespace | `spire-server` |
| Agent/CSI namespace | `spire-system` |
| Helm management namespace | `spire-mgmt` |
`trustDomain`、`clusterName`、issuer URL 与 Controller Manager class 都进入身份或
下游信任配置。修改它们不是普通 rename,必须按 trust-domain migration 处理。
## 3. 身份与授权模型
Kubernetes workload 的默认 SPIFFE ID 约定为:
```text
spiffe://ddupan.top/ns/<namespace>/sa/<service-account>
```
身份必须同时在两侧声明:
1. SPIRE `ClusterSPIFFEID` 决定哪些 Pod 可以取得该身份;
2. OpenBao JWT role 决定该 `sub`、`aud` 能换取哪些 policy。
这两个声明是有意的双重门:只有 SPIRE entry 而没有 Bao role 时,workload 能取得
SVID,但不能登录 Bao;只有 Bao role 而没有 SPIRE entry 时,没有 workload 能铸造
满足条件的 JWT。
禁止使用以下宽泛规则:
- 给所有 Pod 启用 fallback `ClusterSPIFFEID`;
- OpenBao role 接受整个 `spiffe://ddupan.top/*`;
- 仅按 namespace 匹配高权限身份,却不限制 ServiceAccount 和 Pod labels;
- 多个安全边界不同的 workload 共用同一个 ServiceAccount;
- 给 workload token 附带 `default` 或 `admin` policy。
## 4. 新 workload 接入流程
以下示例为 namespace `example` 中的 ServiceAccount `example-worker`。
### 4.1 创建专用 ServiceAccount
```yaml
apiVersion: v1
kind: ServiceAccount
metadata:
name: example-worker
namespace: example
```
不要使用 namespace 的 `default` ServiceAccount。
### 4.2 声明 ClusterSPIFFEID
```yaml
apiVersion: spire.spiffe.io/v1alpha1
kind: ClusterSPIFFEID
metadata:
name: example-worker
labels:
spire.spiffe.io/class-name: spire-mgmt-spire
spec:
className: spire-mgmt-spire
namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: example
podSelector:
matchLabels:
app.kubernetes.io/name: example-worker
spiffeIDTemplate: "spiffe://{{ .TrustDomain }}/ns/{{ .PodMeta.Namespace }}/sa/{{ .PodSpec.ServiceAccountName }}"
```
Controller Manager 最终会为具体 Pod UID 创建 registration entry。检查状态:
```bash
sudo k3s kubectl get clusterspiffeid example-worker -o yaml
sudo k3s kubectl -n spire-server exec statefulset/spire-server -c spire-server -- \
/opt/spire/bin/spire-server entry show \
-spiffeID spiffe://ddupan.top/ns/example/sa/example-worker
```
`status.stats.entryFailures` 必须是 `0`。Pod 重建后 UID 会变化,短暂看到旧 entry
属于正常收敛过程。
### 4.3 挂载 Workload API
```yaml
spec:
serviceAccountName: example-worker
containers:
- name: worker
volumeMounts:
- name: spiffe-workload-api
mountPath: /spiffe-workload-api
readOnly: true
env:
- name: SPIFFE_ENDPOINT_SOCKET
value: unix:///spiffe-workload-api/spire-agent.sock
volumes:
- name: spiffe-workload-api
csi:
driver: csi.spiffe.io
readOnly: true
```
挂载 socket 不会自动获得身份。Agent 会对调用进程执行 workload attestation,只有
selector 命中 registration entry 才签发 SVID。
应用应优先使用 SPIFFE SDK,通过 Workload API 按需取得并自动轮换 SVID。不要把
JWT-SVID 写入 Kubernetes Secret、镜像、持久卷、CI artifact 或日志。
### 4.4 声明 OpenBao policy
在 `infrastructure/openbao/terraform/policies/` 中为 workload 建独立 policy。例如:
```hcl
path "kv/data/apps/example/*" {
capabilities = ["read"]
}
path "auth/token/lookup-self" {
capabilities = ["read"]
}
path "auth/token/revoke-self" {
capabilities = ["update"]
}
```
不要直接复用 `admin`。动态数据库凭据、SSH 签名和 KV 应分别授权到精确 path。
### 4.5 声明 OpenBao JWT role
在 `infrastructure/openbao/terraform/auth-spire.tf` 增加 role:
```hcl
resource "vault_jwt_auth_backend_role" "example_worker" {
backend = vault_jwt_auth_backend.spire.path
role_name = "example-worker"
role_type = "jwt"
user_claim = "sub"
bound_audiences = ["openbao"]
bound_claims = {
sub = "spiffe://ddupan.top/ns/example/sa/example-worker"
}
token_policies = [vault_policy.example_worker.name]
token_no_default_policy = true
token_ttl = 300
token_max_ttl = 900
}
```
同一个 JWT 可以请求多个 audience,但 Bao role 只接受 `openbao`。未来接入其他服务时
应给对应服务使用独立 audience,不能把 `openbao` 当作通用 audience。
Terraform apply 必须遵循 `infrastructure/openbao/README.md` 的 remote-state 和认证
流程。任何包含 destroy/replace 的 plan 都应停止审查;JWT role/policy 的正常新增应为
纯 `add`。
## 5. JWT-SVID 交换流程
逻辑请求如下:
```http
POST /v1/auth/jwt-spire/login
Content-Type: application/json
{
"role": "example-worker",
"jwt": "<aud=openbao 的 JWT-SVID>"
}
```
成功响应中的 `auth.client_token` 是短期 Bao token。它只应存在于进程内存或
job-scoped `tmpfs`;通常无需主动续期,过期前重新用 Workload API 获取 JWT-SVID 并
登录即可。
如果使用 SPIRE CLI 调试,`1.15.3` 的 JSON 输出顶层是数组,JWT 位于:
```text
.[0].svids[0].svid
```
调试脚本必须把 JSON 捕获到变量中,禁止直接输出:
```bash
set -euo pipefail
JWT_RESPONSE="$(spire-agent api fetch jwt \
-audience openbao \
-socketPath /spiffe-workload-api/spire-agent.sock \
-output json)"
JWT_SVID="$(printf '%s' "$JWT_RESPONSE" | jq -er '.[0].svids[0].svid')"
LOGIN_PAYLOAD="$(jq -nc \
--arg role example-worker \
--arg jwt "$JWT_SVID" \
'{role:$role,jwt:$jwt}')"
LOGIN_RESPONSE="$(curl --fail-with-body --silent --show-error \
-H 'Content-Type: application/json' \
--data "$LOGIN_PAYLOAD" \
https://bao.ad.ddupan.top:8200/v1/auth/jwt-spire/login)"
export BAO_ADDR=https://bao.ad.ddupan.top:8200
export BAO_TOKEN="$(printf '%s' "$LOGIN_RESPONSE" | jq -er '.auth.client_token')"
```
不要在 shell 中启用 `set -x`,不要 `echo "$JWT_SVID"` 或输出完整 login response。
cleanup 阶段可以尽力主动吊销;短 TTL 仍是主要安全边界:
```bash
bao token revoke -self
unset BAO_TOKEN JWT_SVID JWT_RESPONSE LOGIN_RESPONSE LOGIN_PAYLOAD
```
## 6. CI 与 AI Agent 使用方式
CI job/Agent 不应接收长期 `BAO_TOKEN`。标准启动顺序是:
1. 调度到带 SPIRE Agent 与 CSI Driver 的节点;
2. 以专用 ServiceAccount 启动,挂载 Workload API socket;
3. 获取目标 audience 的 JWT-SVID;
4. 用对应 Bao role 换取短期 token;
5. 在同一进程树中以环境变量调用 `tofu`、Ansible 或其他工具;
6. cleanup 尝试 `revoke-self`,随后销毁 job/VM/容器。
通用 credential-exec 包装器未来应负责步骤 3–6。它必须满足:
- 不把 JWT-SVID 或 Bao token 写到 stdout/stderr;
- 不把凭据传入命令行参数,避免出现在进程列表;
- 子进程退出后清除环境和临时文件;
- 不尝试把短期 token 上传到 Actions Secret 或 artifact;
- role、audience 和目标命令由受审查的 pipeline 配置决定。
Kubernetes 以外的执行环境不能伪造 ServiceAccount。未来应分别使用 host SPIRE
Agent、TPM/DevID、cloud instance identity、GitHub OIDC 等初始证明接入同一信任模型。
## 7. 日常检查
### Flux 与 Helm
```bash
sudo k3s kubectl -n flux-system get kustomization spire
sudo k3s kubectl -n spire-mgmt get helmrepository,helmrelease
```
### Server、Agent 与 CSI
```bash
sudo k3s kubectl -n spire-server get pods,pvc
sudo k3s kubectl -n spire-system get daemonset,pods
sudo k3s kubectl -n spire-server exec statefulset/spire-server -c spire-server -- \
/opt/spire/bin/spire-server agent list
```
Agent 应显示 `Attestation type: k8s_psat` 与 `Can re-attest: true`。
### OIDC discovery 与 JWKS
```bash
dig @192.168.10.5 spire-oidc.ad.ddupan.top A +short
dig @192.168.10.127 spire-oidc.ad.ddupan.top A +short
curl --fail --silent \
https://spire-oidc.ad.ddupan.top/.well-known/openid-configuration | jq
curl --fail --silent https://spire-oidc.ad.ddupan.top/keys | jq '.keys | length'
```
discovery 的 `issuer` 必须严格等于
`https://spire-oidc.ad.ddupan.top`。路径、scheme、hostname 或尾部 `/` 的差异都会
导致 JWT 验证失败。JWKS 是公开验证材料,不是 secret。
### OpenBao
管理员只检查非敏感配置:
```bash
export BAO_ADDR=https://bao.ad.ddupan.top:8200
bao auth list
bao read auth/jwt-spire/role/<role-name>
bao policy read <policy-name>
```
## 8. 故障排查
### `no identity issued`
依次检查:
```bash
sudo k3s kubectl get clusterspiffeid <name> -o yaml
sudo k3s kubectl -n <namespace> get pod <pod> \
-o custom-columns=NAME:.metadata.name,UID:.metadata.uid,SA:.spec.serviceAccountName,LABELS:.metadata.labels
sudo k3s kubectl -n spire-server exec statefulset/spire-server -c spire-server -- \
/opt/spire/bin/spire-server entry show -spiffeID <expected-spiffe-id>
```
确认 namespace、ServiceAccount、Pod labels 与 entry 中的 `k8s:pod-uid`。刚启动的
一次性 Job 可能在 Controller Manager 建 entry 前请求身份并失败;生产客户端应重试
Workload API,而不是假设 Pod 一启动身份就已可用。
### `auth/jwt-spire/login` 返回 400
常见原因:
- JWT audience 不是 role 的 `bound_audiences`;
- JWT `sub` 与 role 的 `bound_claims.sub` 不完全一致;
- issuer 与 `bound_issuer` 不一致;
- Bao 无法解析或验证 `spire-oidc.ad.ddupan.top`;
- issuer/JWKS route 被 Authelia forward-auth 拦截;
- SVID 已过期,或者节点与 Bao 时钟偏差过大;
- 脚本错误解析 CLI JSON,向 Bao 发送了空 JWT。
排查时只解码 JWT header/claims,禁止记录原始 token。公开 endpoint 可单独验证:
```bash
curl --fail https://spire-oidc.ad.ddupan.top/.well-known/openid-configuration
curl --fail https://spire-oidc.ad.ddupan.top/keys
```
### CSI mount 失败
如果事件包含 `driver name csi.spiffe.io not found`:
```bash
sudo k3s kubectl get csidriver csi.spiffe.io
sudo k3s kubectl -n spire-system get daemonset spire-spiffe-csi-driver
sudo k3s kubectl -n spire-system logs daemonset/spire-spiffe-csi-driver \
-c spiffe-csi-driver --tail=100
```
首次安装时 OIDC Provider/业务 Pod 可能早于 CSI registration,短暂 mount retry 正常;
持续失败才需要处理。
### Server 无法连接 PostgreSQL
```bash
sudo k3s kubectl -n shared-db get cluster shared-postgresql
sudo k3s kubectl -n spire-server get secret spire-postgresql \
-o go-template='{{range $k, $_ := .data}}{{$k}}{{"\n"}}{{end}}'
sudo k3s kubectl -n spire-server logs statefulset/spire-server \
-c spire-server --tail=100
```
Secret 必须有 `password` key。禁止为排障直接输出或提交其值。
## 9. 轮换、备份与恢复
SPIRE 会自动轮换 X.509 CA 与 JWT signing key;OIDC Provider 从 Workload API 获取
当前 JWKS,下游按 `kid` 验证。正常轮换不应要求更新 Bao role。
必须备份两类不同状态:
- PostgreSQL:registration entries、agent state 和 SPIRE metadata;
- `spire-data-spire-server-0` PVC:disk KeyManager 的 trust-domain signing keys。
仅恢复 PostgreSQL 而丢失 PVC,不等于恢复 SPIRE。签名密钥丢失会使既有 SVID/JWKS
信任链失效。恢复顺序:
1. 恢复共享 PostgreSQL;
2. 恢复 signing-key PVC;
3. 启动 SPIRE Server;
4. 确认 bundle/JWKS 后启动或恢复 Agent;
5. 最后恢复依赖 SPIRE 登录 Bao 的 workload。
不要以空数据库或空 PVC“修复”启动失败。若确实需要重建 trust domain,应把它作为
全体下游重新建立信任的灾难恢复事件处理。
## 10. 当前 PoC 结论与后续工作
2026-09-14 已完成并清理一次临时 PoC:
- workload 取得 `aud=openbao` JWT-SVID;
- 精确 subject 成功登录 `jwt-spire/spire-poc`;
- 返回 token 仅含 `spire-poc` policy,TTL 为 300 秒,无 default policy;
- `lookup-self` 成功,随后 `revoke-self` 并验证 token 已失效;
- 临时 Namespace、Pod/Job、ServiceAccount、`ClusterSPIFFEID` 与 registration entries
均已删除;
- Terraform 完整 plan 最终为 `No changes`。
下一步不是重复 PoC,而是为真实 CI/AI Agent 定义:
- 独立 ServiceAccount 和稳定 SPIFFE ID;
- 按能力拆分的 OpenBao policy(例如 SSH CA、S3 state、数据库动态凭据);
- 通用 credential-exec 包装器;
- token 获取失败、过期与 cleanup 的客户端重试语义;
- Kubernetes 外 host/VM/microVM 的 SPIRE Agent attestation 方案。