Author SHA1 Message Date
panxiao81 f00b5e3fc2 排除 OCI HelmRepository 正常缺失 Ready 条件的误报 2026-09-25 19:39:04 +00:00
panxiao81 9a800f55fc 补齐证书、秘密同步、GitOps 与数据服务监控
yaml / yaml (pull_request) Successful in 3m20s
2026-09-25 19:35:38 +00:00
panxiao81 18ded1e4a8 Merge 原子部署解耦后的 dynamic runner 镜像
yaml / yaml (push) Successful in 4m32s
2026-09-25 19:12:07 +00:00
panxiao81 ca0862fe5a fix: 原子部署解耦后的 controller 与 runner
yaml / yaml (pull_request) Failing after 10m54s
2026-09-25 19:11:39 +00:00
panxiao81 a729f01546 Merge 部署修复后的 dynamic runner placement v2
yaml / yaml (push) Successful in 17s
2026-09-25 18:54:41 +00:00
panxiao81 bbcd5d6aa5 fix: 部署合法 consumer 名称的 placement v2
yaml / yaml (pull_request) Successful in 20s
2026-09-25 18:53:23 +00:00
panxiao81 c0c4ff521b Merge pull request '部署 dynamic runner placement v2' (#154) from deploy/workload-placement-v2 into main
yaml / yaml (push) Successful in 31s
Reviewed-on: #154
2026-09-25 18:37:40 +00:00
panxiao81 8bc039b40f Merge pull request #155: 授予 Backstage Kubernetes 只读观察权限
yaml / yaml (push) Successful in 27s
2026-09-25 18:33:32 +00:00
panxiao81 349aa77fb0 ops: 暂停 placement v2 自动 rollout
yaml / yaml (pull_request) Successful in 50s
2026-09-25 18:30:55 +00:00
panxiao81 cae6acbfb0 feat: 授予 Backstage Kubernetes 只读观察权限
yaml / yaml (pull_request) Successful in 43s
2026-09-25 18:30:42 +00:00
panxiao81 4d80f8efd7 deploy: 切换 dynamic runner placement v2
yaml / yaml (pull_request) Successful in 42s
2026-09-25 18:11:05 +00:00
panxiao81 9009e3fa11 启用 Backstage 显式目录位置 (#153)
yaml / yaml (push) Successful in 49s
2026-09-25 18:07:37 +00:00
panxiao81 a252ee5602 fix: 启用 Backstage 显式目录位置
yaml / yaml (pull_request) Successful in 47s
2026-09-25 18:06:59 +00:00
panxiao81 dbf66590ac 修复告警 Source 链接并接入 Grafana Alertmanager 数据源 (#152)
yaml / yaml (push) Successful in 51s
Co-authored-by: panxiao81 <[email protected]>
2026-09-25 18:02:14 +00:00
panxiao81 e8e7bd4b20 优先使用 PrometheusRule 声明基础告警 (#151)
yaml / yaml (push) Successful in 51s
Co-authored-by: panxiao81 <[email protected]>
2026-09-25 17:53:07 +00:00
panxiao81 185e3edaba 修复 Backstage 首次启动 (#150)
yaml / yaml (push) Successful in 1m13s
2026-09-25 17:47:34 +00:00
panxiao81 072936ce03 fix: 修正 Backstage 数据库模式与健康探针
yaml / yaml (pull_request) Successful in 54s
2026-09-25 17:46:49 +00:00
panxiao81 39aeac25ee 补齐集群状态、主机与通知链路基础监控 (#149)
yaml / yaml (push) Successful in 1m10s
Co-authored-by: panxiao81 <[email protected]>
2026-09-25 17:45:56 +00:00
panxiao81 a43b7d3c26 修复 kubelet 超限与旧 Docker 采集目标 (#148)
yaml / yaml (push) Failing after 10m32s
Co-authored-by: panxiao81 <[email protected]>
2026-09-25 17:22:01 +00:00
panxiao81 c06c6f7857 启用 Alertmanager Telegram 频道告警 (#147)
yaml / yaml (push) Successful in 39s
Co-authored-by: panxiao81 <[email protected]>
2026-09-25 17:10:30 +00:00
panxiao81 6a27adc18d 部署 Backstage 门户第一版 (#146)
yaml / yaml (push) Successful in 51s
2026-09-25 17:08:01 +00:00
panxiao81 29b3bdeeb1 feat: 部署 Backstage 门户第一版
yaml / yaml (pull_request) Successful in 22s
2026-09-25 16:09:23 +00:00
panxiao81 7a62dde601 Merge pull request '为 Backstage 镜像仓库配置 SPIFFE 发布权限' (#145) from feat/zot-backstage-access into main
yaml / yaml (push) Successful in 25s
Reviewed-on: #145
2026-09-25 15:26:38 +00:00
panxiao81 521a49599a 为 Backstage 镜像仓库配置 SPIFFE 发布权限
yaml / yaml (pull_request) Successful in 30s
2026-09-25 15:23:56 +00:00
panxiao81 185a47e2bb 合并 zot 本地维护身份删除权限修正
yaml / yaml (push) Successful in 25s
2026-09-25 14:53:58 +00:00
panxiao81 8dabbdce68 修复 zot 本地维护身份在 CI 镜像仓库的删除权限
yaml / yaml (pull_request) Successful in 25s
2026-09-25 14:52:14 +00:00
panxiao81 d5b1bb9640 为 Gitea 添加 Hydra OIDC 登录入口
yaml / yaml (push) Successful in 25s
2026-09-25 13:56:23 +00:00
panxiao81 cfb2fddfa9 feat: 为 Gitea 添加 Hydra OIDC 登录入口
yaml / yaml (pull_request) Successful in 17s
2026-09-25 13:55:03 +00:00
panxiao81 7f218531ed 部署 Hydra 与通用 OIDC 登录适配器
yaml / yaml (push) Successful in 30s
ansible / collection-test (push) Successful in 1m38s
ansible / lint (push) Successful in 3m13s
2026-09-25 13:54:11 +00:00
panxiao81 8f3d1ed04a fix: 使用非 root runner 的 Ansible collections 目录
yaml / yaml (pull_request) Successful in 1m50s
ansible / collection-test (pull_request) Successful in 4m16s
hydra-login / verify (pull_request) Successful in 4m16s
ansible / lint (pull_request) Successful in 4m38s
2026-09-25 13:48:39 +00:00
panxiao81 5435d7d2d5 feat: 部署 Hydra 与通用 OIDC 登录适配器
ansible / lint (pull_request) Failing after 4m34s
hydra-login / verify (pull_request) Successful in 6m9s
ansible / collection-test (pull_request) Successful in 8m24s
2026-09-25 13:42:39 +00:00
panxiao81 e325a703dd Merge 部署内置 SPIRE CLI 的 runner 镜像
yaml / yaml (push) Successful in 16s
2026-09-24 07:29:53 +00:00
panxiao81 130849040c chore(runner): 部署内置 SPIRE CLI 镜像
yaml / yaml (pull_request) Successful in 22s
2026-09-24 07:29:36 +00:00
panxiao81 bcd034355e Merge 部署 runner 工作目录修复镜像
yaml / yaml (push) Failing after 8s
2026-09-24 06:57:48 +00:00
panxiao81 506e0cb983 chore(runner): 部署工作目录修复镜像
yaml / yaml (pull_request) Failing after 8s
2026-09-24 06:57:30 +00:00
panxiao81 c07709e078 Merge 为 OpenSandbox runner 挂载独立工作目录
yaml / yaml (push) Failing after 8s
2026-09-24 06:29:37 +00:00
panxiao81 e1979ebbb8 fix(runner): 挂载独立工作目录
yaml / yaml (pull_request) Failing after 7s
2026-09-24 06:29:02 +00:00
panxiao81 b02b2aec96 Merge 滚动部署最终 Docker runner 镜像
yaml / yaml (push) Successful in 18s
2026-09-24 06:12:24 +00:00
panxiao81 ff92037a28 chore(runner): 滚动部署最终 Docker 运行环境镜像
yaml / yaml (pull_request) Successful in 19s
2026-09-24 06:11:49 +00:00
panxiao81 9360eb2ea0 Merge 自动 Docker runner 镜像 rollout
yaml / yaml (push) Successful in 19s
2026-09-21 17:43:19 +00:00
panxiao81 0c7eb371e5 chore: rollout job-local Docker runner image
yaml / yaml (pull_request) Successful in 28s
2026-09-21 17:40:12 +00:00
panxiao81 757c21fc0b Merge pull request 'feat: 启用生产 VM runner 标签' (#136) from rollout/vm-runner-production into main
yaml / yaml (push) Successful in 12s
Reviewed-on: #136
2026-09-21 15:23:31 +00:00
panxiao81 c8e163f01e feat: 启用生产 VM runner 标签
yaml / yaml (pull_request) Successful in 21s
2026-09-21 15:22:36 +00:00
panxiao81 59fe28faa5 Merge Kata kind executor 最低资源配置
yaml / yaml (push) Successful in 20s
2026-09-21 12:36:20 +00:00
panxiao81 69c90b9fdf fix: 为 Kata kind executor 配置最低资源
yaml / yaml (pull_request) Successful in 19s
2026-09-21 12:33:35 +00:00
panxiao81 a6a861cc7e Merge 统一 Pod 与 VM executor 运行模型
yaml / yaml (push) Successful in 19s
2026-09-21 12:06:02 +00:00
panxiao81 5c9a09edde refactor: 统一 Pod 与 VM executor 运行模型
yaml / yaml (pull_request) Successful in 31s
2026-09-21 12:01:55 +00:00
panxiao81 cfb9c2d150 Merge Kata DinD guest ext4 存储修复
yaml / yaml (push) Successful in 16s
2026-09-21 11:50:26 +00:00
panxiao81 1ddec34149 fix: 为 Kata DinD 使用 guest ext4 数据盘
yaml / yaml (pull_request) Successful in 21s
2026-09-21 11:48:57 +00:00
panxiao81 db63503ec9 Merge Kata kind 资源规格
yaml / yaml (push) Successful in 16s
2026-09-21 11:34:00 +00:00
panxiao81 9953fae93d fix: 为 Kata kind 任务分配明确资源
yaml / yaml (pull_request) Successful in 25s
2026-09-21 11:31:57 +00:00
panxiao81 56598d6c84 Merge Kata DinD cgroup nesting 修复
yaml / yaml (push) Successful in 16s
2026-09-21 11:16:48 +00:00
panxiao81 834fdcb771 fix: 恢复 Kata DinD cgroup 初始化
yaml / yaml (pull_request) Successful in 17s
2026-09-21 11:14:52 +00:00
panxiao81 47ff73cefb Merge pull request '彻底删除旧静态 Gitea Runner' (#130) from retire/static-gitea-runner-phase2 into main
yaml / yaml (push) Successful in 19s
2026-09-21 10:23:05 +00:00
panxiao81 bb24c76704 chore: 删除旧静态 Gitea Runner
yaml / yaml (pull_request) Successful in 18s
2026-09-21 10:21:33 +00:00
panxiao81 b0791e4095 Merge pull request '退役旧静态 Gitea Runner(第一阶段)' (#129) from retire/static-gitea-runner-phase1 into main
yaml / yaml (push) Successful in 19s
2026-09-21 10:18:44 +00:00
panxiao81 26179f700f chore: 开始退役旧静态 Gitea Runner
yaml / yaml (pull_request) Successful in 16s
2026-09-21 10:17:06 +00:00
panxiao81 1ac99f4476 Merge pull request '部署 Runner 结构化生命周期日志' (#128) from deploy/structured-runner-logs into main
yaml / yaml (push) Successful in 23s
2026-09-21 09:46:47 +00:00
panxiao81 891c6b90e7 deploy: 更新 Runner 结构化日志镜像
yaml / yaml (pull_request) Successful in 21s
2026-09-21 09:45:33 +00:00
panxiao81 c2314256cc Merge pull request '为 sandbox executor 暴露内网 Runner facade' (#127) from fix/expose-runner-facade-internal into main
yaml / yaml (push) Successful in 26s
2026-09-21 09:18:16 +00:00
panxiao81 8244cdad79 fix: 为 sandbox 暴露 Runner facade
yaml / yaml (pull_request) Successful in 16s
2026-09-21 09:11:29 +00:00
panxiao81 23841556d3 Merge VM 集成测试入口隔离
yaml / yaml (push) Successful in 24s
2026-09-21 09:00:23 +00:00
panxiao81 4daaa38495 deploy: 隔离 VM 集成测试入口
yaml / yaml (pull_request) Successful in 12s
2026-09-21 08:58:21 +00:00
panxiao81 da0630cd29 Merge OpenSandbox Runner 标记修复部署
yaml / yaml (push) Successful in 46s
2026-09-21 08:35:24 +00:00
panxiao81 d0fcfdeddc deploy: 补回 OpenSandbox Runner 标记
yaml / yaml (pull_request) Successful in 23s
2026-09-21 08:34:37 +00:00
panxiao81 ecf395cc8f Merge OpenSandbox metadata 修复部署
yaml / yaml (push) Successful in 58s
2026-09-21 08:07:33 +00:00
panxiao81 9f646d51e7 deploy: 修复 OpenSandbox metadata 编码
yaml / yaml (pull_request) Successful in 56s
2026-09-21 08:04:15 +00:00
panxiao81 5c7e7ff8db Merge Runner 后端独立容量池部署
yaml / yaml (push) Successful in 34s
2026-09-21 07:35:32 +00:00
panxiao81 d6b97355d6 deploy: 启用 Runner 后端容量池
yaml / yaml (pull_request) Successful in 48s
2026-09-21 07:33:31 +00:00
panxiao81 b82ba4d5a0 Merge Pod Runner 四并发恢复
yaml / yaml (push) Successful in 22s
2026-09-21 06:57:37 +00:00
panxiao81 9dc7aabc94 恢复 Pod Runner 四并发
yaml / yaml (pull_request) Successful in 21s
2026-09-21 06:57:09 +00:00
panxiao81 b7c95fd0b8 Merge Runner 无中断滚动与容量隔离
yaml / yaml (push) Successful in 24s
2026-09-21 06:51:39 +00:00
panxiao81 098ff4e27f 更新 Runner 容量隔离镜像
yaml / yaml (pull_request) Successful in 27s
2026-09-21 06:51:37 +00:00
panxiao81 dfc55e1d9e 启用 Runner 无中断滚动更新
yaml / yaml (pull_request) Successful in 29s
2026-09-21 06:39:11 +00:00
panxiao81 18d9f83413 Merge Runner claim 恢复镜像
yaml / yaml (push) Successful in 34s
2026-09-21 06:24:31 +00:00
panxiao81 e5e2ab1b89 更新 Runner claim 恢复镜像
yaml / yaml (pull_request) Successful in 15s
2026-09-21 06:21:20 +00:00
panxiao81 0b236c43d9 Merge OpenSandbox VM Runner canary
yaml / yaml (push) Successful in 14s
2026-09-21 06:12:35 +00:00
panxiao81 762ce6458d 启用 OpenSandbox VM Runner canary
yaml / yaml (pull_request) Successful in 17s
2026-09-21 06:10:59 +00:00
72 changed files with 2670 additions and 378 deletions
+2 -4
View File
@@ -16,9 +16,6 @@ on:
- '.ansible-lint'
- '.gitea/workflows/ansible.yml'
env:
ANSIBLE_COLLECTIONS_PATH: /root/.ansible/collections
jobs:
lint:
runs-on: [self-hosted, pod]
@@ -30,6 +27,7 @@ jobs:
python3 -m pip install --user --break-system-packages \
--index-url https://pypi.org/simple --quiet uv==0.11.7
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
echo "ANSIBLE_COLLECTIONS_PATH=$HOME/.ansible/collections" >> "$GITHUB_ENV"
- name: Install ansible-lint and collections
run: |
@@ -50,7 +48,7 @@ jobs:
- name: ansible-lint
run: |
export PATH="$HOME/.local/bin:$PATH"
export ANSIBLE_COLLECTIONS_PATH="$PWD/infrastructure/samba-ad/ansible/collections:/root/.ansible/collections"
export ANSIBLE_COLLECTIONS_PATH="$PWD/infrastructure/samba-ad/ansible/collections:$HOME/.ansible/collections"
# 静态检查不应依赖生产 vault 凭据。一次性 checkout 可以去掉加密变量文件;
# syntax-check 只验证结构,不需要解析变量的运行时值。
rm -f \
+22
View File
@@ -0,0 +1,22 @@
name: hydra-login
on:
pull_request:
paths:
- 'apps/hydra/login-consent/**'
- '.gitea/workflows/hydra.yml'
workflow_dispatch:
jobs:
verify:
runs-on: [self-hosted, pod]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: apps/hydra/login-consent/go.mod
cache-dependency-path: apps/hydra/login-consent/go.sum
- name: Test authentication boundaries
working-directory: apps/hydra/login-consent
run: |
go test -race ./...
go vet ./...
CGO_ENABLED=0 go build -trimpath .
+20
View File
@@ -0,0 +1,20 @@
# Backstage
Backstage 作为 homelab 的只读开发者门户运行。应用源码、插件、测试和镜像构建归
`panxiao81/backstage` 仓库管理;本目录只保存 Kubernetes/Flux 部署声明,并通过
OCI digest 固定镜像。
## 外部前置
- OpenBao `kv/k8s/backstage` 必须包含以下与容器环境变量同名的字段:
`BACKEND_SECRET`、`AUTH_SESSION_SECRET`、`AUTH_OIDC_CLIENT_ID`、
`AUTH_OIDC_CLIENT_SECRET`、`POSTGRES_PASSWORD`、`GITEA_TOKEN`。
- PostgreSQL 需要在共享集群中预先创建由 `backstage` 角色拥有的 `backstage`
数据库;密码必须与 OpenBao 中的 `POSTGRES_PASSWORD` 一致。
- Authelia OIDC 客户端回调地址为
`https://backstage.ad.ddupan.top/api/auth/oidc/handler/frame`。
- AD DNS 需要将 `backstage.ad.ddupan.top` 指向 Envoy Gateway
`192.168.10.127`。
Flux 等待 ExternalSecret 和 Deployment 就绪;任何前置缺失都会使该
Kustomization 保持 NotReady,而不会回退到明文 Secret。
+85
View File
@@ -0,0 +1,85 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: backstage
namespace: backstage
labels:
app.kubernetes.io/name: backstage
backstage.io/kubernetes-id: homelab-backstage
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app.kubernetes.io/name: backstage
template:
metadata:
labels:
app.kubernetes.io/name: backstage
backstage.io/kubernetes-id: homelab-backstage
spec:
serviceAccountName: backstage
securityContext:
fsGroup: 1000
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
containers:
- name: backstage
image: zot.ad.ddupan.top/panxiao81/backstage@sha256:e5a12550726f19a680bc7c40e2cc07cc624318f9279ee121814d293a006ef210
imagePullPolicy: IfNotPresent
env:
- name: BACKSTAGE_BASE_URL
value: https://backstage.ad.ddupan.top
- name: POSTGRES_HOST
value: shared-postgresql-rw.shared-db.svc.cluster.local
- name: POSTGRES_PORT
value: "5432"
- name: POSTGRES_USER
value: backstage
- name: POSTGRES_DATABASE
value: backstage
- name: GITEA_HOST
value: git.ddupan.top
envFrom:
- secretRef:
name: backstage
ports:
- containerPort: 7007
name: http
protocol: TCP
readinessProbe:
httpGet:
path: /.backstage/health/v1/readiness
port: http
initialDelaySeconds: 10
periodSeconds: 10
timeoutSeconds: 3
livenessProbe:
httpGet:
path: /.backstage/health/v1/liveness
port: http
initialDelaySeconds: 30
periodSeconds: 20
timeoutSeconds: 3
resources:
requests:
cpu: 100m
memory: 512Mi
limits:
cpu: "1"
memory: 1Gi
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: [ALL]
readOnlyRootFilesystem: true
runAsNonRoot: true
runAsUser: 1000
volumeMounts:
- mountPath: /tmp
name: tmp
volumes:
- emptyDir: {}
name: tmp
+16
View File
@@ -0,0 +1,16 @@
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: backstage
namespace: backstage
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: openbao
target:
creationPolicy: Owner
name: backstage
dataFrom:
- extract:
key: k8s/backstage
+20
View File
@@ -0,0 +1,20 @@
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: backstage
namespace: backstage
spec:
parentRefs:
- name: eg
namespace: envoy-gateway-system
sectionName: https
hostnames:
- backstage.ad.ddupan.top
rules:
- matches:
- path:
type: PathPrefix
value: /
backendRefs:
- name: backstage
port: 7007
+11
View File
@@ -0,0 +1,11 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- serviceaccount.yaml
- rbac.yaml
- external-secret.yaml
- deployment.yaml
- service.yaml
- networkpolicy.yaml
- httproute.yaml
+7
View File
@@ -0,0 +1,7 @@
apiVersion: v1
kind: Namespace
metadata:
name: backstage
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/enforce-version: v1.36
+24
View File
@@ -0,0 +1,24 @@
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: backstage-ingress
namespace: backstage
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: backstage
policyTypes: [Ingress]
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: envoy-gateway-system
ports:
- port: 7007
protocol: TCP
- from:
- ipBlock:
cidr: 192.168.10.127/32
ports:
- port: 7007
protocol: TCP
+51
View File
@@ -0,0 +1,51 @@
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: backstage-read-only
rules:
- apiGroups: [""]
resources:
- configmaps
- limitranges
- pods
- pods/log
- resourcequotas
- services
verbs: [get, list, watch]
- apiGroups: [apps]
resources:
- daemonsets
- deployments
- replicasets
- statefulsets
verbs: [get, list, watch]
- apiGroups: [autoscaling]
resources:
- horizontalpodautoscalers
verbs: [get, list, watch]
- apiGroups: [batch]
resources:
- cronjobs
- jobs
verbs: [get, list, watch]
- apiGroups: [networking.k8s.io]
resources:
- ingresses
verbs: [get, list, watch]
- apiGroups: [metrics.k8s.io]
resources:
- pods
verbs: [get, list]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: backstage-read-only
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: backstage-read-only
subjects:
- kind: ServiceAccount
name: backstage
namespace: backstage
+15
View File
@@ -0,0 +1,15 @@
apiVersion: v1
kind: Service
metadata:
name: backstage
namespace: backstage
labels:
backstage.io/kubernetes-id: homelab-backstage
spec:
selector:
app.kubernetes.io/name: backstage
ports:
- name: http
port: 7007
protocol: TCP
targetPort: http
+6
View File
@@ -0,0 +1,6 @@
apiVersion: v1
kind: ServiceAccount
metadata:
name: backstage
namespace: backstage
automountServiceAccountToken: true
+12
View File
@@ -44,3 +44,15 @@ API、OIDC、Git/Flux 和 runner 均已验证。第二跳按明确决定跳过
结构,或者 release notes 指出相关 breaking migration 时,才恢复停机一致备份、分阶段
suspend、详细日志审计和扩展验收。出现启动失败或 migration error 时也立即升级为完整
故障流程。
## Hydra 人类登录 PoC
新增 `hydra` OIDC 登录源,旧 `authelia` 入口保留。Hydra 通过通用 OIDC Login/Consent
适配器转到现有 Authelia 完成人类认证;不是 Gitea 直接验证 LDAP 或 SPIFFE。
入口为 `https://git.ddupan.top/user/oauth2/hydra`,需 LAN/Tailscale 可达 Hydra 内网域名。
新 client secret 通过 `ExternalSecret/gitea-hydra-oidc` 从 OpenBao 投射。沿用
preferred_username、已验证邮箱与 groups;当前仍映射 gitea-admins,不在本轮切换组模型。
先部署并验证 Hydra discovery 后再接入本配置,避免 Gitea init 因上游不可达而失败。
实际登录验收与部署状态见 wiki;依赖和回退见 [Hydra README](../hydra/README.md)。
+7
View File
@@ -116,6 +116,13 @@ gitea:
scopes: openid profile email groups
groupClaimName: groups
adminGroup: gitea-admins
- name: hydra
provider: openidConnect
existingSecret: gitea-hydra-oidc
autoDiscoverUrl: https://hydra.ad.ddupan.top/.well-known/openid-configuration
scopes: openid profile email groups
groupClaimName: groups
adminGroup: gitea-admins
persistence:
size: 20Gi
+22
View File
@@ -0,0 +1,22 @@
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: gitea-hydra-oidc
namespace: gitea
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: openbao
target:
name: gitea-hydra-oidc
creationPolicy: Owner
template:
data:
key: gitea
secret: "{{ .client_secret }}"
data:
- secretKey: client_secret
remoteRef:
key: k8s/hydra
property: gitea_client_secret
+1
View File
@@ -16,3 +16,4 @@ resources:
- helmrepository.yaml
- helmrelease.yaml
- httproute.yaml
- hydra-external-secret.yaml
+102
View File
@@ -0,0 +1,102 @@
# Hydra 与 OIDC Login/Consent PoC
本目录提供独立 Hydra 签发服务,以及一个薄的 **OIDC 上游适配器**。当前上游配置为
Authelia;适配器不连接 LDAP,也不管理用户目录。Samba AD、密码和 MFA 继续由现有
Authelia 链路负责。第一轮只接入人类和 Gitea,不实现 agent 动态授权。
```text
Gitea → Hydra → OIDC Login/Consent → Authelia → Samba AD
← OIDC ← 经验证的上游身份 ← OIDC callback
```
目标入口:
- `https://hydra.ad.ddupan.top`:Hydra 公共 OAuth2/OIDC endpoint。
- `https://hydra-login.ad.ddupan.top`:上游 OIDC 登录及 consent 适配器。
- `hydra-admin.hydra.svc.cluster.local:4445`:仅集群内管理接口,无 HTTPRoute。
均为 LAN/Tailscale 入口,复用 Envoy `eg/https` wildcard TLS。没有增加公网 tunnel。
部署及实际验收状态以 wiki 和对应 PR 为准,文件存在不表示登录已验收。
## 首次使用与边界
在 Gitea 登录页选择 `hydra`,跳转到 Authelia 完成现有人类认证,再返回原有 Gitea
账号。旧 `authelia` 登录源保留。Gitea 的账号关联和资源权限仍由 Gitea 维护。
适配器要求验证上游 issuer、audience、签名、过期时间和 nonce,使用 PKCE S256,
并把单次 state 绑定到 Secure/HttpOnly/SameSite=Lax cookie。短期登录事务只存内存,
最多 1024 个、10 分钟过期;单副本重启后正在登录的用户需重试,不存人类密码或 token。
Hydra subject 为上游 `(issuer, sub)` 的 SHA-256 加 `human:` 前缀,与可变邮箱/用户名
分离。第一轮要求上游返回经过验证的 email 及 preferred_username;这些 claims 必须
明确配置进 ID token。更换 issuer 会改变本 PoC 的 subject,正式迁移前需要身份绑定设计。
仅为显式 `ALLOWED_CLIENTS=gitea` 自动 consent,scope 限于 openid/profile/email/groups;
拒绝额外 access-token audience,不发 refresh token。只按实际请求 scope 释放 claims。
这不是通用的无人确认授权服务。组当前透传,沿用 Gitea 的 gitea-admins 映射;统一组
模型和 agent 认证均在后续阶段。不存在对 Authelia 专有协议的调用。
NetworkPolicy 限制公共端口只接收 Envoy 流量,Hydra admin 只允许适配器访问。
Hydra 使用正式模式,TLS 由 Envoy 终止;不使用 `--dev`。管理操作使用受控
`kubectl port-forward`,不要将 admin 接口暴露到 Gateway。
## 依赖、秘密与初始化
依赖共享 CloudNativePG、OpenBao/ESO、Authelia OIDC、Envoy、Samba DNS、zot 镜像仓库。
Hydra 使用独立 `hydra` database/role,不与其他应用共享数据库角色。
`kv/k8s/hydra` 保存 dsn、system_secret、upstream_client_secret、upstream_client_digest、
gitea_client_secret;通过 ExternalSecret 投射,值不写入 Git。Bootstrap 创建角色及数据库
后才启动 Hydra migration。system_secret 必须持久保存,不得在重启时随机重建。
Authelia 中新增 confidential client `hydra-login`:
- redirect URI:`https://hydra-login.ad.ddupan.top/callback`;
- authorization policy:two_factor;grant:authorization_code;PKCE:S256;
- token endpoint auth:client_secret_basic;scope:openid/profile/email/groups;
- claims policy:把 preferred_username、name、email、email_verified、groups 放入 ID token;
- client secret 的 PBKDF2 digest 存入 Authelia,原值仅供适配器使用。
Authelia 尚非 Flux 管理。修改 Helm values 时保留所有已有 clients 与 secret 引用,
通过 `--reuse-values` 和最小 overlay 增加客户端,不能以本目录配置覆盖其完整 values。
Hydra 中注册 confidential client `gitea`,redirect URI 为
`https://git.ddupan.top/user/oauth2/hydra/callback`,grant/response 为 authorization_code/code,
scope 为 openid/profile/email/groups,token endpoint auth 为 client_secret_basic。
Gitea 启动时读取 OIDC discovery,所以应先确认 Hydra 健康和 discovery 可达,再接入 Gitea。
## 构建与检查
```bash
cd apps/hydra/login-consent
go test -race ./...
go vet ./...
CGO_ENABLED=0 go build -trimpath -ldflags='-s -w' -o login-consent .
docker build -t hydra-login-consent:VERSION .
```
Go module 独立,依赖由 go.sum 锁定;Dockerfile 固定基础镜像 digest。
使用已授权的短期 SPIFFE zot 凭据发布镜像,部署使用匿名拉取入口与不可变 digest。
不把 registry 凭据写入源码或 build args。
```bash
kubectl kustomize apps/hydra
sudo k3s kubectl -n hydra get deployment,pod,externalsecret,httproute
sudo k3s kubectl -n hydra logs deployment/hydra -c migrate
sudo k3s kubectl -n hydra logs deployment/hydra-login
```
日志不输出上游 token、授权 code、challenge 或秘密。登录失败先查两端 Pod 状态、
DNS/discovery 连通性、client redirect URI 和 scope,再由用户重新发起登录。
不要在故障排查中关闭签名验证、MFA 或 state/nonce 校验。
## 恢复与回退
保留共享 PostgreSQL 中 Hydra 数据及 OpenBao 秘密;数据库持有 clients、会话及签名密钥,
单独重建 Deployment 不能替代恢复数据库。先恢复依赖,再启动 Hydra 和适配器。
当前恢复仍依赖 homelab 共享基础设施,不能声称已完成独立灾备。
第一轮不切换 Authelia 的主入口。撤回 Gitea 的新增 Hydra 登录源即可回到旧入口;
先撤消费者,再考虑停用 Hydra。不要删除旧 Authelia 登录源、用户或数据库作为回退手段。
跨服务设计见 [独立 IAM 草案](https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/architecture/independent-iam-draft.md)。
+95
View File
@@ -0,0 +1,95 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: hydra
namespace: hydra
spec:
replicas: 1
selector:
matchLabels:
app: hydra
template:
metadata:
labels:
app: hydra
spec:
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
runAsUser: 1000
runAsGroup: 1000
seccompProfile:
type: RuntimeDefault
initContainers:
- name: migrate
image: docker.io/oryd/hydra:v26.2.0@sha256:ff67c7fb5f95074fa53374d41151713554960504b340cd3f95b09e65deaea2a9
args:
- migrate
- sql
- -e
- --yes
env:
- name: DSN
valueFrom:
secretKeyRef:
name: hydra
key: dsn
securityContext: &id002
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
resources: &id001
requests:
cpu: 50m
memory: 64Mi
limits:
memory: 256Mi
containers:
- name: hydra
image: docker.io/oryd/hydra:v26.2.0@sha256:ff67c7fb5f95074fa53374d41151713554960504b340cd3f95b09e65deaea2a9
args:
- serve
- all
- --config
- /etc/hydra/hydra.yaml
- --sqa-opt-out
env:
- name: DSN
valueFrom:
secretKeyRef:
name: hydra
key: dsn
- name: SECRETS_SYSTEM
valueFrom:
secretKeyRef:
name: hydra
key: system_secret
ports:
- name: public
containerPort: 4444
- name: admin
containerPort: 4445
resources: *id001
securityContext: *id002
volumeMounts:
- name: config
mountPath: /etc/hydra
readOnly: true
readinessProbe:
httpGet:
path: /health/ready
port: admin
initialDelaySeconds: 5
periodSeconds: 5
livenessProbe:
httpGet:
path: /health/alive
port: admin
initialDelaySeconds: 20
periodSeconds: 20
volumes:
- name: config
configMap:
name: hydra-config
+16
View File
@@ -0,0 +1,16 @@
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: hydra
namespace: hydra
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: openbao
target:
name: hydra
creationPolicy: Owner
dataFrom:
- extract:
key: k8s/hydra
+33
View File
@@ -0,0 +1,33 @@
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: hydra-public
namespace: hydra
spec:
parentRefs:
- name: eg
namespace: envoy-gateway-system
sectionName: https
hostnames:
- hydra.ad.ddupan.top
rules:
- backendRefs:
- name: hydra-public
port: 4444
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: hydra-login
namespace: hydra
spec:
parentRefs:
- name: eg
namespace: envoy-gateway-system
sectionName: https
hostnames:
- hydra-login.ad.ddupan.top
rules:
- backendRefs:
- name: hydra-login
port: 8080
+25
View File
@@ -0,0 +1,25 @@
serve:
public:
port: 4444
admin:
port: 4445
tls:
allow_termination_from:
- 10.42.0.0/16
cookies:
same_site_mode: Lax
urls:
self:
issuer: https://hydra.ad.ddupan.top
public: https://hydra.ad.ddupan.top
login: https://hydra-login.ad.ddupan.top/login
consent: https://hydra-login.ad.ddupan.top/consent
ttl:
access_token: 15m
id_token: 15m
auth_code: 5m
log:
level: info
leak_sensitive_values: false
oauth2:
expose_internal_errors: false
+15
View File
@@ -0,0 +1,15 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- external-secret.yaml
- deployment.yaml
- login-deployment.yaml
- services.yaml
- httproutes.yaml
- networkpolicy.yaml
configMapGenerator:
- name: hydra-config
namespace: hydra
files:
- hydra.yaml
+1
View File
@@ -0,0 +1 @@
/login-consent
+4
View File
@@ -0,0 +1,4 @@
FROM gcr.io/distroless/static-debian12:nonroot@sha256:afa5c872c891853ca7fcf1f12c3edb23f7eeef36189728842dd51042ff57f7ab
COPY login-consent /login-consent
USER 65532:65532
ENTRYPOINT ["/login-consent"]
+13
View File
@@ -0,0 +1,13 @@
module git.ddupan.top/panxiao81/homelab-infra/apps/hydra/login-consent
go 1.26.0
require (
github.com/coreos/go-oidc/v3 v3.14.1
golang.org/x/oauth2 v0.37.0
)
require (
github.com/go-jose/go-jose/v4 v4.0.5 // indirect
golang.org/x/crypto v0.36.0 // indirect
)
+18
View File
@@ -0,0 +1,18 @@
github.com/coreos/go-oidc/v3 v3.14.1 h1:9ePWwfdwC4QKRlCXsJGou56adA/owXczOzwKdOumLqk=
github.com/coreos/go-oidc/v3 v3.14.1/go.mod h1:HaZ3szPaZ0e4r6ebqvsLWlk2Tn+aejfmrfah6hnSYEU=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/go-jose/go-jose/v4 v4.0.5 h1:M6T8+mKZl/+fNNuFHvGIzDz7BTLQPIounk/b9dw3AaE=
github.com/go-jose/go-jose/v4 v4.0.5/go.mod h1:s3P1lRrkT8igV8D9OjyL4WRyHvjB6a4JSllnOrmmBOA=
github.com/google/go-cmp v0.6.0 h1:ofyhxvXcZhMsU5ulbFiLKl/XBFqE1GSq7atu8tAmTRI=
github.com/google/go-cmp v0.6.0/go.mod h1:17dUlkBOakJ0+DkrSSNjCkIjxS6bF9zb3elmeNGIjoY=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/stretchr/testify v1.10.0 h1:Xv5erBjTwe/5IxqUQTdXv5kgmIvbHo3QQyRwhJsOfJA=
github.com/stretchr/testify v1.10.0/go.mod h1:r2ic/lqez/lEtzL7wO/rwa5dbSLXVDPFyf8C91i36aY=
golang.org/x/crypto v0.36.0 h1:AnAEvhDddvBdpY+uR+MyHmuZzzNqXSe/GvuDeob5L34=
golang.org/x/crypto v0.36.0/go.mod h1:Y4J0ReaxCR1IMaabaSMugxJES1EpwhBHhv2bDHklZvc=
golang.org/x/oauth2 v0.37.0 h1:JUlcxA8oAtauLfiH8FX2/FkAWHAdi0QtGCGc+hofE98=
golang.org/x/oauth2 v0.37.0/go.mod h1:IxwZNxUULJmpBFf9K/9NTMSIfZZuvuTy1gGxhigP/58=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
+285
View File
@@ -0,0 +1,285 @@
// Login/Consent adapter for a single trusted upstream and first-party clients.
package main
import (
"bytes"
"context"
"crypto/rand"
"crypto/sha256"
"crypto/subtle"
"encoding/base64"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"io"
"log"
"net/http"
"net/url"
"os"
"strings"
"sync"
"time"
"github.com/coreos/go-oidc/v3/oidc"
"golang.org/x/oauth2"
)
const cookieName = "__Host-hydra-login"
type pending struct {
Challenge, Nonce, Verifier string
Expires time.Time
}
type claims struct {
Username string `json:"preferred_username"`
Email string `json:"email"`
EmailVerified bool `json:"email_verified"`
Name string `json:"name"`
Groups []string `json:"groups"`
}
type flowRequest struct {
Client struct {
ID string `json:"client_id"`
} `json:"client"`
Subject string `json:"subject"`
Scopes []string `json:"requested_scope"`
Audience []string `json:"requested_access_token_audience"`
Context claims `json:"context"`
}
type app struct {
admin, public string
client *http.Client
oauth oauth2.Config
verifier *oidc.IDTokenVerifier
allowed map[string]bool
mu sync.Mutex
pending map[string]pending
}
func required(key string) string {
v := os.Getenv(key)
if v == "" {
log.Fatalf("missing %s", key)
}
return v
}
func random() string {
b := make([]byte, 32)
if _, err := rand.Read(b); err != nil {
panic(err)
}
return base64.RawURLEncoding.EncodeToString(b)
}
func (a *app) api(ctx context.Context, method, path string, in, out any) error {
var body io.Reader
if in != nil {
b, err := json.Marshal(in)
if err != nil {
return err
}
body = bytes.NewReader(b)
}
req, err := http.NewRequestWithContext(ctx, method, a.admin+path, body)
if err != nil {
return err
}
req.Header.Set("Content-Type", "application/json")
resp, err := a.client.Do(req)
if err != nil {
return errors.New("Hydra unavailable")
}
defer resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("Hydra status %d", resp.StatusCode)
}
if out != nil {
return json.NewDecoder(io.LimitReader(resp.Body, 1<<20)).Decode(out)
}
return nil
}
func (a *app) request(r *http.Request, kind, challenge string) (flowRequest, error) {
var f flowRequest
if challenge == "" || len(challenge) > 8192 {
return f, errors.New("missing or invalid challenge")
}
err := a.api(r.Context(), http.MethodGet, "/admin/oauth2/auth/requests/"+kind+"?"+kind+"_challenge="+url.QueryEscape(challenge), nil, &f)
if err != nil {
return f, err
}
if !a.allowed[f.Client.ID] {
return f, errors.New("client not allowed")
}
return f, nil
}
func (a *app) accept(w http.ResponseWriter, r *http.Request, kind, challenge string, body any) {
var result struct {
Redirect string `json:"redirect_to"`
}
if err := a.api(r.Context(), http.MethodPut, "/admin/oauth2/auth/requests/"+kind+"/accept?"+kind+"_challenge="+url.QueryEscape(challenge), body, &result); err != nil {
fail(w, 502)
return
}
// Only Hydra's own authorization endpoint can receive a challenge verifier.
u, err := url.Parse(result.Redirect)
p, _ := url.Parse(a.public)
if err != nil || u.Scheme != p.Scheme || u.Host != p.Host || u.User != nil || u.Path != "/oauth2/auth" {
fail(w, 502)
return
}
http.Redirect(w, r, result.Redirect, http.StatusSeeOther)
}
func fail(w http.ResponseWriter, status int) { http.Error(w, http.StatusText(status), status) }
func (a *app) login(w http.ResponseWriter, r *http.Request) {
challenge := r.URL.Query().Get("login_challenge")
if _, err := a.request(r, "login", challenge); err != nil {
fail(w, 403)
return
}
state := random()
p := pending{challenge, random(), oauth2.GenerateVerifier(), time.Now().Add(10 * time.Minute)}
a.mu.Lock()
for k, v := range a.pending {
if time.Now().After(v.Expires) {
delete(a.pending, k)
}
}
if len(a.pending) >= 1024 {
a.mu.Unlock()
fail(w, 503)
return
}
a.pending[state] = p
a.mu.Unlock()
http.SetCookie(w, &http.Cookie{Name: cookieName, Value: state, Path: "/", Secure: true, HttpOnly: true, SameSite: http.SameSiteLaxMode, MaxAge: 600})
http.Redirect(w, r, a.oauth.AuthCodeURL(state, oidc.Nonce(p.Nonce), oauth2.S256ChallengeOption(p.Verifier)), http.StatusSeeOther)
}
func (a *app) take(r *http.Request) (pending, error) {
state := r.URL.Query().Get("state")
cookie, err := r.Cookie(cookieName)
if err != nil || state == "" || subtle.ConstantTimeCompare([]byte(cookie.Value), []byte(state)) != 1 {
return pending{}, errors.New("state mismatch")
}
a.mu.Lock()
defer a.mu.Unlock()
p, ok := a.pending[state]
delete(a.pending, state)
if !ok || time.Now().After(p.Expires) {
return pending{}, errors.New("expired or used state")
}
return p, nil
}
func (a *app) callback(w http.ResponseWriter, r *http.Request) {
p, err := a.take(r)
if err != nil {
fail(w, 403)
return
}
http.SetCookie(w, &http.Cookie{Name: cookieName, Path: "/", Secure: true, HttpOnly: true, SameSite: http.SameSiteLaxMode, MaxAge: -1})
if r.URL.Query().Get("error") != "" || r.URL.Query().Get("code") == "" {
fail(w, 403)
return
}
ctx := oidc.ClientContext(r.Context(), a.client)
token, err := a.oauth.Exchange(ctx, r.URL.Query().Get("code"), oauth2.VerifierOption(p.Verifier))
if err != nil {
fail(w, 502)
return
}
raw, ok := token.Extra("id_token").(string)
if !ok {
fail(w, 502)
return
}
id, err := a.verifier.Verify(ctx, raw)
if err != nil || id.Nonce != p.Nonce || id.Subject == "" {
fail(w, 403)
return
}
var c claims
if id.Claims(&c) != nil || c.Username == "" || c.Email == "" || !c.EmailVerified {
fail(w, 403)
return
}
if _, err := a.request(r, "login", p.Challenge); err != nil {
fail(w, 403)
return
}
// Stable identity is tied to the verified upstream issuer+subject, never email.
sum := sha256.Sum256([]byte(id.Issuer + "\x00" + id.Subject))
a.accept(w, r, "login", p.Challenge, map[string]any{"subject": "human:" + hex.EncodeToString(sum[:]), "remember": false, "context": c})
}
func consentSession(f flowRequest) (map[string]any, error) {
if !strings.HasPrefix(f.Subject, "human:") || f.Context.Username == "" || f.Context.Email == "" || !f.Context.EmailVerified {
return nil, errors.New("invalid identity context")
}
allowed := map[string]bool{"openid": true, "profile": true, "email": true, "groups": true}
session := map[string]any{"principal_type": "human"}
for _, scope := range f.Scopes {
if !allowed[scope] {
return nil, errors.New("scope not allowed")
}
switch scope {
case "profile":
session["preferred_username"] = f.Context.Username
session["name"] = f.Context.Name
case "email":
session["email"] = f.Context.Email
session["email_verified"] = true
case "groups":
session["groups"] = f.Context.Groups
}
}
if len(f.Audience) > 0 {
return nil, errors.New("access token audience not allowed")
}
return session, nil
}
func (a *app) consent(w http.ResponseWriter, r *http.Request) {
challenge := r.URL.Query().Get("consent_challenge")
f, err := a.request(r, "consent", challenge)
if err != nil {
fail(w, 403)
return
}
session, err := consentSession(f)
if err != nil {
fail(w, 403)
return
}
// Explicit policy for pre-approved first-party clients only; no generic auto-consent.
a.accept(w, r, "consent", challenge, map[string]any{"grant_scope": f.Scopes, "remember": false, "session": map[string]any{"id_token": session}})
}
func (a *app) handler() http.Handler {
mux := http.NewServeMux()
mux.HandleFunc("GET /healthz", func(w http.ResponseWriter, r *http.Request) { w.WriteHeader(200) })
mux.HandleFunc("GET /login", a.login)
mux.HandleFunc("GET /callback", a.callback)
mux.HandleFunc("GET /consent", a.consent)
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Cache-Control", "no-store")
w.Header().Set("Referrer-Policy", "no-referrer")
w.Header().Set("X-Content-Type-Options", "nosniff")
w.Header().Set("Content-Security-Policy", "default-src 'none'; frame-ancestors 'none'")
mux.ServeHTTP(w, r)
})
}
func main() {
client := &http.Client{Timeout: 15 * time.Second, CheckRedirect: func(r *http.Request, via []*http.Request) error { return http.ErrUseLastResponse }}
issuer := required("UPSTREAM_ISSUER")
ctx := oidc.ClientContext(context.Background(), client)
provider, err := oidc.NewProvider(ctx, issuer)
if err != nil {
log.Fatal("upstream discovery failed")
}
clientID := required("UPSTREAM_CLIENT_ID")
a := &app{admin: required("HYDRA_ADMIN_URL"), public: required("HYDRA_PUBLIC_URL"), client: client, allowed: map[string]bool{}, pending: map[string]pending{},
oauth: oauth2.Config{ClientID: clientID, ClientSecret: required("UPSTREAM_CLIENT_SECRET"), RedirectURL: required("CALLBACK_URL"), Endpoint: provider.Endpoint(), Scopes: []string{"openid", "profile", "email", "groups"}},
verifier: provider.Verifier(&oidc.Config{ClientID: clientID})}
for _, id := range strings.Split(required("ALLOWED_CLIENTS"), ",") {
a.allowed[id] = true
}
s := http.Server{Addr: ":8080", Handler: a.handler(), ReadHeaderTimeout: 5 * time.Second, ReadTimeout: 20 * time.Second, WriteTimeout: 45 * time.Second, IdleTimeout: 60 * time.Second, MaxHeaderBytes: 16384}
log.Print("login/consent adapter listening on :8080")
log.Fatal(s.ListenAndServe())
}
+107
View File
@@ -0,0 +1,107 @@
package main
import (
"encoding/json"
"net/http"
"net/http/httptest"
"net/url"
"strings"
"testing"
"time"
"golang.org/x/oauth2"
)
func TestStateBoundToCookieSingleUseAndExpiry(t *testing.T) {
a := &app{pending: map[string]pending{"valid": {Challenge: "challenge", Expires: time.Now().Add(time.Minute)}, "expired": {Expires: time.Now().Add(-time.Minute)}}}
request := func(state, cookie string) *http.Request {
r := httptest.NewRequest("GET", "https://login.example/callback?state="+state, nil)
if cookie != "" {
r.AddCookie(&http.Cookie{Name: cookieName, Value: cookie})
}
return r
}
for _, r := range []*http.Request{request("valid", ""), request("valid", "other"), request("expired", "expired")} {
if _, err := a.take(r); err == nil {
t.Fatal("invalid state accepted")
}
}
if p, err := a.take(request("valid", "valid")); err != nil || p.Challenge != "challenge" {
t.Fatal("valid state rejected")
}
if _, err := a.take(request("valid", "valid")); err == nil {
t.Fatal("replayed state accepted")
}
}
func TestConsentRejectsPrivilegeExpansionAndFiltersClaims(t *testing.T) {
f := flowRequest{Subject: "human:known", Scopes: []string{"openid", "email"}, Context: claims{Username: "alice", Email: "[email protected]", EmailVerified: true, Groups: []string{"operators"}}}
s, err := consentSession(f)
if err != nil {
t.Fatal(err)
}
if _, ok := s["groups"]; ok {
t.Fatal("groups leaked without scope")
}
if _, ok := s["preferred_username"]; ok {
t.Fatal("profile leaked without scope")
}
for _, scope := range []string{"admin", "offline_access", "unknown"} {
bad := f
bad.Scopes = append([]string{"openid"}, scope)
if _, err := consentSession(bad); err == nil {
t.Fatalf("accepted %s", scope)
}
}
f.Audience = []string{"other-service"}
if _, err := consentSession(f); err == nil {
t.Fatal("unexpected audience accepted")
}
f.Audience = nil
f.Context.EmailVerified = false
if _, err := consentSession(f); err == nil {
t.Fatal("unverified email accepted")
}
}
func TestLoginValidatesClientAndUsesPKCEAndNonce(t *testing.T) {
admin := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
json.NewEncoder(w).Encode(map[string]any{"client": map[string]string{"client_id": r.URL.Query().Get("login_challenge")}})
}))
defer admin.Close()
a := &app{admin: admin.URL, client: admin.Client(), allowed: map[string]bool{"gitea": true}, pending: map[string]pending{}, oauth: oauth2.Config{ClientID: "hydra-login", RedirectURL: "https://login.example/callback", Endpoint: oauth2.Endpoint{AuthURL: "https://upstream.example/authorize"}}}
w := httptest.NewRecorder()
a.handler().ServeHTTP(w, httptest.NewRequest("GET", "https://login.example/login?login_challenge=rogue", nil))
if w.Code != 403 {
t.Fatal("unknown client accepted")
}
w = httptest.NewRecorder()
a.handler().ServeHTTP(w, httptest.NewRequest("GET", "https://login.example/login?login_challenge=gitea", nil))
if w.Code != 303 {
t.Fatalf("status %d", w.Code)
}
u, _ := url.Parse(w.Header().Get("Location"))
q := u.Query()
if q.Get("code_challenge_method") != "S256" || q.Get("code_challenge") == "" || q.Get("nonce") == "" || q.Get("state") == "" {
t.Fatal("missing protocol binding")
}
cookies := w.Result().Cookies()
if len(cookies) != 1 || !cookies[0].Secure || !cookies[0].HttpOnly || cookies[0].SameSite != http.SameSiteLaxMode || cookies[0].Value != q.Get("state") {
t.Fatal("unsafe cookie")
}
if w.Header().Get("Cache-Control") != "no-store" {
t.Fatal("missing cache protection")
}
}
func TestHydraRedirectCannotLeaveTrustedOrigin(t *testing.T) {
for _, target := range []string{"https://evil.example/oauth2/auth", "https://[email protected]/oauth2/auth", "https://hydra.example/other"} {
admin := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
json.NewEncoder(w).Encode(map[string]string{"redirect_to": target})
}))
a := &app{admin: admin.URL, public: "https://hydra.example", client: admin.Client()}
w := httptest.NewRecorder()
a.accept(w, httptest.NewRequest("GET", "https://login.example/login", nil), "login", "challenge", map[string]string{"subject": "human:test"})
if w.Code != 502 || strings.Contains(w.Header().Get("Location"), "evil") {
t.Fatal("untrusted redirect accepted")
}
admin.Close()
}
}
+69
View File
@@ -0,0 +1,69 @@
apiVersion: apps/v1
kind: Deployment
metadata:
name: hydra-login
namespace: hydra
spec:
replicas: 1
strategy:
type: Recreate
selector:
matchLabels:
app: hydra-login
template:
metadata:
labels:
app: hydra-login
spec:
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
runAsUser: 65532
runAsGroup: 65532
seccompProfile:
type: RuntimeDefault
containers:
- name: login-consent
image: zot.ad.ddupan.top/iam/oidc-login-consent@sha256:fede9b9e93c457c4b7a8a6022d9df86ff5d5900d3f6b6c4f3439851a0ae1e944
env:
- name: HYDRA_ADMIN_URL
value: http://hydra-admin.hydra.svc.cluster.local:4445
- name: HYDRA_PUBLIC_URL
value: https://hydra.ad.ddupan.top
- name: UPSTREAM_ISSUER
value: https://auth.ddupan.top
- name: UPSTREAM_CLIENT_ID
value: hydra-login
- name: CALLBACK_URL
value: https://hydra-login.ad.ddupan.top/callback
- name: ALLOWED_CLIENTS
value: gitea
- name: UPSTREAM_CLIENT_SECRET
valueFrom:
secretKeyRef:
name: hydra
key: upstream_client_secret
ports:
- name: http
containerPort: 8080
resources:
requests:
cpu: 50m
memory: 64Mi
limits:
memory: 256Mi
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop:
- ALL
readinessProbe:
httpGet:
path: /healthz
port: http
livenessProbe:
httpGet:
path: /healthz
port: http
initialDelaySeconds: 10
+6
View File
@@ -0,0 +1,6 @@
apiVersion: v1
kind: Namespace
metadata:
name: hydra
labels:
pod-security.kubernetes.io/enforce: restricted
+46
View File
@@ -0,0 +1,46 @@
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: hydra
namespace: hydra
spec:
podSelector:
matchLabels:
app: hydra
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: envoy-gateway-system
ports:
- port: 4444
protocol: TCP
- from:
- podSelector:
matchLabels:
app: hydra-login
ports:
- port: 4445
protocol: TCP
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: hydra-login
namespace: hydra
spec:
podSelector:
matchLabels:
app: hydra-login
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: envoy-gateway-system
ports:
- port: 8080
protocol: TCP
+35
View File
@@ -0,0 +1,35 @@
apiVersion: v1
kind: Service
metadata:
name: hydra-public
namespace: hydra
spec:
selector:
app: hydra
ports:
- port: 4444
targetPort: 4444
---
apiVersion: v1
kind: Service
metadata:
name: hydra-admin
namespace: hydra
spec:
selector:
app: hydra
ports:
- port: 4445
targetPort: 4445
---
apiVersion: v1
kind: Service
metadata:
name: hydra-login
namespace: hydra
spec:
selector:
app: hydra-login
ports:
- port: 8080
targetPort: 8080
+29 -5
View File
@@ -70,11 +70,11 @@ configFiles:
},
"accessControl": {
"repositories": {
"panxiao81/gitea-dynamic-runner-controller": {
"panxiao81/backstage": {
"policies": [
{
"users": [
"spiffe://ddupan.top/ci/panxiao81/gitea-dynamic-runner/publish-images",
"spiffe://ddupan.top/ci/panxiao81/backstage/image",
"spiffe://ddupan.top/dev/panxiao81"
],
"actions": [
@@ -88,18 +88,42 @@ configFiles:
"read"
]
},
"panxiao81/gitea-dynamic-runner-runner": {
"panxiao81/gitea-dynamic-runner-controller": {
"policies": [
{
"users": [
"spiffe://ddupan.top/ci/panxiao81/gitea-dynamic-runner/publish-images",
"spiffe://ddupan.top/dev/panxiao81"
"spiffe://ddupan.top/ci/panxiao81/gitea-dynamic-runner/publish-images"
],
"actions": [
"read",
"create",
"update"
]
},
{
"users": ["spiffe://ddupan.top/dev/panxiao81"],
"actions": ["read", "create", "update", "delete"]
}
],
"defaultPolicy": [
"read"
]
},
"panxiao81/gitea-dynamic-runner-runner": {
"policies": [
{
"users": [
"spiffe://ddupan.top/ci/panxiao81/gitea-dynamic-runner/publish-images"
],
"actions": [
"read",
"create",
"update"
]
},
{
"users": ["spiffe://ddupan.top/dev/panxiao81"],
"actions": ["read", "create", "update", "delete"]
}
],
"defaultPolicy": [
+22
View File
@@ -0,0 +1,22 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: backstage
namespace: flux-system
spec:
dependsOn:
- name: envoy-gateway
- name: external-secrets
healthChecks:
- apiVersion: apps/v1
kind: Deployment
name: backstage
namespace: backstage
interval: 10m
path: ./apps/backstage
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 5m
wait: true
@@ -1,14 +1,17 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: gitea-actions
name: hydra
namespace: flux-system
spec:
dependsOn:
- name: envoy-gateway
- name: external-secrets
interval: 10m
path: ./platform/gitea-runner
path: ./apps/hydra
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
wait: false
timeout: 5m
wait: true
+2 -1
View File
@@ -7,7 +7,7 @@ resources:
- apps/envoy-gateway.yaml
- apps/external-secrets.yaml
- apps/gitea.yaml
- apps/gitea-actions.yaml
- apps/backstage.yaml
- apps/http-echo.yaml
- apps/openebs.yaml
- apps/nats.yaml
@@ -16,3 +16,4 @@ resources:
- apps/observability.yaml
- apps/zot.yaml
- apps/nexus.yaml
- apps/hydra.yaml
+2
View File
@@ -20,6 +20,8 @@ homelab_dns:
- { zone: ad.ddupan.top, name: nats, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: nexus, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: s3, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: hydra, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: hydra-login, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-oidc, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-server, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: zot, type: A, values: [192.168.10.127] }
+32 -21
View File
@@ -1,11 +1,10 @@
# Gitea dynamic runner
本目录部署单副本 Go controller,默认在同一进程运行 RunnerService scheduler、原生
Kubernetes Pod worker 和 SPIFFE mTLS facade。首轮生产 canary 只启用
`scheduler,pod-worker`,OpenSandbox VM worker 保持关闭:
本目录部署单副本 Go controller,在同一进程运行 RunnerService scheduler、原生
Kubernetes worker、OpenSandbox worker 和 SPIFFE mTLS facade:
```text
Gitea RunnerService -> scheduler -> JetStream ci.runner.pod
Gitea RunnerService -> scheduler -> JetStream ci.runner.container.kubernetes
|
v
homelab Kubernetes Pod
@@ -16,28 +15,34 @@ Gitea RunnerService -> scheduler -> JetStream ci.runner.pod
Gitea
```
Pod backend 不经过 OpenSandbox。VM backend 后续启用时才访问 VyOS 暴露的
`container+kubernetes` placement 不经过 OpenSandbox;`vm+opensandbox` placement
访问 VyOS 暴露的
OpenSandbox Lifecycle API;本目录不修改 sandbox 平台侧 ESO、Bao Terraform 或
OpenSandbox chart 所有权边界。
## Canary 安全边界
- Deployment 为单副本,滚动策略固定 `maxSurge: 0`、`maxUnavailable: 1`,避免两个
scheduler 共享同一个 Gitea runner 身份。
- controller 的 single-flight gate 只允许一个已领取 task 运行;Gitea 接受终态
`UpdateTask` 后才领取下一条。`POD_CAPACITY=1` 同时限制 worker 创建并发。
- Deployment 为单副本,滚动策略固定 `maxSurge: 1`、`maxUnavailable: 0`,保证新
facade Ready 后才终止旧实例。scheduler 必须通过 Kubernetes Lease 保持单 leader,
不能依赖 Recreate 避免重复领取。
- scheduler 使用一个 runner registration,并按总容量启动并发 `FetchTask` goroutine;
`POD_CAPACITY=4` 与 `VM_CAPACITY=1` 分别限制两个 durable consumer 和 backend
admission pool。池满时 assignment 保持 JetStream pending,任一 backend 不占用
另一方的执行槽位,也不会创建超出容量的 workload。
- rollout 重叠期间只有持有 `Lease/dynamic-runner-scheduler` 的 controller 执行
`FetchTask`;所有 Ready 实例都可通过 backend metadata 恢复 claim 并服务 facade。
- executor 镜像使用 digest;Pod 以 UID 2000 运行,SPIRE `ClusterStaticEntry` 同时绑定
具体 Pod UID 与 `unix:uid:2000`。
- facade 只有 ClusterIP,executor 通过
`dynamic-runner-controller.dynamic-runner.svc:8443` 访问;双方使用 Workload API
X509-SVID mTLS,不再保留公网 webhook/token endpoint。
- facade 通过仅内网可路由的 `192.168.10.127:30443` NodePort 提供给 sandbox executor;
双方使用 Workload API X509-SVID mTLS,并按 SPIFFE ID 而不是 IP/DNS 名验证服务端。
该入口不经过公网或 Cloudflare Tunnel。
- controller 的 SPIFFE ID 固定为
`spiffe://ddupan.top/ns/dynamic-runner/sa/dynamic-runner-controller`。
- NATS 保留既有最小权限分离:`ci-producer` 仅 publish,`ci-worker` 仅 pull/ACK。
当前 facade claim registry 与 single-flight gate 是进程内状态。异常重启时应先检查
现存 assignment Pod 和 Gitea task,再人工决定是否恢复 scheduler;不能通过扩大副本数
规避。完成 backend-driven recovery 前保持 `replicas: 1`。
facade claim registry 在启动时从 Pod/OpenSandbox metadata 恢复;scheduler leadership
由 Kubernetes Lease 持久化协调。Deployment 仍保持 `replicas: 1`,滚动更新期间允许
一个额外 Pod 提供 facade 连续性。
## Secret 边界
@@ -78,17 +83,23 @@ kubectl -n dynamic-runner rollout status deploy/dynamic-runner-controller --time
kubectl -n dynamic-runner logs deploy/dynamic-runner-controller -f
```
确认 scheduler 只领取一条 `[self-hosted,pod]` task,然后验证:
确认 scheduler 只领取一条 `[self-hosted,container,kubernetes]` task,然后验证:
1. assignment message 进入并离开 durable `pod` consumer;
1. assignment message 进入并离开 durable `container.kubernetes` consumer;
2. `gitea-task-<task-id>` Pod 创建,取得实际 Pod UID;
3. 同名 `ClusterStaticEntry` 的 parent ID 包含该 UID,SPIFFE ID 使用 repository/job key;
4. 官方 runner v3.5.0 经 facade claim 精确 task,Gitea 实时收到日志和终态;
5. terminal gate 释放后才领取下一条任务。
5. Pod consumer 达到 capacity 时 VM consumer 仍可独立接受任务。
当前代码尚未完成 ACK 后 backend lifecycle reconciler,因此 canary 成功后可能留下
Completed Pod/ClusterStaticEntry。首次测试要人工核对并删除已终态资源;在完整自动清理
通过前不得提高并发或启用 VM worker。
Pod 与 VM 都在 Gitea 接受终态后先把 terminal marker 写入各自 backend metadata,再由
lifecycle reconciler 删除执行器。首次 VM 测试仍须观察 BatchSandbox、Pod 与
ClusterStaticEntry 全部消失;完整自动清理通过前不得提高 `POD_CAPACITY` 或
`VM_CAPACITY`。
VM backend 已通过 `vm-dev` canary 完成 Docker、kind、SPIFFE 和完整生命周期测试。
生产规范标签为 `[self-hosted, vm, opensandbox]`,旧 `[self-hosted, vm]` 仍严格映射到
同一 placement。初始保持 `VM_CAPACITY=1`;扩容前先观察实际任务的
资源水位、等待时间以及 OpenSandbox 是否存在 terminal sandbox 残留。
## 回滚
+18 -6
View File
@@ -8,8 +8,8 @@ spec:
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 0
maxUnavailable: 1
maxSurge: 1
maxUnavailable: 0
selector:
matchLabels:
app.kubernetes.io/name: dynamic-runner-controller
@@ -38,12 +38,12 @@ spec:
mountPath: /trust
containers:
- name: controller
image: zot.ad.ddupan.top/panxiao81/gitea-dynamic-runner-controller@sha256:91ef82f49e0117d8973a3c22fbfcc5eab8ac74dc6eb22d70d871ac2998fd0795
image: zot.ad.ddupan.top/panxiao81/gitea-dynamic-runner-controller@sha256:15735667e8990974a30d9a5e6a1030d63f791ed6d80fa2b0dc9e39959ffcf334
imagePullPolicy: IfNotPresent
args: [controller]
env:
- name: COMPONENTS
value: scheduler,pod-worker
value: scheduler,kubernetes-worker,opensandbox-worker
- name: GITEA_INSTANCE_URL
value: https://git.ddupan.top
- name: GITEA_RUNNER_UUID_FILE
@@ -63,7 +63,7 @@ spec:
- name: RUNNER_FACADE_LISTEN
value: :8443
- name: RUNNER_FACADE_URL
value: https://dynamic-runner-controller.dynamic-runner.svc:8443
value: https://192.168.10.127:30443
- name: RUNNER_FACADE_SPIFFE_ID
value: spiffe://ddupan.top/ns/dynamic-runner/sa/dynamic-runner-controller
- name: SPIFFE_ENDPOINT_SOCKET
@@ -73,13 +73,25 @@ spec:
fieldRef:
fieldPath: metadata.namespace
- name: POD_EXECUTOR_IMAGE
value: zot.ad.ddupan.top/panxiao81/gitea-dynamic-runner-runner@sha256:08f83c9993645dce376140fdde7e29e809795166b2bb32e50cebdbefaf3cf303
value: zot.ad.ddupan.top/panxiao81/gitea-dynamic-runner-runner@sha256:83e96c3959af663a75a7229cb04cbbb97a642d4ef3704ac5cdceb1d502a72ed3
- name: POD_SERVICE_ACCOUNT
value: gitea-dynamic-runner
- name: POD_EXECUTOR_UID
value: "2000"
- name: POD_CAPACITY
value: "4"
- name: OPENSANDBOX_API
value: http://10.60.0.13:8080
- name: OPENSANDBOX_API_KEY_FILE
value: /run/dynamic-runner-secrets/opensandbox-api-key
- name: OPENSANDBOX_POOL
value: ci-vm
- name: VM_CAPACITY
value: "1"
- name: VM_RUNNER_LABEL
value: vm
- name: VM_TIMEOUT_SECONDS
value: "14400"
- name: SPIRE_CLUSTER
value: homelab
- name: SPIRE_CLASS
+3
View File
@@ -19,6 +19,9 @@ rules:
- apiGroups: [""]
resources: [pods]
verbs: [create, get, list, watch, patch, delete]
- apiGroups: [coordination.k8s.io]
resources: [leases]
verbs: [create, get, list, watch, update, patch]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
+2 -1
View File
@@ -4,10 +4,11 @@ metadata:
name: dynamic-runner-controller
namespace: dynamic-runner
spec:
type: ClusterIP
type: NodePort
selector:
app.kubernetes.io/name: dynamic-runner-controller
ports:
- name: facade
port: 8443
targetPort: facade
nodePort: 30443
-114
View File
@@ -1,114 +0,0 @@
# Gitea Actions runner
This is the bootstrap runner for Gitea Actions. One persistent runner Pod accepts
up to four jobs; each job runs in a dynamically created container inside a
Docker-in-Docker daemon. The official chart runs DinD privileged. Rootless DinD
would still be privileged and is blocked by the node's AppArmor user-namespace
policy, so this deployment uses regular DinD instead of weakening that host-wide
policy. Only trusted workflows may target this runner.
DinD 同时使用 `--mtu=1450` 和
`--default-network-opt=bridge=com.docker.network.driver.mtu=1450`,与 k3s Pod 的
`eth0` 一致。前者只覆盖 Docker 默认 bridge;act 为每个 job 创建 user-defined
bridge,必须由后者设置默认 MTU。不要在未验证节点 Pod MTU 的情况下删除或修改这
两个参数:MTU 1500 的 job 容器虽然能够解析 GitHub、甚至建立 TCP 连接,但较大的
TLS 数据包会在嵌套网络路径中丢失,表现为 `github.com` / `api.github.com` 超时或
`setup-go` 每次请求卡满 6 分钟后重试。Pod 网络和默认 Docker bridge 正常不代表
Actions job bridge 正常。
The runner is registered at instance scope so it is available to every repository
on this Gitea instance. Repository permissions and protected-branch review are
therefore the security boundary; do not enable Actions for untrusted repositories.
The runner registration token is authoritative in OpenBao at
`kv/k8s/gitea-runner`. External Secrets Operator projects its `token` property to
the `gitea-runner-token` Secret. Never put the token in this directory or a Helm
command line.
## SPIRE 与 OCI 发布
runner Pod 使用专用 ServiceAccount `gitea-actions`,并由精确匹配 namespace、
ServiceAccount 隐含的 Pod、以及 chart labels 的 `ClusterSPIFFEID` 获得:
```text
spiffe://ddupan.top/ci/gitea-actions
```
SPIFFE CSI socket 同时只读挂载到 runner 和 DinD。act 的 volume allowlist 只允许
`/run/spire/agent-sockets`;workflow 仍必须在 job container 中显式请求该 bind
mount。原因是 bind mount 由 DinD 内的 dockerd 解析,只挂 runner 容器无法让 job
访问 Workload API。
该身份不是通用 registry 管理员。zot 只对明确列出的 CI 镜像仓库授予
`read/create/update`,不授予 delete 或其他仓库写入。workflow 应获取
`aud=zot` 的短期 JWT-SVID,并经 stdin 传给 registry client,不得把 JWT、X.509
SVID 或 Docker auth 写入 workspace/artifact。
## Flux 接管状态
该 release 最初通过下述 review-first 流程手动 bootstrap。下一个 GitOps 阶段将
使用 Flux `HelmRelease` 接管它,并首先固定现有 chart `0.1.1`,不在接管 PR 中升级。
迁移前审计发现:Helm 保存的 user-supplied values 和 release manifest 仍描述失败的
rootless DinD 尝试,但 live StatefulSet 与本目录 `values.yaml` 都已经使用 regular
DinD。首次 reconcile 的验收条件是修正 Helm 存储状态,同时 live Pod spec、PVC
identity、runner capacity 和在线状态保持不变。接管稳定后再用独立 PR 升级 chart。
接管分两阶段:第一阶段提交 `suspend: true` 的 HelmRelease、HelmRepository 和由
`values.yaml` 生成的 ConfigMap。Flux 只登记这些对象,不执行 Helm action。合并后
检查 HelmRepository Ready,并用固定 chart 重复比较期望清单与 live StatefulSet;
第二阶段解除 suspend。第一阶段已经确认 source Ready、完整 chart render 与 live
资源零差异,且登记过程中现有 runner 没有 rollout。失败重试使用
`RetryOnFailure`,不会用 stored rootless release 做 rollback。
## 历史 review-first bootstrap
这是 Flux 安装前执行过的一次性手动部署流程,保留用于恢复和审计:
1. Merge the reviewed PR.
2. As a Gitea site administrator, create an instance-scoped runner registration
token under **Site Administration → Actions → Runners**.
3. Store it as the `token` property at `kv/k8s/gitea-runner` without exposing it
in shell history:
```bash
read -rsp 'Runner token: ' runner_token
printf '%s' "$runner_token" | bao kv put kv/k8s/gitea-runner token=-
unset runner_token
```
4. From the updated `main`, create the namespace and ExternalSecret, then wait
for `SecretSynced=True`:
```bash
KUBECONFIG="$HOME/.kube/config" k3s kubectl apply \
-f platform/gitea-runner/namespace.yaml
KUBECONFIG="$HOME/.kube/config" k3s kubectl apply \
-f platform/gitea-runner/external-secret.yaml
KUBECONFIG="$HOME/.kube/config" k3s kubectl wait \
--namespace gitea-actions \
--for=condition=Ready externalsecret/gitea-runner-token \
--timeout=60s
```
5. Install chart `actions` version `0.1.1` from
`https://dl.gitea.com/charts/` with this `values.yaml`:
```bash
helm repo add gitea-charts https://dl.gitea.com/charts/
helm repo update gitea-charts
helm upgrade --install gitea-actions gitea-charts/actions \
--namespace gitea-actions \
--version 0.1.1 \
--values platform/gitea-runner/values.yaml \
--wait --timeout 10m
```
6. Confirm the runner is online, then re-run the queued lint workflow.
Do not deploy from an unmerged feature branch. Do not use `--set` for the token.
The 1 GiB PVC preserves `.runner` identity. Docker image layers are ephemeral;
the Pod has a 20 GiB ephemeral-storage limit. Terraform apply jobs must use a
workflow concurrency group because runner capacity does not serialize access to
a shared state.
@@ -1,17 +0,0 @@
apiVersion: spire.spiffe.io/v1alpha1
kind: ClusterSPIFFEID
metadata:
name: gitea-actions
spec:
className: spire-mgmt-spire
spiffeIDTemplate: spiffe://{{ .TrustDomain }}/ci/gitea-actions
namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: gitea-actions
podSelector:
matchLabels:
app.kubernetes.io/instance: gitea-actions
app.kubernetes.io/name: actions-runner
workloadSelectorTemplates:
- k8s:ns:gitea-actions
- k8s:sa:gitea-actions
-31
View File
@@ -1,31 +0,0 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: gitea-actions
namespace: gitea-actions
spec:
chart:
spec:
chart: actions
interval: 1h
sourceRef:
kind: HelmRepository
name: gitea-charts
version: 0.1.1
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: gitea-actions
targetNamespace: gitea-actions
timeout: 10m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: gitea-actions-values
@@ -1,8 +0,0 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: gitea-charts
namespace: gitea-actions
spec:
interval: 1h
url: https://dl.gitea.com/charts/
-18
View File
@@ -1,18 +0,0 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: gitea-actions-values
namespace: gitea-actions
files:
- values.yaml=values.yaml
resources:
- namespace.yaml
- serviceaccount.yaml
- clusterspiffeid.yaml
- external-secret.yaml
- helmrepository.yaml
- helmrelease.yaml
-4
View File
@@ -1,4 +0,0 @@
apiVersion: v1
kind: Namespace
metadata:
name: gitea-actions
@@ -1,6 +0,0 @@
apiVersion: v1
kind: ServiceAccount
metadata:
name: gitea-actions
namespace: gitea-actions
automountServiceAccountToken: false
-81
View File
@@ -1,81 +0,0 @@
enabled: true
giteaRootURL: http://gitea-http.gitea.svc.cluster.local:3000
existingSecret: gitea-runner-token
existingSecretKey: token
statefulset:
replicas: 1
timezone: Etc/UTC
serviceAccountName: gitea-actions
extraVolumes:
- name: spiffe-workload-api
csi:
driver: csi.spiffe.io
readOnly: true
securityContext:
fsGroup: 1000
# Chart 0.1.1 applies this block to both runner and DinD containers.
resources:
requests:
cpu: 250m
memory: 512Mi
ephemeral-storage: 2Gi
limits:
cpu: "4"
memory: 6Gi
ephemeral-storage: 20Gi
persistence:
size: 1Gi
runner:
registry: docker.io
repository: gitea/runner
tag: 2.3.0
pullPolicy: IfNotPresent
extraVolumeMounts:
- name: spiffe-workload-api
mountPath: /run/spire/agent-sockets
readOnly: true
config: |
log:
level: info
runner:
file: .runner
capacity: 4
timeout: 3h
shutdown_timeout: 3h
labels:
- self-hosted:docker://docker.gitea.com/runner-images:ubuntu-latest
cache:
enabled: false
container:
require_docker: true
docker_timeout: 300s
# Workflows must still request this exact bind mount explicitly. The
# allowlist prevents arbitrary host paths from reaching job containers.
valid_volumes:
- /run/spire/agent-sockets
dind:
# The node enforces AppArmor's unprivileged-userns restriction, which blocks
# rootlesskit even though this chart must run DinD privileged either way.
rootless: false
registry: docker.io
repository: docker
tag: 29.7.1-dind
pullPolicy: IfNotPresent
# Bind mounts are resolved by dockerd, so the CSI socket must exist in the
# DinD container as well as in the runner container.
extraVolumeMounts:
- name: spiffe-workload-api
mountPath: /run/spire/agent-sockets
readOnly: true
# k3s uses a 1450-byte pod MTU. Without matching it here, nested Actions
# networks advertise 1500 and GitHub TLS packets disappear on the outer
# overlay path while direct pod traffic remains healthy.
extraArgs:
- --mtu=1450
# --mtu only changes Docker's default bridge. act creates a user-defined
# bridge per job, so give every new bridge the same explicit default.
- --default-network-opt=bridge=com.docker.network.driver.mtu=1450
+4 -3
View File
@@ -35,10 +35,11 @@ ci_worker_password
## CI stream 约定
controller 首次启动时幂等创建 `CI_RUNNER` stream:`ci.runner.>`、
controller 使用 `CI_RUNNER` stream:`ci.runner.>`、
`WorkQueuePolicy`、file storage、24h/10000 条/256 MiB 上限。每类 runner 使用独立
subject 和 durable pull consumer;`ci.runner.<backend>.binding` 传递 runner 实际
领取任务后的身份绑定。同类型的多个 worker 共享 durable consumer。ACK 后消息立即
placement subject 和 durable pull consumer;assignment v2 subject 为
`ci.runner.<workload-class>.<driver>`,当前为 `ci.runner.container.kubernetes` 与
`ci.runner.vm.opensandbox`。同 placement 的多个 worker 共享 durable consumer。ACK 后消息立即
删除,不保存 CI 历史。
## 验证
@@ -32,6 +32,14 @@ datasources:
isDefault: true
jsonData:
prometheusType: Prometheus
- name: Alertmanager
uid: alertmanager
type: alertmanager
access: proxy
url: http://vmalertmanager-main.monitoring.svc:9093
jsonData:
implementation: prometheus
handleGrafanaManagedAlerts: false
- name: VictoriaLogs
type: victoriametrics-logs-datasource
access: proxy
@@ -1,18 +1,18 @@
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: gitea-runner-token
namespace: gitea-actions
name: alertmanager-telegram
namespace: monitoring
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: openbao
target:
name: alertmanager-telegram
creationPolicy: Owner
name: gitea-runner-token
data:
- secretKey: token
- secretKey: telegram_bot_token
remoteRef:
key: k8s/gitea-runner
property: token
key: k8s/alertmanager
property: telegram_bot_token
@@ -0,0 +1,257 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: kube-state-metrics
namespace: monitoring
spec:
interval: 30m
timeout: 5m
dependsOn:
- name: vm-operator
- name: prometheus-operator-crds
chart:
spec:
chart: kube-state-metrics
version: 8.5.0
sourceRef:
kind: HelmRepository
name: prometheus-community
namespace: monitoring
interval: 1h
driftDetection:
mode: enabled
install:
remediation:
retries: 3
upgrade:
remediation:
retries: 3
values:
fullnameOverride: kube-state-metrics
collectors:
- cronjobs
- daemonsets
- deployments
- jobs
- namespaces
- nodes
- persistentvolumeclaims
- persistentvolumes
- pods
- statefulsets
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 384Mi
prometheus:
monitor:
enabled: true
jobLabel: app.kubernetes.io/name
http:
interval: 30s
honorLabels: true
metrics:
interval: 30s
honorLabels: true
selfMonitor:
enabled: true
rbac:
extraRules:
- apiGroups:
- kustomize.toolkit.fluxcd.io
resources:
- kustomizations
verbs:
- list
- watch
- apiGroups:
- helm.toolkit.fluxcd.io
resources:
- helmreleases
verbs:
- list
- watch
- apiGroups:
- source.toolkit.fluxcd.io
resources:
- gitrepositories
- helmrepositories
- helmcharts
- ocirepositories
verbs:
- list
- watch
customResourceState:
enabled: true
config:
kind: CustomResourceStateMetrics
spec:
resources:
- groupVersionKind:
group: kustomize.toolkit.fluxcd.io
version: v1
kind: Kustomization
metricNamePrefix: gotk
metrics:
- name: resource_info
help: The current state of a GitOps Toolkit resource.
each:
type: Info
info:
labelsFromPath:
name:
- metadata
- name
labelsFromPath:
resource_namespace:
- metadata
- namespace
suspended:
- spec
- suspend
ready:
- status
- conditions
- '[type=Ready]'
- status
- groupVersionKind:
group: helm.toolkit.fluxcd.io
version: v2
kind: HelmRelease
metricNamePrefix: gotk
metrics:
- name: resource_info
help: The current state of a GitOps Toolkit resource.
each:
type: Info
info:
labelsFromPath:
name:
- metadata
- name
labelsFromPath:
resource_namespace:
- metadata
- namespace
suspended:
- spec
- suspend
ready:
- status
- conditions
- '[type=Ready]'
- status
- groupVersionKind:
group: source.toolkit.fluxcd.io
version: v1
kind: GitRepository
metricNamePrefix: gotk
metrics:
- name: resource_info
help: The current state of a GitOps Toolkit resource.
each:
type: Info
info:
labelsFromPath:
name:
- metadata
- name
labelsFromPath:
resource_namespace:
- metadata
- namespace
suspended:
- spec
- suspend
ready:
- status
- conditions
- '[type=Ready]'
- status
- groupVersionKind:
group: source.toolkit.fluxcd.io
version: v1
kind: HelmRepository
metricNamePrefix: gotk
metrics:
- name: resource_info
help: The current state of a GitOps Toolkit resource.
each:
type: Info
info:
labelsFromPath:
name:
- metadata
- name
labelsFromPath:
resource_namespace:
- metadata
- namespace
suspended:
- spec
- suspend
ready:
- status
- conditions
- '[type=Ready]'
- status
repository_type:
- spec
- type
- groupVersionKind:
group: source.toolkit.fluxcd.io
version: v1
kind: HelmChart
metricNamePrefix: gotk
metrics:
- name: resource_info
help: The current state of a GitOps Toolkit resource.
each:
type: Info
info:
labelsFromPath:
name:
- metadata
- name
labelsFromPath:
resource_namespace:
- metadata
- namespace
suspended:
- spec
- suspend
ready:
- status
- conditions
- '[type=Ready]'
- status
- groupVersionKind:
group: source.toolkit.fluxcd.io
version: v1
kind: OCIRepository
metricNamePrefix: gotk
metrics:
- name: resource_info
help: The current state of a GitOps Toolkit resource.
each:
type: Info
info:
labelsFromPath:
name:
- metadata
- name
labelsFromPath:
resource_namespace:
- metadata
- namespace
suspended:
- spec
- suspend
ready:
- status
- conditions
- '[type=Ready]'
- status
@@ -6,6 +6,7 @@ resources:
- vmagent.yaml
- vmalert.yaml
- vmalertmanager.yaml
- alertmanager-external-secret.yaml
- reload-rbac.yaml
- rules/vm-health.yaml
- rules/vmagent.yaml
@@ -16,3 +17,12 @@ resources:
- scrapes/kubelet.yaml
- exporters/node-exporter.yaml
- exporters/process-exporter.yaml
- exporters/kube-state-metrics-helmrelease.yaml
- rules/kubernetes-health.yaml
- rules/host-health.yaml
- rules/monitoring-delivery.yaml
- scrapes/platform-controllers.yaml
- rules/platform-controllers.yaml
- scrapes/cnpg.yaml
- scrapes/seaweedfs.yaml
- rules/data-services.yaml
@@ -0,0 +1,148 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: data-services
namespace: monitoring
spec:
groups:
- name: data-services
interval: 30s
rules:
- alert: DataServiceMetricsUnavailable
expr: up{job="cnpg"} == 0 or absent(up{job="cnpg"})
for: 5m
labels:
severity: critical
job: cnpg
annotations:
summary: 数据服务 {{ $labels.job }} 指标不可用
description: 检查监控对象、Pod 与 exporter;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: DataServiceMetricsUnavailable
expr: up{job="seaweedfs-master"} == 0 or absent(up{job="seaweedfs-master"})
for: 5m
labels:
severity: critical
job: seaweedfs-master
annotations:
summary: 数据服务 {{ $labels.job }} 指标不可用
description: 检查监控对象、Pod 与 exporter;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: DataServiceMetricsUnavailable
expr: up{job="seaweedfs-volume"} == 0 or absent(up{job="seaweedfs-volume"})
for: 5m
labels:
severity: critical
job: seaweedfs-volume
annotations:
summary: 数据服务 {{ $labels.job }} 指标不可用
description: 检查监控对象、Pod 与 exporter;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: DataServiceMetricsUnavailable
expr: up{job="seaweedfs-filer"} == 0 or absent(up{job="seaweedfs-filer"})
for: 5m
labels:
severity: critical
job: seaweedfs-filer
annotations:
summary: 数据服务 {{ $labels.job }} 指标不可用
description: 检查监控对象、Pod 与 exporter;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: CnpgPostgresDown
expr: cnpg_collector_up{job="cnpg"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: CNPG {{ $labels.pod }} PostgreSQL 不可用
description: 检查数据库 Pod 状态、实例日志和存储;exporter up 不代表数据库可用。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: CnpgMetricsCollectionFailed
expr: cnpg_collector_last_collection_error{job="cnpg"} == 1
for: 5m
labels:
severity: warning
annotations:
summary: CNPG {{ $labels.pod }} SQL 指标采集失败
description: 检查内置 exporter 日志、数据库连接与查询兼容性。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: CnpgConnectionsHigh
expr: sum by (job,namespace,pod) (cnpg_backends_total{job="cnpg"}) / on(job,namespace,pod)
max by(job,namespace,pod) (cnpg_pg_settings_setting{job="cnpg",name="max_connections"})
> 0.8
for: 10m
labels:
severity: warning
annotations:
summary: CNPG {{ $labels.pod }} 连接使用率超过 80%
description: 检查连接池、空闲连接和连接泄漏,不直接提高 max_connections。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: CnpgLongTransaction
expr: cnpg_backends_max_tx_duration_seconds{job="cnpg",state!="idle"} > 300
for: 5m
labels:
severity: warning
annotations:
summary: CNPG {{ $labels.pod }} 存在长事务
description: 事务持续超过五分钟并维持五分钟,检查阻塞与 idle in transaction;先确认业务再处理。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: NatsJetStreamStorageHigh
expr: nats_varz_jetstream_stats_storage{job="nats/nats"} / nats_varz_jetstream_config_max_storage{job="nats/nats"}
> 0.8
for: 10m
labels:
severity: warning
annotations:
summary: NATS JetStream 文件容量超过 80%
description: 检查服务端配额与消息保留策略;该指标不代表单个 Account 或 stream 配额。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: NatsJetStreamMemoryHigh
expr: nats_varz_jetstream_stats_memory{job="nats/nats"} / nats_varz_jetstream_config_max_memory{job="nats/nats"}
> 0.8
for: 10m
labels:
severity: warning
annotations:
summary: NATS JetStream 内存容量超过 80%
description: 检查 memory store 消息保留与服务端内存配额。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: NatsSlowConsumers
expr: increase(nats_varz_slow_consumers{job="nats/nats"}[10m]) > 0
for: 5m
labels:
severity: warning
annotations:
summary: NATS 检测到慢消费者
description: 检查消费者处理速率与网络;此规则不能代替 JetStream consumer 积压监控。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: SeaweedVolumeDiskLow
expr: SeaweedFS_volumeServer_resource{job="seaweedfs-volume",type="avail"} /
ignoring(type) SeaweedFS_volumeServer_resource{job="seaweedfs-volume",type="all"}
< 0.1
for: 10m
labels:
severity: warning
annotations:
summary: SeaweedFS {{ $labels.name }} 可用空间不足 10%
description: 检查 volume 文件系统容量、保留策略与底层 ZFS;不要直接删除卷文件。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: SeaweedVolumeDiskReadOnly
expr: SeaweedFS_volumeServer_read_only_volumes{job="seaweedfs-volume",type="isDiskSpaceLow"}
> 0
for: 5m
labels:
severity: critical
annotations:
summary: SeaweedFS 因磁盘空间不足限制卷写入
description: 检查磁盘容量和 volume 日志;主动只读或满卷切换不按此故障处理。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- alert: SeaweedVolumeWriteFailures
expr: increase(SeaweedFS_volumeServer_file_write_failures{job="seaweedfs-volume"}[5m])
> 0
for: 2m
labels:
severity: warning
annotations:
summary: SeaweedFS 卷写入失败
description: 检查 volume 日志、磁盘状态及上游请求;不代表已经验证 S3 端到端可用。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
@@ -0,0 +1,66 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: host-health
namespace: monitoring
labels:
app.kubernetes.io/part-of: victoria-metrics
spec:
groups:
- name: host-health
interval: 30s
rules:
- alert: HostMemoryLow
expr: (node_memory_MemAvailable_bytes{job="node-exporter"} / node_memory_MemTotal_bytes{job="node-exporter"})
< 0.10
for: 10m
labels:
severity: warning
annotations:
summary: 主机 {{ $labels.instance }} 可用内存不足 10%
description: 检查主机内存、Swap、进程 PSS;仅使用 node-exporter job,避免重复采集统计。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: HostFilesystemSpaceLow
expr: (node_filesystem_avail_bytes{job="node-exporter",fstype!~"tmpfs|devtmpfs|overlay|squashfs|nsfs|fuse.*",mountpoint!~"/run(/.*)?|/var/lib/kubelet/.*|/var/lib/docker/.*"}
/ node_filesystem_size_bytes{job="node-exporter",fstype!~"tmpfs|devtmpfs|overlay|squashfs|nsfs|fuse.*",mountpoint!~"/run(/.*)?|/var/lib/kubelet/.*|/var/lib/docker/.*"}
< 0.10) and on(instance,device,mountpoint) (node_filesystem_readonly{job="node-exporter",fstype!~"tmpfs|devtmpfs|overlay|squashfs|nsfs|fuse.*",mountpoint!~"/run(/.*)?|/var/lib/kubelet/.*|/var/lib/docker/.*"}
== 0)
for: 15m
labels:
severity: warning
annotations:
summary: 主机 {{ $labels.instance }} {{ $labels.mountpoint }} 可用空间不足 10%
description: 检查文件系统和 ZFS dataset 容量,清理前确认数据归属。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: HostFilesystemInodesLow
expr: (node_filesystem_files_free{job="node-exporter",fstype!~"tmpfs|devtmpfs|overlay|squashfs|nsfs|fuse.*",mountpoint!~"/run(/.*)?|/var/lib/kubelet/.*|/var/lib/docker/.*"}
/ node_filesystem_files{job="node-exporter",fstype!~"tmpfs|devtmpfs|overlay|squashfs|nsfs|fuse.*",mountpoint!~"/run(/.*)?|/var/lib/kubelet/.*|/var/lib/docker/.*"}
< 0.10) and on(instance,device,mountpoint) (node_filesystem_files{job="node-exporter",fstype!~"tmpfs|devtmpfs|overlay|squashfs|nsfs|fuse.*",mountpoint!~"/run(/.*)?|/var/lib/kubelet/.*|/var/lib/docker/.*"}
> 0) and on(instance,device,mountpoint) (node_filesystem_readonly{job="node-exporter",fstype!~"tmpfs|devtmpfs|overlay|squashfs|nsfs|fuse.*",mountpoint!~"/run(/.*)?|/var/lib/kubelet/.*|/var/lib/docker/.*"}
== 0)
for: 15m
labels:
severity: warning
annotations:
summary: 主机 {{ $labels.instance }} {{ $labels.mountpoint }} inode 不足 10%
description: 检查小文件数量,确认 filesystem 是否耗尽 inode。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: HostFilesystemReadOnly
expr: node_filesystem_readonly{job="node-exporter",fstype!~"tmpfs|devtmpfs|overlay|squashfs|nsfs|fuse.*",mountpoint!~"/run(/.*)?|/var/lib/kubelet/.*|/var/lib/docker/.*"}
== 1
for: 5m
labels:
severity: critical
annotations:
summary: 主机 {{ $labels.instance }} {{ $labels.mountpoint }} 变为只读
description: 排查磁盘和文件系统错误;不要直接强制重新挂载。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: HostZpoolUnhealthy
expr: node_zfs_zpool_state{job="node-exporter",state!="online"} == 1
for: 5m
labels:
severity: critical
annotations:
summary: 主机 {{ $labels.instance }} ZFS 池 {{ $labels.zpool }} 状态 {{ $labels.state }}
description: 检查 zpool status、磁盘和冗余;零值状态不触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
@@ -0,0 +1,121 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: kubernetes-health
namespace: monitoring
labels:
app.kubernetes.io/part-of: victoria-metrics
spec:
groups:
- name: kubernetes-health
interval: 30s
rules:
- alert: KubeNodeNotReady
expr: max by (node) (kube_node_status_condition{job="kube-state-metrics",condition="Ready",status="true"})
== 0
for: 5m
labels:
severity: critical
annotations:
summary: 节点 {{ $labels.node }} 未就绪
description: 检查节点、kubelet 和网络;维护时按 node 精确静默。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubeNodePressure
expr: max by (node, condition) (kube_node_status_condition{job="kube-state-metrics",condition=~"MemoryPressure|DiskPressure|PIDPressure",status="true"})
== 1
for: 5m
labels:
severity: warning
annotations:
summary: 节点 {{ $labels.node }} 出现 {{ $labels.condition }}
description: 检查内存、磁盘/inode 或 PID 资源压力。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubePodCrashLooping
expr: max by (namespace,pod,container) (kube_pod_container_status_waiting_reason{job="kube-state-metrics",reason="CrashLoopBackOff"})
== 1
for: 10m
labels:
severity: warning
annotations:
summary: 容器 {{ $labels.namespace }}/{{ $labels.pod }}/{{ $labels.container }} 持续崩溃
description: 检查容器退出原因和日志;持续 10 分钟后通知,避免短暂重启噪声。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubePodFrequentRestarts
expr: sum by (namespace,pod,container) (increase(kube_pod_container_status_restarts_total{job="kube-state-metrics"}[15m]))
> 3
for: 5m
labels:
severity: warning
annotations:
summary: 容器 {{ $labels.namespace }}/{{ $labels.pod }}/{{ $labels.container }} 频繁重启
description: 15 分钟内重启超过 3 次,检查退出原因和资源限制。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubePodOOMKilled
expr: (max by (namespace,pod,container) (kube_pod_container_status_last_terminated_reason{job="kube-state-metrics",reason="OOMKilled"})
== 1) and on(namespace,pod,container) (sum by (namespace,pod,container) (increase(kube_pod_container_status_restarts_total{job="kube-state-metrics"}[10m]))
> 0)
for: 1m
labels:
severity: warning
annotations:
summary: 容器 {{ $labels.namespace }}/{{ $labels.pod }}/{{ $labels.container }} 发生 OOM 重启
description: 仅在近期有重启且最近终止原因为 OOMKilled 时触发,检查内存峰值和限制。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubePodPending
expr: max by (namespace,pod) (kube_pod_status_phase{job="kube-state-metrics",phase="Pending"}) == 1
for: 15m
labels:
severity: warning
annotations:
summary: Pod {{ $labels.namespace }}/{{ $labels.pod }} 长时间 Pending
description: 检查调度事件、资源、卷和镜像拉取。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubeDeploymentUnavailable
expr: (max by(namespace,deployment) (kube_deployment_spec_replicas{job="kube-state-metrics"}) - max by(namespace,deployment)
(kube_deployment_status_replicas_available{job="kube-state-metrics"})) > 0
for: 10m
labels:
severity: warning
annotations:
summary: Deployment {{ $labels.namespace }}/{{ $labels.deployment }} 副本不足
description: 可用副本持续 10 分钟低于声明值;缩容至零不会触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubeStatefulSetUnavailable
expr: (max by(namespace,statefulset) (kube_statefulset_replicas{job="kube-state-metrics"}) - max by(namespace,statefulset)
(kube_statefulset_status_replicas_ready{job="kube-state-metrics"})) > 0
for: 10m
labels:
severity: warning
annotations:
summary: StatefulSet {{ $labels.namespace }}/{{ $labels.statefulset }} 副本不足
description: 检查 Pod 就绪、存储和依赖;已知维护按工作负载精确静默。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubeDaemonSetUnavailable
expr: (max by(namespace,daemonset) (kube_daemonset_status_desired_number_scheduled{job="kube-state-metrics"})
- max by(namespace,daemonset) (kube_daemonset_status_number_ready{job="kube-state-metrics"})) > 0
for: 10m
labels:
severity: warning
annotations:
summary: DaemonSet {{ $labels.namespace }}/{{ $labels.daemonset }} 副本不足
description: 检查节点和 Pod 就绪;不按整个 namespace 屏蔽。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubePVCPending
expr: max by(namespace,persistentvolumeclaim) (kube_persistentvolumeclaim_status_phase{job="kube-state-metrics",phase="Pending"})
== 1
for: 15m
labels:
severity: warning
annotations:
summary: PVC {{ $labels.namespace }}/{{ $labels.persistentvolumeclaim }} 长时间 Pending
description: 检查 StorageClass、调度拓扑、容量与 provisioner。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: KubeJobFailed
expr: max by(namespace,job_name) (kube_job_failed{job="kube-state-metrics",condition="true"}) == 1
for: 5m
labels:
severity: warning
annotations:
summary: Job {{ $labels.namespace }}/{{ $labels.job_name }} 已失败
description: 检查失败 Job 的 Pod 和任务日志;成功重试中的失败 Pod 计数不作为失败 Job。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
@@ -0,0 +1,58 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: monitoring-delivery
namespace: monitoring
labels:
app.kubernetes.io/part-of: victoria-metrics
spec:
groups:
- name: monitoring-delivery
interval: 30s
rules:
- alert: KubeStateMetricsUnavailable
expr: up{job="kube-state-metrics"} == 0 or absent(up{job="kube-state-metrics"})
for: 5m
labels:
severity: critical
annotations:
summary: kube-state-metrics 采集不可用
description: 检查 HelmRelease、ServiceMonitor 转换、targets 与 exporter;目标消失也会触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: NodeExporterUnavailable
expr: up{job="node-exporter"} == 0 or absent(up{job="node-exporter"})
for: 5m
labels:
severity: critical
annotations:
summary: node-exporter 采集不可用
description: 检查节点与 exporter,主机健康规则依赖此采集。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: NatsMetricsUnavailable
expr: up{job="nats/nats"} == 0 or absent(up{job="nats/nats"})
for: 5m
labels:
severity: warning
annotations:
summary: NATS 指标采集不可用
description: 依次检查 PodMonitor、转换后的 VMPodScrape、:7777 exporter 及 targets。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: AlertmanagerTelegramDeliveryFailed
expr: sum by(job,instance,integration) (increase(alertmanager_notifications_failed_total{job="vmalertmanager-main",integration="telegram"}[5m]))
> 0
for: 1m
labels:
severity: critical
annotations:
summary: Alertmanager 向 Telegram 发送失败
description: 检查网络、bot token 和频道权限。若 Telegram 完全不可达,本告警也无法经同一渠道送达;需在 Alertmanager/Grafana 排查,集群外心跳另行建设。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- alert: AlertmanagerConfigurationReloadFailed
expr: alertmanager_config_last_reload_successful{job="vmalertmanager-main"} == 0
for: 5m
labels:
severity: critical
annotations:
summary: Alertmanager 配置加载失败
description: 检查 operator 和 Alertmanager 日志、配置格式及 Secret 引用。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
@@ -0,0 +1,161 @@
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: platform-controllers
namespace: monitoring
spec:
groups:
- name: platform-controllers
interval: 30s
rules:
- alert: PlatformControllerMetricsUnavailable
expr: up{job="cert-manager"} == 0 or absent(up{job="cert-manager"})
for: 5m
labels:
severity: warning
job: cert-manager
annotations:
summary: 控制器 {{ $labels.job }} 指标不可用
description: 检查对应 Pod、监控 CR 转换与 targets;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: PlatformControllerMetricsUnavailable
expr: up{job="external-secrets"} == 0 or absent(up{job="external-secrets"})
for: 5m
labels:
severity: warning
job: external-secrets
annotations:
summary: 控制器 {{ $labels.job }} 指标不可用
description: 检查对应 Pod、监控 CR 转换与 targets;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: PlatformControllerMetricsUnavailable
expr: up{job="external-secrets-cert-controller"} == 0 or absent(up{job="external-secrets-cert-controller"})
for: 5m
labels:
severity: warning
job: external-secrets-cert-controller
annotations:
summary: 控制器 {{ $labels.job }} 指标不可用
description: 检查对应 Pod、监控 CR 转换与 targets;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: PlatformControllerMetricsUnavailable
expr: up{job="external-secrets-webhook"} == 0 or absent(up{job="external-secrets-webhook"})
for: 5m
labels:
severity: warning
job: external-secrets-webhook
annotations:
summary: 控制器 {{ $labels.job }} 指标不可用
description: 检查对应 Pod、监控 CR 转换与 targets;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: PlatformControllerMetricsUnavailable
expr: up{job="helm-controller"} == 0 or absent(up{job="helm-controller"})
for: 5m
labels:
severity: warning
job: helm-controller
annotations:
summary: 控制器 {{ $labels.job }} 指标不可用
description: 检查对应 Pod、监控 CR 转换与 targets;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: PlatformControllerMetricsUnavailable
expr: up{job="kustomize-controller"} == 0 or absent(up{job="kustomize-controller"})
for: 5m
labels:
severity: warning
job: kustomize-controller
annotations:
summary: 控制器 {{ $labels.job }} 指标不可用
description: 检查对应 Pod、监控 CR 转换与 targets;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: PlatformControllerMetricsUnavailable
expr: up{job="source-controller"} == 0 or absent(up{job="source-controller"})
for: 5m
labels:
severity: warning
job: source-controller
annotations:
summary: 控制器 {{ $labels.job }} 指标不可用
description: 检查对应 Pod、监控 CR 转换与 targets;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: PlatformControllerMetricsUnavailable
expr: up{job="notification-controller"} == 0 or absent(up{job="notification-controller"})
for: 5m
labels:
severity: warning
job: notification-controller
annotations:
summary: 控制器 {{ $labels.job }} 指标不可用
description: 检查对应 Pod、监控 CR 转换与 targets;目标消失同样触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: CertificateNotReady
expr: certmanager_certificate_ready_status{job="cert-manager",condition="True"}
== 0
for: 15m
labels:
severity: warning
annotations:
summary: 证书 {{ $labels.namespace }}/{{ $labels.name }} 未就绪
description: 检查 Certificate、CertificateRequest、Issuer、Order 与 Challenge;允许短暂签发过程。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: CertificateExpiringSoon
expr: (certmanager_certificate_expiration_timestamp_seconds{job="cert-manager"}
- time() < 7 * 86400) and (certmanager_certificate_expiration_timestamp_seconds{job="cert-manager"}
- time() >= 86400)
for: 15m
labels:
severity: warning
annotations:
summary: 证书 {{ $labels.namespace }}/{{ $labels.name }} 将在七天内到期
description: 检查自动续期链路及实际入口证书;一天内到期由 critical 规则接替。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: CertificateExpiryCritical
expr: certmanager_certificate_expiration_timestamp_seconds{job="cert-manager"}
- time() < 86400
for: 5m
labels:
severity: critical
annotations:
summary: 证书 {{ $labels.namespace }}/{{ $labels.name }} 即将到期或已过期
description: 不足一天或已经过期;立即检查续期与入口证书。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: ExternalSecretNotReady
expr: externalsecret_status_condition{job="external-secrets",condition="Ready",status="True"}
== 0
for: 10m
labels:
severity: warning
annotations:
summary: Secret 同步 {{ $labels.namespace }}/{{ $labels.name }} 未就绪
description: 检查 ExternalSecret 状态及 provider 权限、网络;不要输出 Secret 内容。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: ClusterSecretStoreNotReady
expr: clustersecretstore_status_condition{job="external-secrets",condition="Ready",status="True"}
== 0
for: 5m
labels:
severity: critical
annotations:
summary: ClusterSecretStore {{ $labels.name }} 未就绪
description: 检查 OpenBao 可达性与 ESO 身份鉴权;可能影响多个命名空间。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: FluxResourceNotReady
expr: gotk_resource_info{job="kube-state-metrics",ready!="True",suspended!="true",repository_type!="oci"}
== 1
for: 15m
labels:
severity: warning
annotations:
summary: Flux {{ $labels.customresource_kind }} {{ $labels.resource_namespace
}}/{{ $labels.name }} 未就绪
description: 检查对象 conditions、依赖和 controller 日志;主动 suspend 不触发,不自动解除暂停。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- alert: FluxStateMetricsMissing
expr: absent(gotk_resource_info{job="kube-state-metrics",customresource_kind="Kustomization",resource_namespace="flux-system",name="flux-system"})
for: 5m
labels:
severity: warning
annotations:
summary: Flux 对象状态指标缺失
description: 检查 kube-state-metrics 自定义资源配置、RBAC 与采集;不能把空结果当作健康。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
@@ -0,0 +1,22 @@
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: cnpg
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- shared-db
selector:
matchLabels:
cnpg.io/cluster: shared-postgresql
cnpg.io/podRole: instance
podTargetLabels:
- cnpg.io/cluster
podMetricsEndpoints:
- port: metrics
interval: 30s
honorLabels: true
relabelings:
- targetLabel: job
replacement: cnpg
@@ -1,6 +1,6 @@
# Pull metrics from Docker hosts running ../../docker-hosts/compose.yaml.
# vmagent scrapes each host's node-exporter (:9100) and cAdvisor (:8080) over the LAN.
# Add one target block per host; keep the `host` label in sync with its HOST_LABEL.
# 仅抓取已存在的宿主机 node-exporter。旧 :8080 目标实际返回 nginx 404,
# 独立 Docker cAdvisor 尚未部署;部署并确认端口后再添加其目标。
# Kubernetes 容器指标由 kubelet.yaml 的 /metrics/cadvisor 独立采集。
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMStaticScrape
metadata:
@@ -17,13 +17,6 @@ spec:
labels:
job: node-exporter
host: docker-01
- targets:
- "192.168.10.127:8080"
labels:
job: cadvisor
host: docker-01
# --- add more hosts below, mirroring the two blocks above ---
# --- 新增主机前确认 exporter 已部署且 /metrics 返回成功 ---
# - targets: ["192.168.10.x:9100"]
# labels: { job: node-exporter, host: docker-02 }
# - targets: ["192.168.10.x:8080"]
# labels: { job: cadvisor, host: docker-02 }
@@ -13,6 +13,9 @@ spec:
honorLabels: true
honorTimestamps: false
interval: 30s
# k3s 此端点包含 apiserver/etcd 指标,响应已超过默认 16 MiB。
# 仅放宽 kubelet 任务,其他采集目标保留默认限制。
max_scrape_size: "32MiB"
tlsConfig:
insecureSkipVerify: true
caFile: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
@@ -0,0 +1,57 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: cert-manager
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- cert-manager
selector:
matchLabels:
app.kubernetes.io/name: cert-manager
app.kubernetes.io/component: controller
jobLabel: app.kubernetes.io/name
endpoints:
- port: http-metrics
interval: 30s
scrapeTimeout: 10s
honorLabels: true
---
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: external-secrets
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- external-secrets
selector:
matchLabels:
app.kubernetes.io/instance: external-secrets
jobLabel: app.kubernetes.io/name
podMetricsEndpoints:
- port: metrics
interval: 30s
scrapeTimeout: 10s
honorLabels: true
---
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: flux-controllers
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- flux-system
selector:
matchLabels:
app.kubernetes.io/part-of: flux
jobLabel: app
podMetricsEndpoints:
- port: http-prom
interval: 30s
scrapeTimeout: 10s
honorLabels: true
@@ -0,0 +1,25 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: seaweedfs
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- seaweedfs
selector:
matchLabels:
app.kubernetes.io/name: seaweedfs
app.kubernetes.io/instance: seaweedfs
endpoints:
- port: metrics
interval: 30s
honorLabels: true
relabelings:
- sourceLabels:
- __meta_kubernetes_service_name
regex: seaweedfs-(master|volume|filer)
action: keep
- sourceLabels:
- __meta_kubernetes_service_name
targetLabel: job
@@ -0,0 +1,20 @@
#!/usr/bin/env bash
# 依赖 Python 3 + PyYAML,以及 PATH 中的 promtool(验证版本 3.5.0)。
set -euo pipefail
test_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
check_dir="$(mktemp -d)"
trap 'rm -rf "$check_dir"' EXIT
python3 - "$test_dir/../rules" "$check_dir/rules.yaml" <<'PY'
import pathlib
import sys
import yaml
groups = []
for name in ("kubernetes-health", "host-health", "monitoring-delivery", "platform-controllers", "data-services"):
path = pathlib.Path(sys.argv[1]) / f"{name}.yaml"
groups.extend(yaml.safe_load(path.read_text())["spec"]["groups"])
pathlib.Path(sys.argv[2]).write_text(yaml.safe_dump({"groups": groups}, allow_unicode=True))
PY
cp "$test_dir/"*.test.yaml "$check_dir/"
promtool check rules "$check_dir/rules.yaml"
promtool test rules "$check_dir/"*.test.yaml
@@ -0,0 +1,137 @@
rule_files:
- rules.yaml
evaluation_interval: 1m
tests:
- name: 健康 ZFS 的零值故障状态不报警
interval: 1m
input_series:
- series: node_zfs_zpool_state{job="node-exporter",instance="laptop",zpool="data",state="degraded"}
values: '0x20'
- series: node_zfs_zpool_state{job="node-exporter",instance="laptop",zpool="data",state="online"}
values: 1x20
alert_rule_test:
- eval_time: 10m
alertname: HostZpoolUnhealthy
exp_alerts: &id001 []
- name: ZFS 降级持续五分钟触发
interval: 1m
input_series:
- series: node_zfs_zpool_state{job="node-exporter",instance="laptop",zpool="data",state="degraded"}
values: 1x20
alert_rule_test:
- eval_time: 4m
alertname: HostZpoolUnhealthy
exp_alerts: *id001
- eval_time: 6m
alertname: HostZpoolUnhealthy
exp_alerts:
- exp_labels:
job: node-exporter
instance: laptop
zpool: data
state: degraded
severity: critical
exp_annotations:
summary: 主机 laptop ZFS 池 data 状态 degraded
description: 检查 zpool status、磁盘和冗余;零值状态不触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- name: 忽略旧 docker-hosts 重复主机指标
interval: 1m
input_series:
- series: node_memory_MemAvailable_bytes{job="node-exporter",instance="laptop"}
values: 50x20
- series: node_memory_MemTotal_bytes{job="node-exporter",instance="laptop"}
values: 100x20
- series: node_memory_MemAvailable_bytes{job="docker-hosts",instance="laptop"}
values: 1x20
- series: node_memory_MemTotal_bytes{job="docker-hosts",instance="laptop"}
values: 100x20
alert_rule_test:
- eval_time: 15m
alertname: HostMemoryLow
exp_alerts: *id001
- name: 临时文件系统耗尽不触发持久磁盘告警
interval: 1m
input_series:
- series: node_filesystem_avail_bytes{job="node-exporter",instance="laptop",device="tmpfs",mountpoint="/run",fstype="tmpfs"}
values: 1x20
- series: node_filesystem_size_bytes{job="node-exporter",instance="laptop",device="tmpfs",mountpoint="/run",fstype="tmpfs"}
values: 100x20
- series: node_filesystem_readonly{job="node-exporter",instance="laptop",device="tmpfs",mountpoint="/run",fstype="tmpfs"}
values: '0x20'
alert_rule_test:
- eval_time: 20m
alertname: HostFilesystemSpaceLow
exp_alerts: *id001
- name: Deployment 允许短暂滚动且缩零不报错
interval: 1m
input_series:
- series: kube_deployment_spec_replicas{job="kube-state-metrics",namespace="app",deployment="web"}
values: 2x20
- series: kube_deployment_status_replicas_available{job="kube-state-metrics",namespace="app",deployment="web"}
values: 1x20
- series: kube_deployment_spec_replicas{job="kube-state-metrics",namespace="app",deployment="scaled-down"}
values: '0x20'
- series: kube_deployment_status_replicas_available{job="kube-state-metrics",namespace="app",deployment="scaled-down"}
values: '0x20'
alert_rule_test:
- eval_time: 5m
alertname: KubeDeploymentUnavailable
exp_alerts: *id001
- eval_time: 11m
alertname: KubeDeploymentUnavailable
exp_alerts:
- exp_labels:
namespace: app
deployment: web
severity: warning
exp_annotations:
summary: Deployment app/web 副本不足
description: 可用副本持续 10 分钟低于声明值;缩容至零不会触发。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- name: NATS 发现集合完全消失也报警
interval: 1m
input_series: []
alert_rule_test:
- eval_time: 6m
alertname: NatsMetricsUnavailable
exp_alerts:
- exp_labels:
job: nats/nats
severity: warning
exp_annotations:
summary: NATS 指标采集不可用
description: 依次检查 PodMonitor、转换后的 VMPodScrape、:7777 exporter 及 targets。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
- name: 历史 OOM 没有新重启不反复报警
interval: 1m
input_series:
- series: kube_pod_container_status_last_terminated_reason{job="kube-state-metrics",namespace="app",pod="web",container="web",reason="OOMKilled"}
values: 1x20
- series: kube_pod_container_status_restarts_total{job="kube-state-metrics",namespace="app",pod="web",container="web"}
values: 5x20
alert_rule_test:
- eval_time: 10m
alertname: KubePodOOMKilled
exp_alerts: *id001
- name: Telegram 失败跨 reason 汇总为一条
interval: 1m
input_series:
- series: alertmanager_notifications_failed_total{job="vmalertmanager-main",instance="am",integration="telegram",reason="clientError"}
values: 0+1x20
- series: alertmanager_notifications_failed_total{job="vmalertmanager-main",instance="am",integration="telegram",reason="serverError"}
values: 0+1x20
alert_rule_test:
- eval_time: 6m
alertname: AlertmanagerTelegramDeliveryFailed
exp_alerts:
- exp_labels:
job: vmalertmanager-main
instance: am
integration: telegram
severity: critical
exp_annotations:
summary: Alertmanager 向 Telegram 发送失败
description: 检查网络、bot token 和频道权限。若 Telegram 完全不可达,本告警也无法经同一渠道送达;需在 Alertmanager/Grafana
排查,集群外心跳另行建设。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-foundation.md
@@ -0,0 +1,185 @@
rule_files:
- rules.yaml
evaluation_interval: 1m
tests:
- name: Flux 暂停的失败对象不报警
interval: 1m
input_series:
- series: gotk_resource_info{job="kube-state-metrics",ready="False",suspended="true",customresource_kind="HelmRelease",resource_namespace="app",name="paused"}
values: 1x30
alert_rule_test:
- eval_time: 20m
alertname: FluxResourceNotReady
exp_alerts: []
- name: Flux 缺省 suspend 仍检测失败
interval: 1m
input_series:
- series: gotk_resource_info{job="kube-state-metrics",ready="False",customresource_kind="HelmRelease",resource_namespace="app",name="broken"}
values: 1x30
alert_rule_test:
- eval_time: 10m
alertname: FluxResourceNotReady
exp_alerts: []
- eval_time: 16m
alertname: FluxResourceNotReady
exp_alerts:
- exp_labels:
job: kube-state-metrics
ready: 'False'
customresource_kind: HelmRelease
resource_namespace: app
name: broken
severity: warning
exp_annotations:
summary: Flux HelmRelease app/broken 未就绪
description: 检查对象 conditions、依赖和 controller 日志;主动 suspend 不触发,不自动解除暂停。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- name: 证书一天内到期只由 critical 接替
interval: 1m
input_series:
- series: certmanager_certificate_expiration_timestamp_seconds{job="cert-manager",namespace="app",name="tls"}
values: 3600x30
alert_rule_test:
- eval_time: 20m
alertname: CertificateExpiringSoon
exp_alerts: []
- name: 过期证书仍持续 critical
interval: 1m
input_series:
- series: certmanager_certificate_expiration_timestamp_seconds{job="cert-manager",namespace="app",name="tls"}
values: '0x30'
alert_rule_test:
- eval_time: 4m
alertname: CertificateExpiryCritical
exp_alerts: []
- eval_time: 6m
alertname: CertificateExpiryCritical
exp_alerts:
- exp_labels:
job: cert-manager
namespace: app
name: tls
severity: critical
exp_annotations:
summary: 证书 app/tls 即将到期或已过期
description: 不足一天或已经过期;立即检查续期与入口证书。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- name: ESO False 零值不能导致健康对象误报
interval: 1m
input_series:
- series: externalsecret_status_condition{job="external-secrets",condition="Ready",status="False",namespace="app",name="credentials"}
values: '0x30'
- series: externalsecret_status_condition{job="external-secrets",condition="Ready",status="True",namespace="app",name="credentials"}
values: 1x30
alert_rule_test:
- eval_time: 20m
alertname: ExternalSecretNotReady
exp_alerts: []
- name: ESO 失败持续十分钟报警
interval: 1m
input_series:
- series: externalsecret_status_condition{job="external-secrets",condition="Ready",status="True",namespace="app",name="credentials"}
values: '0x30'
alert_rule_test:
- eval_time: 9m
alertname: ExternalSecretNotReady
exp_alerts: []
- eval_time: 11m
alertname: ExternalSecretNotReady
exp_alerts:
- exp_labels:
job: external-secrets
condition: Ready
status: 'True'
namespace: app
name: credentials
severity: warning
exp_annotations:
summary: Secret 同步 app/credentials 未就绪
description: 检查 ExternalSecret 状态及 provider 权限、网络;不要输出 Secret 内容。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- name: CNPG 聚合数据库连接并对齐实例限额
interval: 1m
input_series:
- series: cnpg_backends_total{job="cnpg",namespace="db",pod="pg-1",datname="a"}
values: 50x30
- series: cnpg_backends_total{job="cnpg",namespace="db",pod="pg-1",datname="b"}
values: 40x30
- series: cnpg_pg_settings_setting{job="cnpg",namespace="db",pod="pg-1",name="max_connections"}
values: 100x30
alert_rule_test:
- eval_time: 9m
alertname: CnpgConnectionsHigh
exp_alerts: []
- eval_time: 11m
alertname: CnpgConnectionsHigh
exp_alerts:
- exp_labels:
job: cnpg
namespace: db
pod: pg-1
severity: warning
exp_annotations:
summary: CNPG pg-1 连接使用率超过 80%
description: 检查连接池、空闲连接和连接泄漏,不直接提高 max_connections。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- name: SeaweedFS 磁盘分母忽略 type 标签
interval: 1m
input_series:
- series: SeaweedFS_volumeServer_resource{job="seaweedfs-volume",name="/data",type="avail"}
values: 5x30
- series: SeaweedFS_volumeServer_resource{job="seaweedfs-volume",name="/data",type="all"}
values: 100x30
alert_rule_test:
- eval_time: 9m
alertname: SeaweedVolumeDiskLow
exp_alerts: []
- eval_time: 11m
alertname: SeaweedVolumeDiskLow
exp_alerts:
- exp_labels:
job: seaweedfs-volume
name: /data
severity: warning
exp_annotations:
summary: SeaweedFS /data 可用空间不足 10%
description: 检查 volume 文件系统容量、保留策略与底层 ZFS;不要直接删除卷文件。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-data-services.md
- name: NATS 历史慢消费者计数不持续报警
interval: 1m
input_series:
- series: nats_varz_slow_consumers{job="nats/nats"}
values: 10x30
alert_rule_test:
- eval_time: 20m
alertname: NatsSlowConsumers
exp_alerts: []
- name: Flux 根对象状态指标缺失报警
interval: 1m
input_series: []
alert_rule_test:
- eval_time: 4m
alertname: FluxStateMetricsMissing
exp_alerts: []
- eval_time: 6m
alertname: FluxStateMetricsMissing
exp_alerts:
- exp_labels:
job: kube-state-metrics
customresource_kind: Kustomization
resource_namespace: flux-system
name: flux-system
severity: warning
exp_annotations:
summary: Flux 对象状态指标缺失
description: 检查 kube-state-metrics 自定义资源配置、RBAC 与采集;不能把空结果当作健康。
runbook_url: https://git.ddupan.top/panxiao81/homelab-wiki/src/branch/main/guides/monitoring-controllers.md
- name: OCI HelmRepository 没有 Ready 条件仍不误报
interval: 1m
input_series:
- series: gotk_resource_info{job="kube-state-metrics",customresource_kind="HelmRepository",repository_type="oci",name="envoy-gateway"}
values: 1x30
alert_rule_test:
- eval_time: 20m
alertname: FluxResourceNotReady
exp_alerts: []
@@ -10,6 +10,11 @@ spec:
replicaCount: 1
selectAllByDefault: true
evaluationInterval: 30s
# 告警 Source 复用已认证的 Grafana 入口,携带表达式和触发时间。
# 整个 Explore JSON 一次性 URL 编码,避免表达式中的引号、&、+ 损坏链接。
extraArgs:
external.url: https://grafana.ad.ddupan.top
external.alert.source: 'explore?left={{ printf "{\"datasource\":\"VictoriaMetrics\",\"queries\":[{\"expr\":%s,\"refId\":\"A\"}],\"range\":{\"from\":\"%d\",\"to\":\"now\"}}" (.Expr | jsonEscape) .ActiveAt.UnixMilli | queryEscape }}'
datasource:
url: http://vmsingle-main.monitoring.svc:8428
remoteWrite:
@@ -1,6 +1,4 @@
# Alert router/notifier. Migrated from the retired Compose stack (see Git history),
# which currently blackholes everything. Wire real receivers here (email via the
# in-cluster smtp-relay, or a webhook) when you want notifications.
# Telegram token 由 ExternalSecret 从 OpenBao 投射,配置中仅引用文件。
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMAlertmanager
metadata:
@@ -8,19 +6,37 @@ metadata:
namespace: monitoring
spec:
replicaCount: 1
secrets:
- alertmanager-telegram
configRawYaml: |
route:
receiver: blackhole
group_by: [alertname, cluster, job, severity]
group_wait: 1m
group_interval: 5m
repeat_interval: 4h
routes:
- receiver: telegram
matchers:
- severity="critical"
group_wait: 10s
repeat_interval: 1h
- receiver: telegram
matchers:
- severity="warning"
inhibit_rules:
- source_matchers:
- alertname="KubePodCrashLooping"
target_matchers:
- alertname="KubePodFrequentRestarts"
equal: [namespace, pod, container]
receivers:
- name: blackhole
# Example email receiver via the in-cluster Postfix relay (smtp-relay/):
# receivers:
# - name: email
# email_configs:
# - to: '[email protected]'
# from: 'Alertmanager <[email protected]>'
# smarthost: 'smtp-relay.smtp-relay.svc.cluster.local:25'
# require_tls: false
- name: telegram
telegram_configs:
- bot_token_file: /etc/vm/secrets/alertmanager-telegram/telegram_bot_token
chat_id: -1003956377923
send_resolved: true
resources:
requests:
cpu: 25m
@@ -4,6 +4,8 @@ metadata:
name: vm-operator
namespace: monitoring
spec:
dependsOn:
- name: prometheus-operator-crds
chart:
spec:
chart: victoria-metrics-operator
+19 -1
View File
@@ -9,6 +9,15 @@ Pod-bound PSAT 向中央 SPIRE 注册;identity controller 从 BatchSandbox all
取得真实 Pod UID,再创建精确的 `ClusterStaticEntry`。runner 只有拿到请求中的完整
repository/task SVID 后才领取一次性 Gitea registration token。
VM 与 Pod backend 使用同一个 `gitea-dynamic-runner-runner` executor 镜像。Pool 不运行
常驻 `docker:dind` sidecar;需要 Docker 的 workflow 在 privileged executor 内按任务启动
daemon。Kata VM 中的 workflow 必须先把稀疏 ext4 镜像 loop-mount 到
`/var/lib/docker`,并完成 cgroup v2 nesting 初始化。
`ci-vm` executor 固定请求并限制为 2 CPU/3GiB。真实 kind canary 表明 1 CPU/约 2GiB
虽然能最终启动全部 control-plane 容器,但无法在 kubeadm 超时前提供可用的 API server;
该规格是 nested Kubernetes 任务的容量下限,不是用资源掩盖存储阻塞。
这些 `ClusterStaticEntry` 位于 sandbox 集群,由 central SPIRE Server 内的
`spire-controller-manager-sandbox` 通过受限 external kubeconfig reconcile。必须在
`platform/spire/values.yaml` 显式启用 external controller-manager 的
@@ -63,4 +72,13 @@ kubectl -n opensandbox logs deploy/opensandbox-identity -f
- `PoolCapacityExhausted`:检查 `ci-vm` 的 `poolMax` 及残留 BatchSandbox;
- runner 等待 SVID:核对 allocation Pod UID、ClusterStaticEntry parentID、guest Agent 日志;
- runner 等待 token:核对 `192.168.10.127:8787` 的 sandbox 到 homelab 路由;
- Docker 任务失败:检查 `docker` sidecar 和 `/run/docker/docker.sock` 的 group 2000。
- Docker 任务失败:检查 workflow 的 job-local dockerd 日志及
`/var/run/docker.sock`;Pool 不提供共享或常驻 daemon。
- kind node 的 systemd 报 `Failed to create /init.scope` 或 `Structure needs
cleaning`:确认 workflow 启动 dockerd 前完成 cgroup v2 nesting 初始化。否则
子容器的私有 cgroup namespace 根会变成 `threaded`,systemd 无法创建 domain
cgroup。
- kind 的 kubeadm 卡在 `CreateContainer`,而 nested containerd goroutine 停在
`bbolt` 的 `fdatasync`:不要直接把 Kata shared mount 用作 Docker 数据目录。
workflow 必须在 guest 内创建稀疏 ext4 镜像并 loop-mount 到
`/var/lib/docker`;该镜像随 Sandbox 删除,不得在节点侧遗留。
+17 -25
View File
@@ -48,7 +48,7 @@ spec:
mountPath: /opt/opensandbox
containers:
- name: sandbox
image: zot.ad.ddupan.top/panxiao81/gitea-dynamic-runner-runner@sha256:a45875fd2d0e67429b0b7bc3914735581669c705bb4e1a05d712134d2bceb86f
image: zot.ad.ddupan.top/panxiao81/gitea-dynamic-runner-runner@sha256:325d75a208b1a1c6f3b4d4705e47bbc58b34733edcd12f689c60769ab6772a59
imagePullPolicy: IfNotPresent
command: [/opt/opensandbox/task-executor]
args:
@@ -66,25 +66,28 @@ spec:
value: https://git.ddupan.top
- name: HOME
value: /data
- name: DOCKER_HOST
value: unix:///run/docker/docker.sock
ports:
- name: task-executor
containerPort: 5758
securityContext:
allowPrivilegeEscalation: false
capabilities:
drop: [ALL]
privileged: true
runAsNonRoot: true
runAsUser: 2000
runAsGroup: 2000
resources:
requests:
cpu: "2"
memory: 3Gi
limits:
cpu: "2"
memory: 3Gi
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: runner-data
mountPath: /data
- name: docker-socket
mountPath: /run/docker
- name: runner-workspace
mountPath: /workspace
- name: spire-socket
mountPath: /run/spire/agent-sockets
- name: spire-agent
@@ -108,28 +111,13 @@ spec:
mountPath: /run/spire/data
- name: spire-socket
mountPath: /run/spire/agent-sockets
- name: docker
image: docker.io/library/docker:29.1.5-dind
command: [/bin/sh, -c]
args:
- test -e /dev/kmsg || mknod /dev/kmsg c 1 11; exec dockerd --host=unix:///run/docker/docker.sock --group=2000 --storage-driver=overlay2
securityContext:
privileged: true
volumeMounts:
- name: docker-socket
mountPath: /run/docker
- name: docker-data
mountPath: /var/lib/docker
volumes:
- name: opensandbox-bin
emptyDir: {}
- name: runner-data
emptyDir: {}
- name: docker-data
- name: runner-workspace
emptyDir: {}
- name: docker-socket
emptyDir:
medium: Memory
- name: spire-data
emptyDir:
medium: Memory
@@ -198,7 +186,7 @@ spec:
mountPath: /opt/opensandbox
containers:
- name: sandbox
image: zot.ad.ddupan.top/panxiao81/gitea-dynamic-runner-runner@sha256:a45875fd2d0e67429b0b7bc3914735581669c705bb4e1a05d712134d2bceb86f
image: zot.ad.ddupan.top/panxiao81/gitea-dynamic-runner-runner@sha256:325d75a208b1a1c6f3b4d4705e47bbc58b34733edcd12f689c60769ab6772a59
imagePullPolicy: IfNotPresent
command: [/opt/opensandbox/task-executor]
args:
@@ -233,6 +221,8 @@ spec:
mountPath: /opt/opensandbox
- name: runner-data
mountPath: /data
- name: runner-workspace
mountPath: /workspace
- name: docker-socket
mountPath: /run/docker
- name: spire-socket
@@ -275,6 +265,8 @@ spec:
emptyDir: {}
- name: runner-data
emptyDir: {}
- name: runner-workspace
emptyDir: {}
- name: docker-data
emptyDir: {}
- name: docker-socket