Author SHA1 Message Date
panxiao81 0e3564bf72 feat: 部署 OpenSandbox API 与双运行时 Pool
yaml / yaml (pull_request) Successful in 20s
2026-09-18 01:10:07 +00:00
panxiao81 c6ec310b0b Merge pull request 切换 SPIRE 到内部 chart fork
yaml / yaml (push) Successful in 19s
合并内部 chart fork 与 sandbox-kata Pod UID PSAT profile。
2026-09-18 00:24:54 +00:00
panxiao81 3b77af8da1 feat: 切换 SPIRE 到内部 chart fork
yaml / yaml (pull_request) Successful in 14s
2026-09-18 00:24:21 +00:00
panxiao81 8af511ecc8 Merge pull request '部署 sandbox Kata Containers' (#92) from feat/sandbox-kata into main
yaml / yaml (push) Successful in 22s
Reviewed-on: #92
2026-09-17 18:06:18 +00:00
panxiao81 94684d0722 feat: 部署 sandbox Kata Containers
yaml / yaml (pull_request) Successful in 17s
2026-09-17 18:02:46 +00:00
panxiao81 9515cde49b Merge pull request '补齐 sandbox PSAT reviewer 权限' (#90) from fix/sandbox-spire-psat-rbac into main
yaml / yaml (push) Successful in 16s
Reviewed-on: #90
2026-09-17 17:47:27 +00:00
panxiao81 accf2d8210 Merge pull request '修复 SPIRE Agent 连接 Server 的端口' (#91) from fix/spire-agent-server-port into main
yaml / yaml (push) Successful in 49s
Reviewed-on: #91
2026-09-17 17:44:50 +00:00
panxiao81 d1ccc99125 修复 SPIRE Agent Server 端口
yaml / yaml (pull_request) Failing after 1m1s
2026-09-17 17:43:39 +00:00
panxiao81 b619f6f681 fix: 补齐 sandbox PSAT reviewer 权限
yaml / yaml (pull_request) Successful in 13s
2026-09-17 17:30:54 +00:00
panxiao81 1806c678a4 Merge pull request '部署 sandbox SPIRE Agent 与 CSI' (#89) from feat/sandbox-spire-agents into main
yaml / yaml (push) Successful in 21s
ansible / collection-test (push) Successful in 1m16s
ansible / lint (push) Successful in 2m17s
Reviewed-on: #89
2026-09-17 17:27:14 +00:00
panxiao81 aeb8c49d0a feat: 部署 sandbox SPIRE Agent 与 CSI
yaml / yaml (pull_request) Successful in 20s
ansible / collection-test (pull_request) Successful in 1m9s
ansible / lint (pull_request) Successful in 2m15s
2026-09-17 17:23:25 +00:00
panxiao81 7b1a98280c Merge pull request '引导 sandbox SPIRE external controller' (#88) from feat/sandbox-spire-controller-bootstrap into main
yaml / yaml (push) Successful in 17s
ansible / collection-test (push) Successful in 1m12s
ansible / lint (push) Successful in 3m3s
Reviewed-on: #88
2026-09-17 17:14:50 +00:00
panxiao81 ea15841d4d feat: 引导 sandbox SPIRE external controller
yaml / yaml (pull_request) Successful in 17s
ansible / collection-test (pull_request) Successful in 1m5s
ansible / lint (pull_request) Successful in 4m4s
2026-09-17 17:00:44 +00:00
panxiao81 4b9aa9e164 Merge pull request '引导 sandbox 跨集群 SPIRE 认证' (#87) from feat/sandbox-spire-bootstrap into main
yaml / yaml (push) Successful in 20s
ansible / collection-test (push) Successful in 1m15s
ansible / lint (push) Successful in 2m12s
Reviewed-on: #87
2026-09-17 16:55:59 +00:00
panxiao81 b7b92b3465 feat: 引导 sandbox 跨集群 SPIRE 认证
yaml / yaml (pull_request) Successful in 19s
ansible / lint (pull_request) Successful in 2m13s
ansible / collection-test (pull_request) Successful in 1m11s
2026-09-17 16:50:02 +00:00
panxiao81 dc2b43f693 Merge pull request '接入 sandbox Kubernetes 监控' (#86) from feat/sandbox-monitoring into main
yaml / yaml (push) Successful in 17s
ansible / collection-test (push) Successful in 1m12s
ansible / lint (push) Successful in 2m11s
Reviewed-on: #86
2026-09-17 16:21:27 +00:00
panxiao81 60836c3360 feat: 接入 sandbox Kubernetes 监控
yaml / yaml (pull_request) Successful in 25s
ansible / collection-test (pull_request) Successful in 1m10s
ansible / lint (pull_request) Successful in 4m33s
2026-09-17 16:02:48 +00:00
61 changed files with 1537 additions and 9 deletions
+2 -1
View File
@@ -50,6 +50,7 @@ sudo k3s kubectl -n flux-system get gitrepositories,kustomizations
- VictoriaMetrics Operator 已固定现有 chart `0.66.2` 并完成分阶段 Flux HelmRelease
接管;Metrics、Logs、Traces 与 Grafana 也已统一完成 Flux 接管;
- External Secrets Operator 已固定 chart `2.8.0` 并完成分阶段接管;
- SPIRE 已按官方 hardened chart `0.30.2`(SPIRE `1.15.3`)声明,使用共享
- SPIRE 已按 hardened chart 内部 fork `0.30.2-ddupan.1`(基于上游 `0.30.2`,SPIRE
`1.15.3`)声明,使用共享
PostgreSQL 与独立 signing-key PVC;首次上线和 OpenBao JWT-SVID PoC 尚待合并后验证;
- root Kustomization 与所有 brownfield 子 Kustomization 继续保持 `prune: false`。
+33 -1
View File
@@ -10,7 +10,7 @@ Ansible 将 homelab CA 注入 `GitRepository/flux-system` 引用的同名 Secret
长期 Git 凭据。root Kustomization 从 `./clusters/sandbox` 开始 reconciliation,
初始保持 `prune: false`。
当前 root 为空,作为 bootstrap canary。后续按依赖顺序分别引入:
Root bootstrap 已完成。后续按依赖顺序分别引入:
1. 监控 CRD、kube-state-metrics 以及 kubelet/cAdvisor 抓取配置;
2. SPIRE Agent、SPIFFE CSI Driver 与 workload registration;
@@ -18,6 +18,38 @@ Ansible 将 homelab CA 注入 `GitRepository/flux-system` 引用的同名 Secret
4. OpenSandbox operator/server 及 `ci-pod`、`ci-vm` Pools。
每一阶段单独合并并等待对应 Flux Kustomization Ready,不在 bootstrap 时一次性部署。
第一阶段监控拆为 `monitoring-operator` 与依赖它的 `monitoring`,防止 VM CR 在
VictoriaMetrics Operator CRD Ready 前进入 reconciliation。
SPIRE 阶段先由 `spire-bootstrap` 安装 CRD,并声明按上游 k8s_psat Server plugin
要求收窄的 reviewer:它可以调用 TokenReview,并只读查询用于证明的 Pod 与 Node。
Agent ServiceAccount 留给后续 HelmRelease 创建,避免两个声明方争夺同一资源。随后运行
`infrastructure/sandbox-cluster/ansible/spire-bootstrap.yml`:playbook 从 sandbox
读取 reviewer token,在内存中组成受限 kubeconfig,再通过 stdin reconcile 到 central
集群的 `spire-server/spire-external-kubeconfigs` Secret。凭据不写入仓库、日志或控制机
文件;该 Secret 准备完成后,才能启用 central external PSAT/controller-manager 和
sandbox Agent/CSI。
External controller-manager 使用独立的 `spire-controller-manager` ServiceAccount;其
RBAC 与上游 controller-manager 所需权限一致,用于读取 workload selectors、维护
SPIFFE CR status/finalizer 和 leader election。它不复用只允许 TokenReview 的 Server
reviewer。Ansible 将两份 kubeconfig 写入同一个 central Secret 的不同 key,便于 central
chart 分别绑定 `sandbox` 与 `sandbox-controller`。
Central SPIRE Server 通过内网 `spire-server.ad.ddupan.top:8081` 接收 sandbox Agent
attestation。Server 使用 external bundle publisher 持续维护 sandbox
`spire-system/spire-bundle`,Agent 不固定或复制 trust bundle。Sandbox HelmRelease
显式关闭 Server 与 OIDC Provider,只部署 Agent DaemonSet 和 SPIFFE CSI Driver;因此
不会产生第二个 trust root。
`spire-smoke` namespace、ServiceAccount 和 `sandbox-spire-smoke` ClusterSPIFFEID 是
普通 Pod 与后续 Kata guest 的回归夹具,稳定身份为
`spiffe://ddupan.top/sandbox/smoke`。测试 Pod 临时创建并在验收后删除,身份声明保留。
Kata 阶段使用官方 4.1.0 `kata-deploy` chart 的短生命周期 `job` 模式,逐节点安装并
重启 K3s。只启用 `kata-clh-runtime-rs`,不创建默认 `kata` 别名;该 handler 的
`emptyDir` 固定使用 `block-plain`,为 Docker/BuildKit overlay2 与 kind 提供 guest
内块设备文件系统。详细限制与上线验收见 `platform/sandbox-kata/README.md`。
## 监控边界
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: kata
namespace: flux-system
spec:
dependsOn:
- name: monitoring-operator
- name: spire-agents
interval: 10m
path: ./platform/sandbox-kata
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 35m
wait: true
@@ -0,0 +1,15 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: monitoring-operator
namespace: flux-system
spec:
interval: 10m
path: ./platform/sandbox-monitoring/operator
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 10m
wait: true
+17
View File
@@ -0,0 +1,17 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: monitoring
namespace: flux-system
spec:
dependsOn:
- name: monitoring-operator
interval: 10m
path: ./platform/sandbox-monitoring/workloads
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 10m
wait: true
@@ -0,0 +1,17 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: opensandbox-pools
namespace: flux-system
spec:
dependsOn:
- name: opensandbox
interval: 10m
path: ./platform/sandbox-opensandbox-pools
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 20m
wait: true
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: opensandbox
namespace: flux-system
spec:
dependsOn:
- name: kata
- name: spire-agents
interval: 10m
path: ./platform/sandbox-opensandbox
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 20m
wait: true
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: spire-agents
namespace: flux-system
spec:
dependsOn:
- name: spire-bootstrap
- name: monitoring-operator
interval: 10m
path: ./platform/sandbox-spire/agents
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 15m
wait: true
@@ -0,0 +1,15 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: spire-bootstrap
namespace: flux-system
spec:
interval: 10m
path: ./platform/sandbox-spire/bootstrap
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 10m
wait: true
+8 -1
View File
@@ -1,4 +1,11 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources: []
resources:
- apps/monitoring-operator.yaml
- apps/monitoring.yaml
- apps/spire-bootstrap.yaml
- apps/spire-agents.yaml
- apps/kata.yaml
- apps/opensandbox.yaml
- apps/opensandbox-pools.yaml
+2
View File
@@ -15,10 +15,12 @@ homelab_dns:
- { zone: ad.ddupan.top, name: sandbox-k8s, type: A, values: [10.60.0.13] }
- { zone: ad.ddupan.top, name: retrolab, type: A, values: [10.60.0.10] }
- { zone: ad.ddupan.top, name: grafana, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: metrics-write, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: netbox, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: nats, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: s3, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-oidc, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-server, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: zot, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: zot-push, type: A, values: [192.168.10.127] }
+22
View File
@@ -108,6 +108,28 @@ homelab CA trust。Flux `v2.9.5` controllers 与 root sync 也由 Ansible 通过
server manifests 管理;root 使用 homelab CA 访问公开 Gitea 仓库,不保存 Git token。
集群内 workload 由 `clusters/sandbox/` 分阶段纳入 Flux。
## SPIRE 跨集群 bootstrap
Sandbox 复用 homelab 的 SPIRE Server 与 `ddupan.top` trust domain。Flux 首先安装
SPIRE CRD,并创建供 k8s_psat 使用的 reviewer。它按上游 Server chart 的权限模型调用
TokenReview,并以 `get/list` 读取用于证明的 Pod 与 Node;不具有修改 workload 的权限。
在该 Kustomization Ready 后运行:
```bash
cd infrastructure/sandbox-cluster/ansible
ansible-playbook spire-bootstrap.yml
```
Playbook 不把 reviewer token 或生成的 kubeconfig 落盘,而是将目标 Secret manifest
通过 stdin 交给本机 homelab `k3s kubectl`。目标 Secret
`spire-server/spire-external-kubeconfigs` 由 Ansible 单独拥有;Flux 和人工操作不得写入。
第二次运行必须为零变更。
Secret 的 `sandbox` key 仅供 SPIRE Server 的 external PSAT plugin 验证 token 与对应的
Pod/Node;
`sandbox-controller` key 供 external controller-manager 读取 Pod/Node、reconcile SPIFFE
CR 及执行 leader election。两者使用不同 ServiceAccount,不得合并权限或互换。
## 已验证的 Kata CI 前置条件
- Cloud Hypervisor 必须报告 `vm.info.config.memory.shared=true`;
@@ -0,0 +1,7 @@
---
sandbox_spire_bootstrap_api_server: https://10.60.0.13:6443
sandbox_spire_bootstrap_source_namespace: spire-system
sandbox_spire_bootstrap_source_secret: spire-server-token-reviewer-token
sandbox_spire_bootstrap_controller_secret: spire-controller-manager-token
sandbox_spire_bootstrap_target_namespace: spire-server
sandbox_spire_bootstrap_target_secret: spire-external-kubeconfigs
@@ -0,0 +1,115 @@
---
- name: Wait for the sandbox SPIRE token reviewer credential
ansible.builtin.command:
argv:
- k3s
- kubectl
- --namespace
- "{{ sandbox_spire_bootstrap_source_namespace }}"
- get
- secret
- "{{ sandbox_spire_bootstrap_source_secret }}"
- --output=json
register: sandbox_spire_bootstrap_reviewer_secret
changed_when: false
retries: 60
delay: 10
until:
- sandbox_spire_bootstrap_reviewer_secret.rc == 0
- (sandbox_spire_bootstrap_reviewer_secret.stdout | from_json).data.token is defined
- (sandbox_spire_bootstrap_reviewer_secret.stdout | from_json).data['ca.crt'] is defined
no_log: true
- name: Wait for the sandbox SPIRE controller credential
ansible.builtin.command:
argv:
- k3s
- kubectl
- --namespace
- "{{ sandbox_spire_bootstrap_source_namespace }}"
- get
- secret
- "{{ sandbox_spire_bootstrap_controller_secret }}"
- --output=json
register: sandbox_spire_bootstrap_controller_secret_result
changed_when: false
retries: 60
delay: 10
until:
- sandbox_spire_bootstrap_controller_secret_result.rc == 0
- (sandbox_spire_bootstrap_controller_secret_result.stdout | from_json).data.token is defined
- (sandbox_spire_bootstrap_controller_secret_result.stdout | from_json).data['ca.crt'] is defined
no_log: true
- name: Extract the sandbox TokenReview credential data
ansible.builtin.set_fact:
sandbox_spire_bootstrap_secret_data: >-
{{ (sandbox_spire_bootstrap_reviewer_secret.stdout | from_json).data }}
sandbox_spire_bootstrap_controller_data: >-
{{ (sandbox_spire_bootstrap_controller_secret_result.stdout | from_json).data }}
no_log: true
- name: Build the restricted sandbox TokenReview kubeconfig
ansible.builtin.set_fact:
sandbox_spire_bootstrap_kubeconfig: |
apiVersion: v1
kind: Config
clusters:
- name: sandbox
cluster:
server: {{ sandbox_spire_bootstrap_api_server }}
certificate-authority-data: {{ sandbox_spire_bootstrap_secret_data['ca.crt'] }}
users:
- name: spire-server-token-reviewer
user:
token: {{ sandbox_spire_bootstrap_secret_data.token | b64decode }}
contexts:
- name: sandbox
context:
cluster: sandbox
user: spire-server-token-reviewer
current-context: sandbox
sandbox_spire_bootstrap_controller_kubeconfig: |
apiVersion: v1
kind: Config
clusters:
- name: sandbox
cluster:
server: {{ sandbox_spire_bootstrap_api_server }}
certificate-authority-data: {{ sandbox_spire_bootstrap_controller_data['ca.crt'] }}
users:
- name: spire-controller-manager
user:
token: {{ sandbox_spire_bootstrap_controller_data.token | b64decode }}
contexts:
- name: sandbox
context:
cluster: sandbox
user: spire-controller-manager
current-context: sandbox
no_log: true
- name: Reconcile the central SPIRE external kubeconfig Secret
ansible.builtin.command:
argv:
- k3s
- kubectl
- apply
- --filename=-
stdin: |
apiVersion: v1
kind: Secret
metadata:
name: {{ sandbox_spire_bootstrap_target_secret }}
namespace: {{ sandbox_spire_bootstrap_target_namespace }}
type: Opaque
data:
sandbox: {{ sandbox_spire_bootstrap_kubeconfig | b64encode }}
sandbox-controller: {{ sandbox_spire_bootstrap_controller_kubeconfig | b64encode }}
delegate_to: localhost
become: true
register: sandbox_spire_bootstrap_target
changed_when: >-
' created' in sandbox_spire_bootstrap_target.stdout or
' configured' in sandbox_spire_bootstrap_target.stdout
no_log: true
@@ -52,3 +52,9 @@
gather_facts: false
roles:
- sandbox_flux
- name: Reconcile central SPIRE access to sandbox Kubernetes
hosts: sandbox1
gather_facts: false
roles:
- sandbox_spire_bootstrap
@@ -0,0 +1,6 @@
---
- name: Reconcile central SPIRE access to sandbox Kubernetes
hosts: sandbox1
gather_facts: false
roles:
- sandbox_spire_bootstrap
@@ -2,6 +2,7 @@ apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- vmsingle.yaml
- vmsingle-write-route.yaml
- vmagent.yaml
- vmalert.yaml
- vmalertmanager.yaml
@@ -0,0 +1,24 @@
---
# Internal-only remote_write ingress for vmagent instances in other homelab
# clusters. Expose only the write endpoint, not VictoriaMetrics query/admin APIs.
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: vmsingle-remote-write
namespace: monitoring
spec:
parentRefs:
- name: eg
namespace: envoy-gateway-system
sectionName: https
hostnames:
- metrics-write.ad.ddupan.top
rules:
- matches:
- method: POST
path:
type: Exact
value: /api/v1/write
backendRefs:
- name: vmsingle-main
port: 8428
+48
View File
@@ -0,0 +1,48 @@
# Sandbox Kata Containers
本目录通过 Flux 安装 Kata Containers 4.1.0,只启用 Cloud Hypervisor 的 Rust
runtime,并创建明确命名的 `kata-clh-runtime-rs` RuntimeClass。OpenSandbox 的 VM
Pool 必须显式选择该 RuntimeClass;不创建含义不明确的 `kata` 默认别名。
OCI chart 同时固定到已验证的 4.1.0 artifact digest,升级时必须重新执行本页验收。
安装使用官方 `kata-deploy` chart 的 `job` 模式。每次 install/upgrade 由 dispatcher
逐个节点创建短生命周期、可修改 host 的安装 Job,写入 Kata artifacts 和 K3s containerd
配置并重启对应节点的 K3s;节点恢复 Ready 后才继续下一节点。安装结束后不保留拥有
host 写权限的 DaemonSet。新增节点后必须触发 HelmRelease upgrade,使 dispatcher
重新枚举节点。
`clh-runtime-rs` 固定使用:
```toml
[runtime]
emptydir_mode = "block-plain"
```
因此普通 `emptyDir` 会在 kubelet volume 目录创建稀疏 backing file,并作为块设备
热插拔给 guest。Docker/BuildKit 可以在 guest ext4 上使用原生 overlay2,避免
virtio-fs 作为 overlayfs upperdir 时的限制。Kata 当前不会用
`emptyDir.sizeLimit` 决定该虚拟盘容量;实际容量取决于节点 rootfs。节点磁盘占用、
Pod 删除后的 backing-file/VMM 回收和真实构建基准必须作为上线验收项目。
`kata-monitor` 常驻每个节点,只读访问 K3s containerd socket 与 Kata sandbox 状态,
由 `VMPodScrape` 写入中央 VictoriaMetrics。它不拥有 Kubernetes API 凭据或 host 写
权限。
## 上线验收
Flux reconciliation 完成后至少确认:
1. 两个节点重新回到 Ready,`RuntimeClass/kata-clh-runtime-rs` 存在;
2. Kata Pod 内核与 LXC host 内核不同,且 `/dev/kvm` 可用;
3. Cloud Hypervisor API 的 `vm.info.config.memory.shared` 为 `true`;
4. `spire-smoke` ServiceAccount 的 Kata Pod 可获得
`spiffe://ddupan.top/sandbox/smoke`,错误 ServiceAccount 无法获得身份;
5. block-backed `emptyDir` 上 Docker 使用 `overlay2`,BuildKit 与 kind smoke test
通过;kind 的 dockerd bootstrap 需要先在 guest 内执行
`mknod /dev/kmsg c 1 11`;
6. 删除测试 Pod 后,Cloud Hypervisor 进程、backing file 和临时数据全部回收,且
LXC 没有新增 OOM 事件;
7. 中央 VictoriaMetrics 中两个 `kata-monitor` target 均为 `up=1`。
不要依赖手工修改 `/opt/kata` 或 K3s containerd 配置;任何修复都必须回写 chart
values 并由 Flux reconciliation。
+40
View File
@@ -0,0 +1,40 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: kata-deploy
namespace: kata-system
spec:
chartRef:
kind: OCIRepository
name: kata-deploy
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
postRenderers:
- kustomize:
patches:
- target:
kind: DaemonSet
name: kata-monitor
patch: |
- op: add
path: /spec/template/spec/automountServiceAccountToken
value: false
- op: add
path: /spec/template/spec/containers/0/ports/0/name
value: metrics
releaseName: kata-deploy
targetNamespace: kata-system
timeout: 30m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: kata-deploy-values
+9
View File
@@ -0,0 +1,9 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- repository.yaml
- values.yaml
- helmrelease.yaml
- monitor-scrape.yaml
+17
View File
@@ -0,0 +1,17 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: kata-monitor
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- kata-system
podMetricsEndpoints:
- interval: 30s
port: metrics
selector:
matchLabels:
app.kubernetes.io/instance: kata-deploy
app.kubernetes.io/name: kata-monitor
+5
View File
@@ -0,0 +1,5 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: kata-system
+11
View File
@@ -0,0 +1,11 @@
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: OCIRepository
metadata:
name: kata-deploy
namespace: kata-system
spec:
interval: 1h
ref:
digest: sha256:33f102f6db70083de4fc8238af4439c4245a601bcb6a72e49bc80529098aefc0
url: oci://ghcr.io/kata-containers/kata-deploy-charts/kata-deploy
+44
View File
@@ -0,0 +1,44 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: kata-deploy-values
namespace: kata-system
data:
values.yaml: |
deploymentMode: job
k8sDistribution: k3s
job:
parallelism: 1
snapshotter:
setup: []
shims:
disableAll: true
clh-runtime-rs:
enabled: true
dropIn: |
[runtime]
emptydir_mode = "block-plain"
defaultShim:
amd64: clh-runtime-rs
runtimeClasses:
enabled: true
createDefault: false
node-feature-discovery:
enabled: false
monitor:
enabled: true
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 200m
memory: 192Mi
+16
View File
@@ -0,0 +1,16 @@
# Sandbox 监控
该目录只采集 sandbox Kubernetes 与 workload 指标:
- kubelet 与 cAdvisor;
- kube-state-metrics;
- vmagent 自身运行状态;
- SPIRE Agent attestation、SVID 与连接状态。
不得在 sandbox LXC 内部署 node_exporter。LXC 的 `/proc/stat` 暴露 PVE 宿主 CPU
视图,会产生重复且语义混合的指标。PVE 宿主监控不属于本目录。
vmagent 为所有远端样本增加 `cluster=sandbox`,并通过仅允许
`POST /api/v1/write` 的 `metrics-write.ad.ddupan.top` 路由写入 homelab VMSingle。
Operator 与 workload 拆成两个 Flux Kustomization,确保 VM CRD Ready 后再创建
VMAgent、VMNodeScrape 和 VMServiceScrape。
@@ -0,0 +1,32 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: vm-operator
namespace: monitoring
spec:
chart:
spec:
chart: victoria-metrics-operator
interval: 1h
sourceRef:
kind: HelmRepository
name: vm
version: 0.66.2
install:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: vm-operator
targetNamespace: monitoring
timeout: 10m
upgrade:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: vm-operator-values
@@ -0,0 +1,8 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- repositories.yaml
- values.yaml
- helmrelease.yaml
@@ -0,0 +1,5 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: monitoring
@@ -0,0 +1,18 @@
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: vm
namespace: monitoring
spec:
interval: 1h
url: https://victoriametrics.github.io/helm-charts/
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: prometheus-community
namespace: monitoring
spec:
interval: 1h
url: https://prometheus-community.github.io/helm-charts
@@ -0,0 +1,23 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: vm-operator-values
namespace: monitoring
data:
values.yaml: |
crds:
plain: true
cleanup:
enabled: false
operator:
disable_prometheus_converter: true
serviceMonitor:
enabled: false
resources:
requests:
cpu: 25m
memory: 96Mi
limits:
cpu: 250m
memory: 256Mi
@@ -0,0 +1,14 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMServiceScrape
metadata:
name: kube-state-metrics
namespace: monitoring
spec:
endpoints:
- interval: 30s
port: http
selector:
matchLabels:
app.kubernetes.io/instance: kube-state-metrics
app.kubernetes.io/name: kube-state-metrics
@@ -0,0 +1,39 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: kube-state-metrics
namespace: monitoring
spec:
chart:
spec:
chart: kube-state-metrics
interval: 1h
sourceRef:
kind: HelmRepository
name: prometheus-community
namespace: monitoring
version: 8.3.0
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: kube-state-metrics
targetNamespace: monitoring
timeout: 10m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
values:
prometheus:
monitor:
enabled: false
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 250m
memory: 256Mi
@@ -0,0 +1,41 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMNodeScrape
metadata:
name: kubelet
namespace: monitoring
spec:
bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token
honorLabels: true
honorTimestamps: false
interval: 30s
scheme: https
tlsConfig:
caFile: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
insecureSkipVerify: true
relabelConfigs:
- action: labelmap
regex: __meta_kubernetes_node_label_(.+)
- targetLabel: job
replacement: kubelet
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMNodeScrape
metadata:
name: cadvisor
namespace: monitoring
spec:
bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token
honorLabels: true
honorTimestamps: false
interval: 30s
path: /metrics/cadvisor
scheme: https
tlsConfig:
caFile: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
insecureSkipVerify: true
relabelConfigs:
- action: labelmap
regex: __meta_kubernetes_node_label_(.+)
- targetLabel: job
replacement: cadvisor
@@ -0,0 +1,9 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- vmagent.yaml
- kubelet-scrapes.yaml
- kube-state-metrics.yaml
- kube-state-metrics-scrape.yaml
- spire-agent-scrape.yaml
@@ -0,0 +1,17 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: spire-agent
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- spire-system
podMetricsEndpoints:
- interval: 30s
port: prom
selector:
matchLabels:
app.kubernetes.io/instance: sandbox-spire
app.kubernetes.io/name: agent
@@ -0,0 +1,21 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMAgent
metadata:
name: sandbox
namespace: monitoring
spec:
externalLabels:
cluster: sandbox
remoteWrite:
- url: https://metrics-write.ad.ddupan.top/api/v1/write
replicaCount: 1
resources:
requests:
cpu: 50m
memory: 128Mi
limits:
cpu: 500m
memory: 512Mi
scrapeInterval: 30s
selectAllByDefault: true
@@ -0,0 +1,4 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- pools.yaml
@@ -0,0 +1,123 @@
---
apiVersion: sandbox.opensandbox.io/v1alpha1
kind: Pool
metadata:
name: ci-pod
namespace: opensandbox
spec:
template:
metadata:
labels:
ci.ddupan.top/backend: pod
spec:
containers:
- name: sandbox
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/code-interpreter:v1.1.0
command: [/opt/opensandbox/task-executor]
args: [-listen-addr=0.0.0.0:5758, -log-dir=/tmp]
env:
- name: SANDBOX_MAIN_CONTAINER
value: sandbox
- name: EXECD_ENVS
value: /opt/opensandbox/.env
- name: EXECD
value: /opt/opensandbox/execd
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "2"
memory: 4Gi
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: sandbox-storage
mountPath: /var/lib/sandbox
initContainers:
- name: task-executor-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/task-executor:v0.1.0
command: [/bin/sh, -c]
args: [cp /workspace/server /opt/opensandbox/task-executor && chmod 0755 /opt/opensandbox/task-executor]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: execd-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22
command: [/bin/sh, -c]
args: [cp ./execd /opt/opensandbox/execd && cp ./bootstrap.sh /opt/opensandbox/bootstrap.sh && chmod 0755 /opt/opensandbox/execd /opt/opensandbox/bootstrap.sh]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
volumes:
- name: opensandbox-bin
emptyDir: {}
- name: sandbox-storage
emptyDir: {}
capacitySpec:
bufferMax: 1
bufferMin: 0
poolMax: 4
poolMin: 0
---
apiVersion: sandbox.opensandbox.io/v1alpha1
kind: Pool
metadata:
name: ci-vm
namespace: opensandbox
spec:
template:
metadata:
labels:
ci.ddupan.top/backend: vm
spec:
runtimeClassName: kata-clh-runtime-rs
containers:
- name: sandbox
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/code-interpreter:v1.1.0
command: [/opt/opensandbox/task-executor]
args: [-listen-addr=0.0.0.0:5758, -log-dir=/tmp]
env:
- name: SANDBOX_MAIN_CONTAINER
value: sandbox
- name: EXECD_ENVS
value: /opt/opensandbox/.env
- name: EXECD
value: /opt/opensandbox/execd
resources:
requests:
cpu: 250m
memory: 512Mi
limits:
cpu: "4"
memory: 8Gi
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: sandbox-storage
mountPath: /var/lib/sandbox
initContainers:
- name: task-executor-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/task-executor:v0.1.0
command: [/bin/sh, -c]
args: [cp /workspace/server /opt/opensandbox/task-executor && chmod 0755 /opt/opensandbox/task-executor]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: execd-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22
command: [/bin/sh, -c]
args: [cp ./execd /opt/opensandbox/execd && cp ./bootstrap.sh /opt/opensandbox/bootstrap.sh && chmod 0755 /opt/opensandbox/execd /opt/opensandbox/bootstrap.sh]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
volumes:
- name: opensandbox-bin
emptyDir: {}
- name: sandbox-storage
emptyDir: {}
capacitySpec:
bufferMax: 1
bufferMin: 0
poolMax: 2
poolMin: 0
+19
View File
@@ -0,0 +1,19 @@
# OpenSandbox
Flux installs the upstream all-in-one OpenSandbox chart pinned to
`helm/opensandbox/0.2.2` (`8f01e935`). The API is cluster-internal and intentionally runs a
single replica until shared server state and HA behaviour have been validated.
`ci-pod` uses `runc`; `ci-vm` uses the separately managed
`kata-clh-runtime-rs` RuntimeClass. Both Pools start at zero and create capacity
on demand. They currently use the upstream interpreter image to validate the
Lifecycle API and Pool allocation independently of the CI scheduler cutover.
The dynamic runner worker, runner image, guest-local SPIRE Agent and Docker
sidecar are introduced only after this layer is Ready. In particular, do not
mount the host SPIFFE CSI socket into `ci-vm`: Unix sockets do not cross the
Kata VM boundary.
Smoke test both backends through the same API by creating sandboxes with
`extensions.poolRef` set to `ci-pod` and `ci-vm`, then confirm their
BatchSandboxes, Pods and VMMs disappear after deletion.
@@ -0,0 +1,31 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: opensandbox
namespace: opensandbox-system
spec:
chart:
spec:
chart: ./kubernetes/charts/opensandbox
interval: 1h
reconcileStrategy: Revision
sourceRef:
kind: GitRepository
name: opensandbox
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: opensandbox
targetNamespace: opensandbox-system
timeout: 15m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: opensandbox-values
@@ -0,0 +1,7 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespaces.yaml
- repository.yaml
- values.yaml
- helmrelease.yaml
@@ -0,0 +1,10 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: opensandbox-system
---
apiVersion: v1
kind: Namespace
metadata:
name: opensandbox
@@ -0,0 +1,11 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: opensandbox
namespace: opensandbox-system
spec:
interval: 1h
ref:
tag: helm/opensandbox/0.2.2
timeout: 60s
url: https://github.com/alibaba/OpenSandbox.git
+62
View File
@@ -0,0 +1,62 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: opensandbox-values
namespace: opensandbox-system
data:
values.yaml: |
opensandbox-controller:
controller:
logLevel: info
replicaCount: 1
metrics:
enabled: true
secure: false
port: 8080
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
opensandbox-server:
server:
replicaCount: 1
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "1"
memory: 1Gi
configToml: |
[server]
host = "0.0.0.0"
port = 80
api_key = ""
[log]
level = "INFO"
[runtime]
type = "kubernetes"
execd_image = "sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22"
[kubernetes]
kubeconfig_path = ""
namespace = "opensandbox"
informer_enabled = true
informer_resync_seconds = 300
informer_watch_timeout_seconds = 60
snapshot_create_timeout_seconds = 900
workload_provider = "batchsandbox"
batchsandbox_template_file = "/etc/opensandbox/example.batchsandbox-template.yaml"
[egress]
image = "sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/egress:v1.1.6"
mode = "dns+nft"
opensandbox-node-agent:
enabled: false
@@ -0,0 +1,7 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- values.yaml
- release.yaml
- smoke-identity.yaml
@@ -0,0 +1,35 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: sandbox-spire
namespace: spire-mgmt
spec:
chart:
spec:
chart: spire
interval: 1h
sourceRef:
kind: HelmRepository
name: spiffe-hardened
version: 0.30.2
dependsOn:
- name: spire-crds
namespace: spire-mgmt
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: sandbox-spire
targetNamespace: spire-mgmt
timeout: 15m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: sandbox-spire-values
@@ -0,0 +1,28 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: spire-smoke
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: spire-smoke
namespace: spire-smoke
---
apiVersion: spire.spiffe.io/v1alpha1
kind: ClusterSPIFFEID
metadata:
name: sandbox-spire-smoke
spec:
className: spire-mgmt-spire
namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: spire-smoke
podSelector:
matchLabels:
app.kubernetes.io/name: spire-smoke
spiffeIDTemplate: spiffe://{{ .TrustDomain }}/sandbox/smoke
workloadSelectorTemplates:
- k8s:ns:spire-smoke
- k8s:sa:spire-smoke
+86
View File
@@ -0,0 +1,86 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: sandbox-spire-values
namespace: spire-mgmt
data:
values.yaml: |
global:
k8s:
clusterDomain: cluster.local
spire:
bundleConfigMap: spire-bundle
clusterName: sandbox
trustDomain: ddupan.top
namespaces:
create: false
system:
name: spire-system
server:
name: spire-server
recommendations:
enabled: true
namespaceLayout: true
namespacePSS: true
priorityClassName: true
strictMode: true
securityContexts: true
prometheus: false
spire-server:
enabled: false
spire-agent:
enabled: true
serviceAccount:
name: spire-agent
server:
address: spire-server.ad.ddupan.top
port: 8081
nodeAttestor:
k8sPSAT:
enabled: true
workloadAttestors:
k8s:
enabled: true
unix:
enabled: false
telemetry:
prometheus:
enabled: true
podMonitor:
enabled: false
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 250m
memory: 192Mi
spiffe-csi-driver:
enabled: true
resources:
requests:
cpu: 10m
memory: 32Mi
limits:
cpu: 100m
memory: 96Mi
spiffe-oidc-discovery-provider:
enabled: false
upstream:
enabled: false
tornjak-frontend:
enabled: false
spire-identity-exchange:
enabled: false
spike-keeper:
enabled: false
spike-nexus:
enabled: false
spike-pilot:
enabled: false
@@ -0,0 +1,162 @@
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: spire-controller-manager
namespace: spire-system
---
apiVersion: v1
kind: Secret
metadata:
name: spire-controller-manager-token
namespace: spire-system
annotations:
kubernetes.io/service-account.name: spire-controller-manager
type: kubernetes.io/service-account-token
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: spire-controller-manager
rules:
- apiGroups:
- ""
resources:
- endpoints
- namespaces
- nodes
- pods
verbs:
- get
- list
- watch
- apiGroups:
- spire.spiffe.io
resources:
- clusterfederatedtrustdomains
- clusterspiffeids
- clusterstaticentries
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- spire.spiffe.io
resources:
- clusterfederatedtrustdomains/finalizers
- clusterspiffeids/finalizers
- clusterstaticentries/finalizers
verbs:
- update
- apiGroups:
- spire.spiffe.io
resources:
- clusterfederatedtrustdomains/status
- clusterspiffeids/status
- clusterstaticentries/status
verbs:
- get
- patch
- update
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: spire-controller-manager
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: spire-controller-manager
subjects:
- kind: ServiceAccount
name: spire-controller-manager
namespace: spire-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: spire-bundle-publisher
namespace: spire-system
rules:
- apiGroups:
- ""
resources:
- configmaps
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: spire-bundle-publisher
namespace: spire-system
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: spire-bundle-publisher
subjects:
- kind: ServiceAccount
name: spire-controller-manager
namespace: spire-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: spire-controller-manager-leader-election
namespace: spire-server
rules:
- apiGroups:
- ""
resources:
- configmaps
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- coordination.k8s.io
resources:
- leases
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
- apiGroups:
- ""
resources:
- events
verbs:
- create
- patch
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: spire-controller-manager-leader-election
namespace: spire-server
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: spire-controller-manager-leader-election
subjects:
- kind: ServiceAccount
name: spire-controller-manager
namespace: spire-system
@@ -0,0 +1,31 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: spire-crds
namespace: spire-mgmt
spec:
chart:
spec:
chart: spire-crds
interval: 1h
sourceRef:
kind: HelmRepository
name: spiffe-hardened
version: 0.6.1
driftDetection:
mode: enabled
install:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: spire-crds
targetNamespace: spire-mgmt
timeout: 10m
upgrade:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
@@ -0,0 +1,9 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespaces.yaml
- repository.yaml
- crds.yaml
- token-reviewer.yaml
- controller-manager.yaml
@@ -0,0 +1,15 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: spire-mgmt
---
apiVersion: v1
kind: Namespace
metadata:
name: spire-system
---
apiVersion: v1
kind: Namespace
metadata:
name: spire-server
@@ -0,0 +1,9 @@
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: spiffe-hardened
namespace: spire-mgmt
spec:
interval: 1h
url: https://spiffe.github.io/helm-charts-hardened/
@@ -0,0 +1,51 @@
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: spire-server-token-reviewer
namespace: spire-system
---
apiVersion: v1
kind: Secret
metadata:
name: spire-server-token-reviewer-token
namespace: spire-system
annotations:
kubernetes.io/service-account.name: spire-server-token-reviewer
type: kubernetes.io/service-account-token
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: spire-server-token-reviewer
rules:
- apiGroups:
- authentication.k8s.io
resources:
- tokenreviews
verbs:
- get
- list
- watch
- create
- apiGroups:
- ""
resources:
- nodes
- pods
verbs:
- get
- list
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: spire-server-token-reviewer
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: spire-server-token-reviewer
subjects:
- kind: ServiceAccount
name: spire-server-token-reviewer
namespace: spire-system
+13 -1
View File
@@ -11,7 +11,8 @@ Authelia 提供;SPIRE 不替代人类 OIDC,也不承担目标服务的资源
Flux 安装 SPIFFE hardened charts:
- `spire-crds` `0.6.1`;
- `spire` `0.30.2`(SPIRE `1.15.3`);
- 内部 fork 的 `spire` `0.30.2-ddupan.1`(SPIRE `1.15.3`),固定 Git tag
`spire-0.30.2-ddupan.1`;
- SPIRE Server、Agent、Controller Manager、SPIFFE CSI Driver;
- OIDC Discovery Provider。
@@ -19,6 +20,17 @@ Flux 安装 SPIFFE hardened charts:
API 或 Broker API。Trust domain 是 `ddupan.top`,Kubernetes cluster name 是
`homelab`。
同一 Server 也接受 cluster name 为 `sandbox` 的 external PSAT attestation。SPIRE gRPC
只通过内网 `spire-server.ad.ddupan.top:8081` 暴露;external PSAT、external
controller-manager 与 bundle publisher 使用由 sandbox Ansible bootstrap 的独立、受限
kubeconfig。Sandbox 不运行第二套 Server 或 OIDC Provider。
内部 fork 仅在上游 `spire-0.30.2` 基础上暴露
`use_pod_uid_for_agent_id`。现有 `sandbox` profile 保持 node UID 模式,供 DaemonSet
Agent 使用;独立的 `sandbox-kata` profile 复用同一 kubeconfig,但启用 Pod UID 模式,
供每个 Kata guest 内的临时 Agent 使用。不得把现有 `sandbox` profile 切换为 Pod UID,
否则会改变常驻 Agent 的 parent ID。
## PostgreSQL bootstrap
SPIRE registration datastore 使用共享 CloudNativePG:
+1 -1
View File
@@ -41,7 +41,7 @@ SPIRE Server(trust domain: ddupan.top)
| 项目 | 当前值 |
|---|---|
| SPIRE chart | `0.30.2` |
| SPIRE chart | `0.30.2-ddupan.1`(内部 fork,基于 `0.30.2`) |
| SPIRE | `1.15.3` |
| SPIRE CRDs chart | `0.6.1` |
| trust domain | `ddupan.top` |
+11
View File
@@ -0,0 +1,11 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: spiffe-hardened-fork
namespace: spire-mgmt
spec:
interval: 1h
ref:
tag: spire-0.30.2-ddupan.1
timeout: 60s
url: http://gitea-http.gitea.svc.cluster.local:3000/panxiao81/helm-charts-hardened.git
+4 -4
View File
@@ -6,12 +6,12 @@ metadata:
spec:
chart:
spec:
chart: spire
chart: ./charts/spire
interval: 1h
reconcileStrategy: Revision
sourceRef:
kind: HelmRepository
name: spiffe-hardened
version: 0.30.2
kind: GitRepository
name: spiffe-hardened-fork
dependsOn:
- name: spire-crds
namespace: spire-mgmt
+1
View File
@@ -12,6 +12,7 @@ configMapGenerator:
resources:
- namespaces.yaml
- helmrepository.yaml
- gitrepository-fork.yaml
- helmrelease-crds.yaml
- helmrelease.yaml
- httproute.yaml
+46
View File
@@ -30,6 +30,47 @@ spire-server:
kind: statefulset
replicaCount: 1
auditLogEnabled: true
service:
type: LoadBalancer
port: 8081
loadBalancerIP: 192.168.10.127
kubeConfigs:
sandbox:
externalSecret:
name: spire-external-kubeconfigs
key: sandbox
sandbox-controller:
externalSecret:
name: spire-external-kubeconfigs
key: sandbox-controller
nodeAttestor:
externalK8sPSAT:
enabled: true
clusters:
sandbox:
kubeConfigName: sandbox
serviceAccountAllowList:
- spire-system:spire-agent
sandbox-kata:
kubeConfigName: sandbox
serviceAccountAllowList:
- spire-smoke:spire-smoke
usePodUIDForAgentID: true
externalControllerManagers:
enabled: true
clusters:
sandbox:
kubeConfigName: sandbox-controller
bundlePublisher:
externalK8sConfigMap:
enabled: true
clusters:
sandbox:
kubeConfigName: sandbox-controller
namespace: spire-system
configMapName: spire-bundle
configMapKey: bundle.spiffe
format: spiffe
persistence:
# PostgreSQL stores registrations, but the disk KeyManager still needs durable
# storage for the trust-domain signing keys.
@@ -65,6 +106,11 @@ spire-server:
enabled: false
spire-agent:
server:
# Keep the Agent endpoint aligned with spire-server.service.port. The
# chart defaults this to 443, which only remained unnoticed while the
# Agent's pre-upgrade gRPC connection stayed alive.
port: 8081
nodeAttestor:
k8sPSAT:
enabled: true