Author SHA1 Message Date
panxiao81 90ba945d85 Merge pull request 部署 OpenSandbox 控制面
yaml / yaml (push) Successful in 18s
2026-09-18 16:38:13 +00:00
panxiao81 d526fd75d3 部署 OpenSandbox 控制面
yaml / yaml (pull_request) Successful in 17s
2026-09-18 16:36:21 +00:00
panxiao81 18cb2858b9 Merge pull request '延长 sandbox Flux 根同步超时' (#93) from fix/sandbox-flux-root-timeout into main
ansible / collection-test (push) Successful in 1m13s
ansible / lint (push) Successful in 2m23s
Reviewed-on: #93
2026-09-18 16:18:46 +00:00
panxiao81 99d1ec1d6f Merge pull request '记录 Kata guest 内 SPIFFE 身份方案' (#95) from docs/kata-inner-spire into main
Reviewed-on: #95
2026-09-18 16:18:29 +00:00
panxiao81 583dab526a docs: 明确 SPIRE chart 字段缺口 2026-09-17 18:54:00 +00:00
panxiao81 3dbd4c5f31 docs: 记录 Kata guest SPIFFE 身份方案 2026-09-17 18:47:01 +00:00
panxiao81 e67bce5121 fix: 延长 sandbox Flux 根同步超时
ansible / collection-test (pull_request) Successful in 1m8s
ansible / lint (pull_request) Successful in 2m56s
2026-09-17 18:20:33 +00:00
14 changed files with 287 additions and 9 deletions
+12 -4
View File
@@ -15,7 +15,7 @@ Root bootstrap 已完成。后续按依赖顺序分别引入:
1. 监控 CRD、kube-state-metrics 以及 kubelet/cAdvisor 抓取配置;
2. SPIRE Agent、SPIFFE CSI Driver 与 workload registration;
3. Kata Containers、`block-plain` RuntimeClass;
4. OpenSandbox operator/server 及 `ci-pod`、`ci-vm` Pools。
4. OpenSandbox controller/server;CI Pool 与 runner 调度器随后独立接入。
每一阶段单独合并并等待对应 Flux Kustomization Ready,不在 bootstrap 时一次性部署。
第一阶段监控拆为 `monitoring-operator` 与依赖它的 `monitoring`,防止 VM CR 在
@@ -42,15 +42,23 @@ attestation。Server 使用 external bundle publisher 持续维护 sandbox
显式关闭 Server 与 OIDC Provider,只部署 Agent DaemonSet 和 SPIFFE CSI Driver;因此
不会产生第二个 trust root。
`spire-smoke` namespace、ServiceAccount 和 `sandbox-spire-smoke` ClusterSPIFFEID 是
普通 Pod 与后续 Kata guest 的回归夹具,稳定身份为
`spiffe://ddupan.top/sandbox/smoke`。测试 Pod 临时创建并在验收后删除,身份声明保留。
`spire-smoke` namespace、ServiceAccount 和 `sandbox-spire-smoke` ClusterSPIFFEID 只用于
普通 Pod 的 CSI 回归夹具,稳定身份为 `spiffe://ddupan.top/sandbox/smoke`。Kata guest
不能复用 node Agent 暴露的 Unix socket;virtio-fs 只能呈现 socket 路径,不能把连接
跨过 VM 边界。Kata workload 必须使用 guest 内 Agent,具体约束见
`platform/sandbox-kata/README.md`。测试 Pod 临时创建并在验收后删除,普通 Pod 的身份
声明保留。
Kata 阶段使用官方 4.1.0 `kata-deploy` chart 的短生命周期 `job` 模式,逐节点安装并
重启 K3s。只启用 `kata-clh-runtime-rs`,不创建默认 `kata` 别名;该 handler 的
`emptyDir` 固定使用 `block-plain`,为 Docker/BuildKit overlay2 与 kind 提供 guest
内块设备文件系统。详细限制与上线验收见 `platform/sandbox-kata/README.md`。
OpenSandbox 阶段固定官方源码 commit 与 umbrella chart `0.2.2`,只部署 controller、
ClusterIP server 和 CRD。server 当前仅能从集群内部访问;在 CI 调度器接入并建立
API key 的 Secret 生命周期前,显式运行于无认证 bootstrap 模式。这个临时边界和切换
步骤记录在 `platform/sandbox-opensandbox/README.md`。
## 监控边界
这里只管理 sandbox LXC 内的 Kubernetes 监控,不负责 PVE 宿主监控。LXC 与宿主共享
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: opensandbox
namespace: flux-system
spec:
dependsOn:
- name: kata
- name: monitoring-operator
interval: 10m
path: ./platform/sandbox-opensandbox
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 15m
wait: true
+1
View File
@@ -7,3 +7,4 @@ resources:
- apps/spire-bootstrap.yaml
- apps/spire-agents.yaml
- apps/kata.yaml
- apps/opensandbox.yaml
+4 -2
View File
@@ -18,7 +18,7 @@ Flux 管理以下 Kubernetes 资源:
- Kata Containers 和 CI 专用的 `block-plain` RuntimeClass;
- SPIRE Agent、SPIFFE CSI Driver 与 workload identity 声明;
- vmagent、kube-state-metrics、kubelet/cAdvisor scrape 配置和告警;
- OpenSandbox operator/server、`ci-pod` 与 `ci-vm` Pools。
- OpenSandbox controller/server;CI Pool 与 runner 调度器由 runner 项目接入。
同一个对象只能有一个 owner。Ansible 不直接部署上述集群内 workload;Flux 不管理
LXC、K3s datastore 或 K3s 本身。
@@ -106,7 +106,9 @@ ansible-playbook verify.yml
当前已经声明 LXC 生命周期、最小 OS baseline、PostgreSQL 和 K3s,包括系统级
homelab CA trust。Flux `v2.9.5` controllers 与 root sync 也由 Ansible 通过 K3s
server manifests 管理;root 使用 homelab CA 访问公开 Gitea 仓库,不保存 Git token。
集群内 workload 由 `clusters/sandbox/` 分阶段纳入 Flux。
集群内 workload 由 `clusters/sandbox/` 分阶段纳入 Flux。root Kustomization 的健康检查
timeout 为 40 分钟,用于覆盖 Kata 等首次安装时会逐节点重启 K3s 的子
Kustomization;各子项仍保留自己的更短 timeout,故障会在对应子项先行暴露。
## SPIRE 跨集群 bootstrap
@@ -35,5 +35,5 @@ spec:
sourceRef:
kind: GitRepository
name: flux-system
timeout: 3m
timeout: 40m
wait: true
+28 -2
View File
@@ -28,6 +28,32 @@ Pod 删除后的 backing-file/VMM 回收和真实构建基准必须作为上线
由 `VMPodScrape` 写入中央 VictoriaMetrics。它不拥有 Kubernetes API 凭据或 host 写
权限。
## Guest 内 SPIFFE 身份
Kata guest 不能直接使用 node SPIRE Agent 的 CSI socket。Unix socket 的路径即使通过
virtio-fs 出现在 guest 中,连接也不能跨 VM 边界。CI Pod 应以 native sidecar 在同一
guest 内启动临时 SPIRE Agent,并满足以下约束:
1. 外层 Pod 挂载 `audience=spire-server` 的 Pod-bound projected ServiceAccount token;
2. 内层 Agent 使用 `k8s_psat` 向中央 Server attestation,Server cluster profile 必须
启用 `use_pod_uid_for_agent_id`,使 Agent ID 包含 Pod UID;
3. 调度器为该具体 Agent 创建 registration entry;SPIFFE ID 使用仓库和任务名等稳定的
业务语义,Pod UID 只作为 parent binding,不进入业务身份;
4. Agent 与 workload 通过 guest 内 `emptyDir` 上的 Unix socket 通信;Pod 必须设置
`shareProcessNamespace: true`,否则 `unix` workload attestor 无法从共享 `/proc`
解析客户端的 `SO_PEERCRED` PID;
5. 调度器必须在 registration entry 已同步到 Agent 后才放行 workload。Pod 删除后同步
删除 entry,并 evict 对应的临时 Agent 记录。
2026-09-17 的 live PoC 已验证:以 Pod UID 作为 parent 的临时 Agent 成功注册,
`unix:uid:2000` workload 从 guest-local socket 获得
`spiffe://ddupan.top/ci/poc/task/build` X.509-SVID,Pod 正常退出。PoC 资源随后全部清理,
中央 SPIRE 恢复声明式配置。
SPIRE 1.15.3 已支持 `use_pod_uid_for_agent_id`,但当前使用的 hardened chart 0.30.2
尚未把它暴露到 values/template。正式部署应先给 chart 补齐该字段并向上游提交,随后
采用包含修复的 release;不得把手工修改 Server ConfigMap 作为运行方案。
## 上线验收
Flux reconciliation 完成后至少确认:
@@ -35,8 +61,8 @@ Flux reconciliation 完成后至少确认:
1. 两个节点重新回到 Ready,`RuntimeClass/kata-clh-runtime-rs` 存在;
2. Kata Pod 内核与 LXC host 内核不同,且 `/dev/kvm` 可用;
3. Cloud Hypervisor API 的 `vm.info.config.memory.shared` 为 `true`;
4. `spire-smoke` ServiceAccount 的 Kata Pod 可获得
`spiffe://ddupan.top/sandbox/smoke`,错误 ServiceAccount 无法获得身份;
4. Kata Pod 内层 Agent 的 ID 包含该 Pod UID;只有调度器创建的业务 entry 所匹配的
workload UID 能从 guest-local socket 获得预期 SPIFFE ID;
5. block-backed `emptyDir` 上 Docker 使用 `overlay2`,BuildKit 与 kind smoke test
通过;kind 的 dockerd bootstrap 需要先在 guest 内执行
`mknod /dev/kmsg c 1 11`;
+44
View File
@@ -0,0 +1,44 @@
# Sandbox OpenSandbox
本目录在独立 sandbox k3s 集群部署 OpenSandbox controller、server 与 CRD。Flux 从
上游 commit `8f01e935c2cabba778cf37a152033fae062fa0f4` 构建官方 umbrella chart
`0.2.2`;该源码渲染结果已与 release `opensandbox-0.2.2.tgz` 对比一致。不要改为跟随
浮动 branch 或 tag。
server 只提供集群内 `opensandbox-server.opensandbox-system.svc:80` ClusterIP,不部署
Gateway、Ingress 或 LoadBalancer。sandbox workload 位于 `opensandbox` namespace,
默认使用 `kata-clh-runtime-rs`;CI Pool、runner 镜像、动态 SPIFFE registration 均由
runner 项目后续声明,本目录不预制。
## 临时认证边界
当前尚无 sandbox 集群内的 Secret 分发链路。为避免把长期凭据提交到 Git,server 暂时
通过 `OPENSANDBOX_INSECURE_SERVER=YES` 显式确认无认证模式;其网络边界严格保持为
ClusterIP。这不是最终认证方案。
runner 接入前必须先创建 `opensandbox-api-key` Secret,并把 Helm values 中的环境变量
改为:
```yaml
- name: OPENSANDBOX_SERVER_API_KEY
valueFrom:
secretKeyRef:
name: opensandbox-api-key
key: api-key
```
随后删除 `OPENSANDBOX_INSECURE_SERVER`。Secret 必须由 OpenBao/SPIFFE 派生的自动化
链路或 Ansible 注入,不得把明文写入仓库。
## 验收
合并后等待 `flux-system/opensandbox` 与 `opensandbox-system/opensandbox` Ready,并确认:
```bash
kubectl get crd batchsandboxes.sandbox.opensandbox.io pools.sandbox.opensandbox.io
kubectl -n opensandbox-system get deploy,pod,svc
kubectl get runtimeclass kata-clh-runtime-rs
```
控制面上线不创建 CI Pool,也不产生 sandbox workload。首个 runner 集成应另行提交 Pool
与完整的 Lifecycle API smoke test。
@@ -0,0 +1,35 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: opensandbox
namespace: opensandbox-system
spec:
chart:
spec:
chart: ./kubernetes/charts/opensandbox
interval: 1h
reconcileStrategy: Revision
sourceRef:
kind: GitRepository
name: opensandbox
driftDetection:
mode: enabled
install:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: opensandbox
targetNamespace: opensandbox-system
timeout: 15m
upgrade:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: opensandbox-values
valuesKey: values.yaml
@@ -0,0 +1,10 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- repository.yaml
- template.yaml
- values.yaml
- helmrelease.yaml
- monitor-scrape.yaml
@@ -0,0 +1,16 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: opensandbox-controller
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- opensandbox-system
podMetricsEndpoints:
- interval: 30s
port: metrics
selector:
matchLabels:
app.kubernetes.io/name: opensandbox
@@ -0,0 +1,10 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: opensandbox-system
---
apiVersion: v1
kind: Namespace
metadata:
name: opensandbox
@@ -0,0 +1,11 @@
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: opensandbox
namespace: opensandbox-system
spec:
interval: 1h
ref:
commit: 8f01e935c2cabba778cf37a152033fae062fa0f4
url: https://github.com/opensandbox-group/OpenSandbox.git
@@ -0,0 +1,17 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: opensandbox-batchsandbox-template
namespace: opensandbox-system
data:
batchsandbox-template.yaml: |
metadata:
labels:
ddupan.top/workload-class: sandbox
spec:
replicas: 1
template:
spec:
restartPolicy: Never
terminationGracePeriodSeconds: 30
+80
View File
@@ -0,0 +1,80 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: opensandbox-values
namespace: opensandbox-system
data:
values.yaml: |
opensandbox-controller:
controller:
metrics:
enabled: true
port: 8080
secure: false
opensandbox-server:
server:
replicaCount: 1
env:
# 临时 bootstrap 边界:Service 仅为 ClusterIP。runner 接入时改为
# secretKeyRef(OPENSANDBOX_SERVER_API_KEY) 并删除本项。
- name: OPENSANDBOX_INSECURE_SERVER
value: "YES"
resources:
limits:
cpu: "1"
memory: 1Gi
requests:
cpu: 100m
memory: 256Mi
volumeMounts:
- name: batchsandbox-template
mountPath: /etc/opensandbox/batchsandbox-template.yaml
subPath: batchsandbox-template.yaml
readOnly: true
volumes:
- name: batchsandbox-template
configMap:
name: opensandbox-batchsandbox-template
configToml: |
[server]
host = "0.0.0.0"
port = 80
api_key = ""
max_sandbox_timeout_seconds = 86400
[log]
level = "INFO"
[runtime]
type = "kubernetes"
execd_image = "sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22"
[storage]
allowed_host_paths = []
volume_default_size = "1Gi"
[kubernetes]
kubeconfig_path = ""
namespace = "opensandbox"
informer_enabled = true
informer_resync_seconds = 300
informer_watch_timeout_seconds = 60
snapshot_create_timeout_seconds = 900
workload_provider = "batchsandbox"
image_pull_policy = "IfNotPresent"
batchsandbox_template_file = "/etc/opensandbox/batchsandbox-template.yaml"
[egress]
image = "sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/egress:v1.1.6"
mode = "dns+nft"
disable_ipv6 = true
[secure_runtime]
type = "kata"
k8s_runtime_class = "kata-clh-runtime-rs"
opensandbox-node-agent:
enabled: false