Author SHA1 Message Date
panxiao81 0e3564bf72 feat: 部署 OpenSandbox API 与双运行时 Pool
yaml / yaml (pull_request) Successful in 20s
2026-09-18 01:10:07 +00:00
panxiao81 c6ec310b0b Merge pull request 切换 SPIRE 到内部 chart fork
yaml / yaml (push) Successful in 19s
合并内部 chart fork 与 sandbox-kata Pod UID PSAT profile。
2026-09-18 00:24:54 +00:00
panxiao81 3b77af8da1 feat: 切换 SPIRE 到内部 chart fork
yaml / yaml (pull_request) Successful in 14s
2026-09-18 00:24:21 +00:00
panxiao81 8af511ecc8 Merge pull request '部署 sandbox Kata Containers' (#92) from feat/sandbox-kata into main
yaml / yaml (push) Successful in 22s
Reviewed-on: #92
2026-09-17 18:06:18 +00:00
panxiao81 94684d0722 feat: 部署 sandbox Kata Containers
yaml / yaml (pull_request) Successful in 17s
2026-09-17 18:02:46 +00:00
27 changed files with 534 additions and 7 deletions
+2 -1
View File
@@ -50,6 +50,7 @@ sudo k3s kubectl -n flux-system get gitrepositories,kustomizations
- VictoriaMetrics Operator 已固定现有 chart `0.66.2` 并完成分阶段 Flux HelmRelease
接管;Metrics、Logs、Traces 与 Grafana 也已统一完成 Flux 接管;
- External Secrets Operator 已固定 chart `2.8.0` 并完成分阶段接管;
- SPIRE 已按官方 hardened chart `0.30.2`(SPIRE `1.15.3`)声明,使用共享
- SPIRE 已按 hardened chart 内部 fork `0.30.2-ddupan.1`(基于上游 `0.30.2`,SPIRE
`1.15.3`)声明,使用共享
PostgreSQL 与独立 signing-key PVC;首次上线和 OpenBao JWT-SVID PoC 尚待合并后验证;
- root Kustomization 与所有 brownfield 子 Kustomization 继续保持 `prune: false`。
+5
View File
@@ -46,6 +46,11 @@ attestation。Server 使用 external bundle publisher 持续维护 sandbox
普通 Pod 与后续 Kata guest 的回归夹具,稳定身份为
`spiffe://ddupan.top/sandbox/smoke`。测试 Pod 临时创建并在验收后删除,身份声明保留。
Kata 阶段使用官方 4.1.0 `kata-deploy` chart 的短生命周期 `job` 模式,逐节点安装并
重启 K3s。只启用 `kata-clh-runtime-rs`,不创建默认 `kata` 别名;该 handler 的
`emptyDir` 固定使用 `block-plain`,为 Docker/BuildKit overlay2 与 kind 提供 guest
内块设备文件系统。详细限制与上线验收见 `platform/sandbox-kata/README.md`。
## 监控边界
这里只管理 sandbox LXC 内的 Kubernetes 监控,不负责 PVE 宿主监控。LXC 与宿主共享
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: kata
namespace: flux-system
spec:
dependsOn:
- name: monitoring-operator
- name: spire-agents
interval: 10m
path: ./platform/sandbox-kata
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 35m
wait: true
@@ -0,0 +1,17 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: opensandbox-pools
namespace: flux-system
spec:
dependsOn:
- name: opensandbox
interval: 10m
path: ./platform/sandbox-opensandbox-pools
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 20m
wait: true
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: opensandbox
namespace: flux-system
spec:
dependsOn:
- name: kata
- name: spire-agents
interval: 10m
path: ./platform/sandbox-opensandbox
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 20m
wait: true
+3
View File
@@ -6,3 +6,6 @@ resources:
- apps/monitoring.yaml
- apps/spire-bootstrap.yaml
- apps/spire-agents.yaml
- apps/kata.yaml
- apps/opensandbox.yaml
- apps/opensandbox-pools.yaml
+48
View File
@@ -0,0 +1,48 @@
# Sandbox Kata Containers
本目录通过 Flux 安装 Kata Containers 4.1.0,只启用 Cloud Hypervisor 的 Rust
runtime,并创建明确命名的 `kata-clh-runtime-rs` RuntimeClass。OpenSandbox 的 VM
Pool 必须显式选择该 RuntimeClass;不创建含义不明确的 `kata` 默认别名。
OCI chart 同时固定到已验证的 4.1.0 artifact digest,升级时必须重新执行本页验收。
安装使用官方 `kata-deploy` chart 的 `job` 模式。每次 install/upgrade 由 dispatcher
逐个节点创建短生命周期、可修改 host 的安装 Job,写入 Kata artifacts 和 K3s containerd
配置并重启对应节点的 K3s;节点恢复 Ready 后才继续下一节点。安装结束后不保留拥有
host 写权限的 DaemonSet。新增节点后必须触发 HelmRelease upgrade,使 dispatcher
重新枚举节点。
`clh-runtime-rs` 固定使用:
```toml
[runtime]
emptydir_mode = "block-plain"
```
因此普通 `emptyDir` 会在 kubelet volume 目录创建稀疏 backing file,并作为块设备
热插拔给 guest。Docker/BuildKit 可以在 guest ext4 上使用原生 overlay2,避免
virtio-fs 作为 overlayfs upperdir 时的限制。Kata 当前不会用
`emptyDir.sizeLimit` 决定该虚拟盘容量;实际容量取决于节点 rootfs。节点磁盘占用、
Pod 删除后的 backing-file/VMM 回收和真实构建基准必须作为上线验收项目。
`kata-monitor` 常驻每个节点,只读访问 K3s containerd socket 与 Kata sandbox 状态,
由 `VMPodScrape` 写入中央 VictoriaMetrics。它不拥有 Kubernetes API 凭据或 host 写
权限。
## 上线验收
Flux reconciliation 完成后至少确认:
1. 两个节点重新回到 Ready,`RuntimeClass/kata-clh-runtime-rs` 存在;
2. Kata Pod 内核与 LXC host 内核不同,且 `/dev/kvm` 可用;
3. Cloud Hypervisor API 的 `vm.info.config.memory.shared` 为 `true`;
4. `spire-smoke` ServiceAccount 的 Kata Pod 可获得
`spiffe://ddupan.top/sandbox/smoke`,错误 ServiceAccount 无法获得身份;
5. block-backed `emptyDir` 上 Docker 使用 `overlay2`,BuildKit 与 kind smoke test
通过;kind 的 dockerd bootstrap 需要先在 guest 内执行
`mknod /dev/kmsg c 1 11`;
6. 删除测试 Pod 后,Cloud Hypervisor 进程、backing file 和临时数据全部回收,且
LXC 没有新增 OOM 事件;
7. 中央 VictoriaMetrics 中两个 `kata-monitor` target 均为 `up=1`。
不要依赖手工修改 `/opt/kata` 或 K3s containerd 配置;任何修复都必须回写 chart
values 并由 Flux reconciliation。
+40
View File
@@ -0,0 +1,40 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: kata-deploy
namespace: kata-system
spec:
chartRef:
kind: OCIRepository
name: kata-deploy
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
postRenderers:
- kustomize:
patches:
- target:
kind: DaemonSet
name: kata-monitor
patch: |
- op: add
path: /spec/template/spec/automountServiceAccountToken
value: false
- op: add
path: /spec/template/spec/containers/0/ports/0/name
value: metrics
releaseName: kata-deploy
targetNamespace: kata-system
timeout: 30m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: kata-deploy-values
+9
View File
@@ -0,0 +1,9 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- repository.yaml
- values.yaml
- helmrelease.yaml
- monitor-scrape.yaml
+17
View File
@@ -0,0 +1,17 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: kata-monitor
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- kata-system
podMetricsEndpoints:
- interval: 30s
port: metrics
selector:
matchLabels:
app.kubernetes.io/instance: kata-deploy
app.kubernetes.io/name: kata-monitor
+5
View File
@@ -0,0 +1,5 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: kata-system
+11
View File
@@ -0,0 +1,11 @@
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: OCIRepository
metadata:
name: kata-deploy
namespace: kata-system
spec:
interval: 1h
ref:
digest: sha256:33f102f6db70083de4fc8238af4439c4245a601bcb6a72e49bc80529098aefc0
url: oci://ghcr.io/kata-containers/kata-deploy-charts/kata-deploy
+44
View File
@@ -0,0 +1,44 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: kata-deploy-values
namespace: kata-system
data:
values.yaml: |
deploymentMode: job
k8sDistribution: k3s
job:
parallelism: 1
snapshotter:
setup: []
shims:
disableAll: true
clh-runtime-rs:
enabled: true
dropIn: |
[runtime]
emptydir_mode = "block-plain"
defaultShim:
amd64: clh-runtime-rs
runtimeClasses:
enabled: true
createDefault: false
node-feature-discovery:
enabled: false
monitor:
enabled: true
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 200m
memory: 192Mi
@@ -0,0 +1,4 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- pools.yaml
@@ -0,0 +1,123 @@
---
apiVersion: sandbox.opensandbox.io/v1alpha1
kind: Pool
metadata:
name: ci-pod
namespace: opensandbox
spec:
template:
metadata:
labels:
ci.ddupan.top/backend: pod
spec:
containers:
- name: sandbox
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/code-interpreter:v1.1.0
command: [/opt/opensandbox/task-executor]
args: [-listen-addr=0.0.0.0:5758, -log-dir=/tmp]
env:
- name: SANDBOX_MAIN_CONTAINER
value: sandbox
- name: EXECD_ENVS
value: /opt/opensandbox/.env
- name: EXECD
value: /opt/opensandbox/execd
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "2"
memory: 4Gi
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: sandbox-storage
mountPath: /var/lib/sandbox
initContainers:
- name: task-executor-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/task-executor:v0.1.0
command: [/bin/sh, -c]
args: [cp /workspace/server /opt/opensandbox/task-executor && chmod 0755 /opt/opensandbox/task-executor]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: execd-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22
command: [/bin/sh, -c]
args: [cp ./execd /opt/opensandbox/execd && cp ./bootstrap.sh /opt/opensandbox/bootstrap.sh && chmod 0755 /opt/opensandbox/execd /opt/opensandbox/bootstrap.sh]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
volumes:
- name: opensandbox-bin
emptyDir: {}
- name: sandbox-storage
emptyDir: {}
capacitySpec:
bufferMax: 1
bufferMin: 0
poolMax: 4
poolMin: 0
---
apiVersion: sandbox.opensandbox.io/v1alpha1
kind: Pool
metadata:
name: ci-vm
namespace: opensandbox
spec:
template:
metadata:
labels:
ci.ddupan.top/backend: vm
spec:
runtimeClassName: kata-clh-runtime-rs
containers:
- name: sandbox
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/code-interpreter:v1.1.0
command: [/opt/opensandbox/task-executor]
args: [-listen-addr=0.0.0.0:5758, -log-dir=/tmp]
env:
- name: SANDBOX_MAIN_CONTAINER
value: sandbox
- name: EXECD_ENVS
value: /opt/opensandbox/.env
- name: EXECD
value: /opt/opensandbox/execd
resources:
requests:
cpu: 250m
memory: 512Mi
limits:
cpu: "4"
memory: 8Gi
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: sandbox-storage
mountPath: /var/lib/sandbox
initContainers:
- name: task-executor-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/task-executor:v0.1.0
command: [/bin/sh, -c]
args: [cp /workspace/server /opt/opensandbox/task-executor && chmod 0755 /opt/opensandbox/task-executor]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: execd-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22
command: [/bin/sh, -c]
args: [cp ./execd /opt/opensandbox/execd && cp ./bootstrap.sh /opt/opensandbox/bootstrap.sh && chmod 0755 /opt/opensandbox/execd /opt/opensandbox/bootstrap.sh]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
volumes:
- name: opensandbox-bin
emptyDir: {}
- name: sandbox-storage
emptyDir: {}
capacitySpec:
bufferMax: 1
bufferMin: 0
poolMax: 2
poolMin: 0
+19
View File
@@ -0,0 +1,19 @@
# OpenSandbox
Flux installs the upstream all-in-one OpenSandbox chart pinned to
`helm/opensandbox/0.2.2` (`8f01e935`). The API is cluster-internal and intentionally runs a
single replica until shared server state and HA behaviour have been validated.
`ci-pod` uses `runc`; `ci-vm` uses the separately managed
`kata-clh-runtime-rs` RuntimeClass. Both Pools start at zero and create capacity
on demand. They currently use the upstream interpreter image to validate the
Lifecycle API and Pool allocation independently of the CI scheduler cutover.
The dynamic runner worker, runner image, guest-local SPIRE Agent and Docker
sidecar are introduced only after this layer is Ready. In particular, do not
mount the host SPIFFE CSI socket into `ci-vm`: Unix sockets do not cross the
Kata VM boundary.
Smoke test both backends through the same API by creating sandboxes with
`extensions.poolRef` set to `ci-pod` and `ci-vm`, then confirm their
BatchSandboxes, Pods and VMMs disappear after deletion.
@@ -0,0 +1,31 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: opensandbox
namespace: opensandbox-system
spec:
chart:
spec:
chart: ./kubernetes/charts/opensandbox
interval: 1h
reconcileStrategy: Revision
sourceRef:
kind: GitRepository
name: opensandbox
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: opensandbox
targetNamespace: opensandbox-system
timeout: 15m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: opensandbox-values
@@ -0,0 +1,7 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespaces.yaml
- repository.yaml
- values.yaml
- helmrelease.yaml
@@ -0,0 +1,10 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: opensandbox-system
---
apiVersion: v1
kind: Namespace
metadata:
name: opensandbox
@@ -0,0 +1,11 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: opensandbox
namespace: opensandbox-system
spec:
interval: 1h
ref:
tag: helm/opensandbox/0.2.2
timeout: 60s
url: https://github.com/alibaba/OpenSandbox.git
+62
View File
@@ -0,0 +1,62 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: opensandbox-values
namespace: opensandbox-system
data:
values.yaml: |
opensandbox-controller:
controller:
logLevel: info
replicaCount: 1
metrics:
enabled: true
secure: false
port: 8080
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
opensandbox-server:
server:
replicaCount: 1
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "1"
memory: 1Gi
configToml: |
[server]
host = "0.0.0.0"
port = 80
api_key = ""
[log]
level = "INFO"
[runtime]
type = "kubernetes"
execd_image = "sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22"
[kubernetes]
kubeconfig_path = ""
namespace = "opensandbox"
informer_enabled = true
informer_resync_seconds = 300
informer_watch_timeout_seconds = 60
snapshot_create_timeout_seconds = 900
workload_provider = "batchsandbox"
batchsandbox_template_file = "/etc/opensandbox/example.batchsandbox-template.yaml"
[egress]
image = "sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/egress:v1.1.6"
mode = "dns+nft"
opensandbox-node-agent:
enabled: false
+8 -1
View File
@@ -11,7 +11,8 @@ Authelia 提供;SPIRE 不替代人类 OIDC,也不承担目标服务的资源
Flux 安装 SPIFFE hardened charts:
- `spire-crds` `0.6.1`;
- `spire` `0.30.2`(SPIRE `1.15.3`);
- 内部 fork 的 `spire` `0.30.2-ddupan.1`(SPIRE `1.15.3`),固定 Git tag
`spire-0.30.2-ddupan.1`;
- SPIRE Server、Agent、Controller Manager、SPIFFE CSI Driver;
- OIDC Discovery Provider。
@@ -24,6 +25,12 @@ API 或 Broker API。Trust domain 是 `ddupan.top`,Kubernetes cluster name 是
controller-manager 与 bundle publisher 使用由 sandbox Ansible bootstrap 的独立、受限
kubeconfig。Sandbox 不运行第二套 Server 或 OIDC Provider。
内部 fork 仅在上游 `spire-0.30.2` 基础上暴露
`use_pod_uid_for_agent_id`。现有 `sandbox` profile 保持 node UID 模式,供 DaemonSet
Agent 使用;独立的 `sandbox-kata` profile 复用同一 kubeconfig,但启用 Pod UID 模式,
供每个 Kata guest 内的临时 Agent 使用。不得把现有 `sandbox` profile 切换为 Pod UID,
否则会改变常驻 Agent 的 parent ID。
## PostgreSQL bootstrap
SPIRE registration datastore 使用共享 CloudNativePG:
+1 -1
View File
@@ -41,7 +41,7 @@ SPIRE Server(trust domain: ddupan.top)
| 项目 | 当前值 |
|---|---|
| SPIRE chart | `0.30.2` |
| SPIRE chart | `0.30.2-ddupan.1`(内部 fork,基于 `0.30.2`) |
| SPIRE | `1.15.3` |
| SPIRE CRDs chart | `0.6.1` |
| trust domain | `ddupan.top` |
+11
View File
@@ -0,0 +1,11 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: spiffe-hardened-fork
namespace: spire-mgmt
spec:
interval: 1h
ref:
tag: spire-0.30.2-ddupan.1
timeout: 60s
url: http://gitea-http.gitea.svc.cluster.local:3000/panxiao81/helm-charts-hardened.git
+4 -4
View File
@@ -6,12 +6,12 @@ metadata:
spec:
chart:
spec:
chart: spire
chart: ./charts/spire
interval: 1h
reconcileStrategy: Revision
sourceRef:
kind: HelmRepository
name: spiffe-hardened
version: 0.30.2
kind: GitRepository
name: spiffe-hardened-fork
dependsOn:
- name: spire-crds
namespace: spire-mgmt
+1
View File
@@ -12,6 +12,7 @@ configMapGenerator:
resources:
- namespaces.yaml
- helmrepository.yaml
- gitrepository-fork.yaml
- helmrelease-crds.yaml
- helmrelease.yaml
- httproute.yaml
+5
View File
@@ -51,6 +51,11 @@ spire-server:
kubeConfigName: sandbox
serviceAccountAllowList:
- spire-system:spire-agent
sandbox-kata:
kubeConfigName: sandbox
serviceAccountAllowList:
- spire-smoke:spire-smoke
usePodUIDForAgentID: true
externalControllerManagers:
enabled: true
clusters: