Author SHA1 Message Date
panxiao81 0e3564bf72 feat: 部署 OpenSandbox API 与双运行时 Pool
yaml / yaml (pull_request) Successful in 20s
2026-09-18 01:10:07 +00:00
panxiao81 c6ec310b0b Merge pull request 切换 SPIRE 到内部 chart fork
yaml / yaml (push) Successful in 19s
合并内部 chart fork 与 sandbox-kata Pod UID PSAT profile。
2026-09-18 00:24:54 +00:00
panxiao81 3b77af8da1 feat: 切换 SPIRE 到内部 chart fork
yaml / yaml (pull_request) Successful in 14s
2026-09-18 00:24:21 +00:00
panxiao81 8af511ecc8 Merge pull request '部署 sandbox Kata Containers' (#92) from feat/sandbox-kata into main
yaml / yaml (push) Successful in 22s
Reviewed-on: #92
2026-09-17 18:06:18 +00:00
panxiao81 94684d0722 feat: 部署 sandbox Kata Containers
yaml / yaml (pull_request) Successful in 17s
2026-09-17 18:02:46 +00:00
panxiao81 9515cde49b Merge pull request '补齐 sandbox PSAT reviewer 权限' (#90) from fix/sandbox-spire-psat-rbac into main
yaml / yaml (push) Successful in 16s
Reviewed-on: #90
2026-09-17 17:47:27 +00:00
panxiao81 accf2d8210 Merge pull request '修复 SPIRE Agent 连接 Server 的端口' (#91) from fix/spire-agent-server-port into main
yaml / yaml (push) Successful in 49s
Reviewed-on: #91
2026-09-17 17:44:50 +00:00
panxiao81 d1ccc99125 修复 SPIRE Agent Server 端口
yaml / yaml (pull_request) Failing after 1m1s
2026-09-17 17:43:39 +00:00
panxiao81 b619f6f681 fix: 补齐 sandbox PSAT reviewer 权限
yaml / yaml (pull_request) Successful in 13s
2026-09-17 17:30:54 +00:00
panxiao81 1806c678a4 Merge pull request '部署 sandbox SPIRE Agent 与 CSI' (#89) from feat/sandbox-spire-agents into main
yaml / yaml (push) Successful in 21s
ansible / collection-test (push) Successful in 1m16s
ansible / lint (push) Successful in 2m17s
Reviewed-on: #89
2026-09-17 17:27:14 +00:00
panxiao81 aeb8c49d0a feat: 部署 sandbox SPIRE Agent 与 CSI
yaml / yaml (pull_request) Successful in 20s
ansible / collection-test (pull_request) Successful in 1m9s
ansible / lint (pull_request) Successful in 2m15s
2026-09-17 17:23:25 +00:00
39 changed files with 837 additions and 13 deletions
+2 -1
View File
@@ -50,6 +50,7 @@ sudo k3s kubectl -n flux-system get gitrepositories,kustomizations
- VictoriaMetrics Operator 已固定现有 chart `0.66.2` 并完成分阶段 Flux HelmRelease - VictoriaMetrics Operator 已固定现有 chart `0.66.2` 并完成分阶段 Flux HelmRelease
接管;Metrics、Logs、Traces 与 Grafana 也已统一完成 Flux 接管; 接管;Metrics、Logs、Traces 与 Grafana 也已统一完成 Flux 接管;
- External Secrets Operator 已固定 chart `2.8.0` 并完成分阶段接管; - External Secrets Operator 已固定 chart `2.8.0` 并完成分阶段接管;
- SPIRE 已按官方 hardened chart `0.30.2`(SPIRE `1.15.3`)声明,使用共享 - SPIRE 已按 hardened chart 内部 fork `0.30.2-ddupan.1`(基于上游 `0.30.2`,SPIRE
`1.15.3`)声明,使用共享
PostgreSQL 与独立 signing-key PVC;首次上线和 OpenBao JWT-SVID PoC 尚待合并后验证; PostgreSQL 与独立 signing-key PVC;首次上线和 OpenBao JWT-SVID PoC 尚待合并后验证;
- root Kustomization 与所有 brownfield 子 Kustomization 继续保持 `prune: false`。 - root Kustomization 与所有 brownfield 子 Kustomization 继续保持 `prune: false`。
+18 -3
View File
@@ -21,9 +21,9 @@ Root bootstrap 已完成。后续按依赖顺序分别引入:
第一阶段监控拆为 `monitoring-operator` 与依赖它的 `monitoring`,防止 VM CR 在 第一阶段监控拆为 `monitoring-operator` 与依赖它的 `monitoring`,防止 VM CR 在
VictoriaMetrics Operator CRD Ready 前进入 reconciliation。 VictoriaMetrics Operator CRD Ready 前进入 reconciliation。
SPIRE 阶段先由 `spire-bootstrap` 安装 CRD,并声明只允许 `tokenreviews.create` 的 SPIRE 阶段先由 `spire-bootstrap` 安装 CRD,并声明按上游 k8s_psat Server plugin
central Server reviewer。Agent ServiceAccount 留给后续 HelmRelease 创建,避免两个 要求收窄的 reviewer:它可以调用 TokenReview,并只读查询用于证明的 Pod 与 Node。
声明方争夺同一资源。随后运行 Agent ServiceAccount 留给后续 HelmRelease 创建,避免两个声明方争夺同一资源。随后运行
`infrastructure/sandbox-cluster/ansible/spire-bootstrap.yml`:playbook 从 sandbox `infrastructure/sandbox-cluster/ansible/spire-bootstrap.yml`:playbook 从 sandbox
读取 reviewer token,在内存中组成受限 kubeconfig,再通过 stdin reconcile 到 central 读取 reviewer token,在内存中组成受限 kubeconfig,再通过 stdin reconcile 到 central
集群的 `spire-server/spire-external-kubeconfigs` Secret。凭据不写入仓库、日志或控制机 集群的 `spire-server/spire-external-kubeconfigs` Secret。凭据不写入仓库、日志或控制机
@@ -36,6 +36,21 @@ SPIFFE CR status/finalizer 和 leader election。它不复用只允许 TokenRevi
reviewer。Ansible 将两份 kubeconfig 写入同一个 central Secret 的不同 key,便于 central reviewer。Ansible 将两份 kubeconfig 写入同一个 central Secret 的不同 key,便于 central
chart 分别绑定 `sandbox` 与 `sandbox-controller`。 chart 分别绑定 `sandbox` 与 `sandbox-controller`。
Central SPIRE Server 通过内网 `spire-server.ad.ddupan.top:8081` 接收 sandbox Agent
attestation。Server 使用 external bundle publisher 持续维护 sandbox
`spire-system/spire-bundle`,Agent 不固定或复制 trust bundle。Sandbox HelmRelease
显式关闭 Server 与 OIDC Provider,只部署 Agent DaemonSet 和 SPIFFE CSI Driver;因此
不会产生第二个 trust root。
`spire-smoke` namespace、ServiceAccount 和 `sandbox-spire-smoke` ClusterSPIFFEID 是
普通 Pod 与后续 Kata guest 的回归夹具,稳定身份为
`spiffe://ddupan.top/sandbox/smoke`。测试 Pod 临时创建并在验收后删除,身份声明保留。
Kata 阶段使用官方 4.1.0 `kata-deploy` chart 的短生命周期 `job` 模式,逐节点安装并
重启 K3s。只启用 `kata-clh-runtime-rs`,不创建默认 `kata` 别名;该 handler 的
`emptyDir` 固定使用 `block-plain`,为 Docker/BuildKit overlay2 与 kind 提供 guest
内块设备文件系统。详细限制与上线验收见 `platform/sandbox-kata/README.md`。
## 监控边界 ## 监控边界
这里只管理 sandbox LXC 内的 Kubernetes 监控,不负责 PVE 宿主监控。LXC 与宿主共享 这里只管理 sandbox LXC 内的 Kubernetes 监控,不负责 PVE 宿主监控。LXC 与宿主共享
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: kata
namespace: flux-system
spec:
dependsOn:
- name: monitoring-operator
- name: spire-agents
interval: 10m
path: ./platform/sandbox-kata
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 35m
wait: true
@@ -0,0 +1,17 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: opensandbox-pools
namespace: flux-system
spec:
dependsOn:
- name: opensandbox
interval: 10m
path: ./platform/sandbox-opensandbox-pools
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 20m
wait: true
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: opensandbox
namespace: flux-system
spec:
dependsOn:
- name: kata
- name: spire-agents
interval: 10m
path: ./platform/sandbox-opensandbox
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 20m
wait: true
+18
View File
@@ -0,0 +1,18 @@
---
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: spire-agents
namespace: flux-system
spec:
dependsOn:
- name: spire-bootstrap
- name: monitoring-operator
interval: 10m
path: ./platform/sandbox-spire/agents
prune: true
sourceRef:
kind: GitRepository
name: flux-system
timeout: 15m
wait: true
+4
View File
@@ -5,3 +5,7 @@ resources:
- apps/monitoring-operator.yaml - apps/monitoring-operator.yaml
- apps/monitoring.yaml - apps/monitoring.yaml
- apps/spire-bootstrap.yaml - apps/spire-bootstrap.yaml
- apps/spire-agents.yaml
- apps/kata.yaml
- apps/opensandbox.yaml
- apps/opensandbox-pools.yaml
+1
View File
@@ -20,6 +20,7 @@ homelab_dns:
- { zone: ad.ddupan.top, name: nats, type: A, values: [192.168.10.127] } - { zone: ad.ddupan.top, name: nats, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: s3, type: A, values: [192.168.10.127] } - { zone: ad.ddupan.top, name: s3, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-oidc, type: A, values: [192.168.10.127] } - { zone: ad.ddupan.top, name: spire-oidc, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-server, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: zot, type: A, values: [192.168.10.127] } - { zone: ad.ddupan.top, name: zot, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: zot-push, type: A, values: [192.168.10.127] } - { zone: ad.ddupan.top, name: zot-push, type: A, values: [192.168.10.127] }
+4 -2
View File
@@ -111,7 +111,8 @@ server manifests 管理;root 使用 homelab CA 访问公开 Gitea 仓库,不
## SPIRE 跨集群 bootstrap ## SPIRE 跨集群 bootstrap
Sandbox 复用 homelab 的 SPIRE Server 与 `ddupan.top` trust domain。Flux 首先安装 Sandbox 复用 homelab 的 SPIRE Server 与 `ddupan.top` trust domain。Flux 首先安装
SPIRE CRD,并创建权限仅为 `authentication.k8s.io/tokenreviews.create` 的 reviewer。 SPIRE CRD,并创建供 k8s_psat 使用的 reviewer。它按上游 Server chart 的权限模型调用
TokenReview,并以 `get/list` 读取用于证明的 Pod 与 Node;不具有修改 workload 的权限。
在该 Kustomization Ready 后运行: 在该 Kustomization Ready 后运行:
```bash ```bash
@@ -124,7 +125,8 @@ Playbook 不把 reviewer token 或生成的 kubeconfig 落盘,而是将目标
`spire-server/spire-external-kubeconfigs` 由 Ansible 单独拥有;Flux 和人工操作不得写入。 `spire-server/spire-external-kubeconfigs` 由 Ansible 单独拥有;Flux 和人工操作不得写入。
第二次运行必须为零变更。 第二次运行必须为零变更。
Secret 的 `sandbox` key 仅供 SPIRE Server 的 external PSAT plugin 执行 TokenReview; Secret 的 `sandbox` key 仅供 SPIRE Server 的 external PSAT plugin 验证 token 与对应的
Pod/Node;
`sandbox-controller` key 供 external controller-manager 读取 Pod/Node、reconcile SPIFFE `sandbox-controller` key 供 external controller-manager 读取 Pod/Node、reconcile SPIFFE
CR 及执行 leader election。两者使用不同 ServiceAccount,不得合并权限或互换。 CR 及执行 leader election。两者使用不同 ServiceAccount,不得合并权限或互换。
+48
View File
@@ -0,0 +1,48 @@
# Sandbox Kata Containers
本目录通过 Flux 安装 Kata Containers 4.1.0,只启用 Cloud Hypervisor 的 Rust
runtime,并创建明确命名的 `kata-clh-runtime-rs` RuntimeClass。OpenSandbox 的 VM
Pool 必须显式选择该 RuntimeClass;不创建含义不明确的 `kata` 默认别名。
OCI chart 同时固定到已验证的 4.1.0 artifact digest,升级时必须重新执行本页验收。
安装使用官方 `kata-deploy` chart 的 `job` 模式。每次 install/upgrade 由 dispatcher
逐个节点创建短生命周期、可修改 host 的安装 Job,写入 Kata artifacts 和 K3s containerd
配置并重启对应节点的 K3s;节点恢复 Ready 后才继续下一节点。安装结束后不保留拥有
host 写权限的 DaemonSet。新增节点后必须触发 HelmRelease upgrade,使 dispatcher
重新枚举节点。
`clh-runtime-rs` 固定使用:
```toml
[runtime]
emptydir_mode = "block-plain"
```
因此普通 `emptyDir` 会在 kubelet volume 目录创建稀疏 backing file,并作为块设备
热插拔给 guest。Docker/BuildKit 可以在 guest ext4 上使用原生 overlay2,避免
virtio-fs 作为 overlayfs upperdir 时的限制。Kata 当前不会用
`emptyDir.sizeLimit` 决定该虚拟盘容量;实际容量取决于节点 rootfs。节点磁盘占用、
Pod 删除后的 backing-file/VMM 回收和真实构建基准必须作为上线验收项目。
`kata-monitor` 常驻每个节点,只读访问 K3s containerd socket 与 Kata sandbox 状态,
由 `VMPodScrape` 写入中央 VictoriaMetrics。它不拥有 Kubernetes API 凭据或 host 写
权限。
## 上线验收
Flux reconciliation 完成后至少确认:
1. 两个节点重新回到 Ready,`RuntimeClass/kata-clh-runtime-rs` 存在;
2. Kata Pod 内核与 LXC host 内核不同,且 `/dev/kvm` 可用;
3. Cloud Hypervisor API 的 `vm.info.config.memory.shared` 为 `true`;
4. `spire-smoke` ServiceAccount 的 Kata Pod 可获得
`spiffe://ddupan.top/sandbox/smoke`,错误 ServiceAccount 无法获得身份;
5. block-backed `emptyDir` 上 Docker 使用 `overlay2`,BuildKit 与 kind smoke test
通过;kind 的 dockerd bootstrap 需要先在 guest 内执行
`mknod /dev/kmsg c 1 11`;
6. 删除测试 Pod 后,Cloud Hypervisor 进程、backing file 和临时数据全部回收,且
LXC 没有新增 OOM 事件;
7. 中央 VictoriaMetrics 中两个 `kata-monitor` target 均为 `up=1`。
不要依赖手工修改 `/opt/kata` 或 K3s containerd 配置;任何修复都必须回写 chart
values 并由 Flux reconciliation。
+40
View File
@@ -0,0 +1,40 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: kata-deploy
namespace: kata-system
spec:
chartRef:
kind: OCIRepository
name: kata-deploy
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
postRenderers:
- kustomize:
patches:
- target:
kind: DaemonSet
name: kata-monitor
patch: |
- op: add
path: /spec/template/spec/automountServiceAccountToken
value: false
- op: add
path: /spec/template/spec/containers/0/ports/0/name
value: metrics
releaseName: kata-deploy
targetNamespace: kata-system
timeout: 30m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: kata-deploy-values
+9
View File
@@ -0,0 +1,9 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespace.yaml
- repository.yaml
- values.yaml
- helmrelease.yaml
- monitor-scrape.yaml
+17
View File
@@ -0,0 +1,17 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: kata-monitor
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- kata-system
podMetricsEndpoints:
- interval: 30s
port: metrics
selector:
matchLabels:
app.kubernetes.io/instance: kata-deploy
app.kubernetes.io/name: kata-monitor
+5
View File
@@ -0,0 +1,5 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: kata-system
+11
View File
@@ -0,0 +1,11 @@
---
apiVersion: source.toolkit.fluxcd.io/v1
kind: OCIRepository
metadata:
name: kata-deploy
namespace: kata-system
spec:
interval: 1h
ref:
digest: sha256:33f102f6db70083de4fc8238af4439c4245a601bcb6a72e49bc80529098aefc0
url: oci://ghcr.io/kata-containers/kata-deploy-charts/kata-deploy
+44
View File
@@ -0,0 +1,44 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: kata-deploy-values
namespace: kata-system
data:
values.yaml: |
deploymentMode: job
k8sDistribution: k3s
job:
parallelism: 1
snapshotter:
setup: []
shims:
disableAll: true
clh-runtime-rs:
enabled: true
dropIn: |
[runtime]
emptydir_mode = "block-plain"
defaultShim:
amd64: clh-runtime-rs
runtimeClasses:
enabled: true
createDefault: false
node-feature-discovery:
enabled: false
monitor:
enabled: true
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 200m
memory: 192Mi
+2 -1
View File
@@ -4,7 +4,8 @@
- kubelet 与 cAdvisor; - kubelet 与 cAdvisor;
- kube-state-metrics; - kube-state-metrics;
- vmagent 自身运行状态。 - vmagent 自身运行状态;
- SPIRE Agent attestation、SVID 与连接状态。
不得在 sandbox LXC 内部署 node_exporter。LXC 的 `/proc/stat` 暴露 PVE 宿主 CPU 不得在 sandbox LXC 内部署 node_exporter。LXC 的 `/proc/stat` 暴露 PVE 宿主 CPU
视图,会产生重复且语义混合的指标。PVE 宿主监控不属于本目录。 视图,会产生重复且语义混合的指标。PVE 宿主监控不属于本目录。
@@ -6,3 +6,4 @@ resources:
- kubelet-scrapes.yaml - kubelet-scrapes.yaml
- kube-state-metrics.yaml - kube-state-metrics.yaml
- kube-state-metrics-scrape.yaml - kube-state-metrics-scrape.yaml
- spire-agent-scrape.yaml
@@ -0,0 +1,17 @@
---
apiVersion: operator.victoriametrics.com/v1beta1
kind: VMPodScrape
metadata:
name: spire-agent
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- spire-system
podMetricsEndpoints:
- interval: 30s
port: prom
selector:
matchLabels:
app.kubernetes.io/instance: sandbox-spire
app.kubernetes.io/name: agent
@@ -0,0 +1,4 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- pools.yaml
@@ -0,0 +1,123 @@
---
apiVersion: sandbox.opensandbox.io/v1alpha1
kind: Pool
metadata:
name: ci-pod
namespace: opensandbox
spec:
template:
metadata:
labels:
ci.ddupan.top/backend: pod
spec:
containers:
- name: sandbox
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/code-interpreter:v1.1.0
command: [/opt/opensandbox/task-executor]
args: [-listen-addr=0.0.0.0:5758, -log-dir=/tmp]
env:
- name: SANDBOX_MAIN_CONTAINER
value: sandbox
- name: EXECD_ENVS
value: /opt/opensandbox/.env
- name: EXECD
value: /opt/opensandbox/execd
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "2"
memory: 4Gi
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: sandbox-storage
mountPath: /var/lib/sandbox
initContainers:
- name: task-executor-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/task-executor:v0.1.0
command: [/bin/sh, -c]
args: [cp /workspace/server /opt/opensandbox/task-executor && chmod 0755 /opt/opensandbox/task-executor]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: execd-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22
command: [/bin/sh, -c]
args: [cp ./execd /opt/opensandbox/execd && cp ./bootstrap.sh /opt/opensandbox/bootstrap.sh && chmod 0755 /opt/opensandbox/execd /opt/opensandbox/bootstrap.sh]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
volumes:
- name: opensandbox-bin
emptyDir: {}
- name: sandbox-storage
emptyDir: {}
capacitySpec:
bufferMax: 1
bufferMin: 0
poolMax: 4
poolMin: 0
---
apiVersion: sandbox.opensandbox.io/v1alpha1
kind: Pool
metadata:
name: ci-vm
namespace: opensandbox
spec:
template:
metadata:
labels:
ci.ddupan.top/backend: vm
spec:
runtimeClassName: kata-clh-runtime-rs
containers:
- name: sandbox
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/code-interpreter:v1.1.0
command: [/opt/opensandbox/task-executor]
args: [-listen-addr=0.0.0.0:5758, -log-dir=/tmp]
env:
- name: SANDBOX_MAIN_CONTAINER
value: sandbox
- name: EXECD_ENVS
value: /opt/opensandbox/.env
- name: EXECD
value: /opt/opensandbox/execd
resources:
requests:
cpu: 250m
memory: 512Mi
limits:
cpu: "4"
memory: 8Gi
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: sandbox-storage
mountPath: /var/lib/sandbox
initContainers:
- name: task-executor-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/task-executor:v0.1.0
command: [/bin/sh, -c]
args: [cp /workspace/server /opt/opensandbox/task-executor && chmod 0755 /opt/opensandbox/task-executor]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
- name: execd-installer
image: sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22
command: [/bin/sh, -c]
args: [cp ./execd /opt/opensandbox/execd && cp ./bootstrap.sh /opt/opensandbox/bootstrap.sh && chmod 0755 /opt/opensandbox/execd /opt/opensandbox/bootstrap.sh]
volumeMounts:
- name: opensandbox-bin
mountPath: /opt/opensandbox
volumes:
- name: opensandbox-bin
emptyDir: {}
- name: sandbox-storage
emptyDir: {}
capacitySpec:
bufferMax: 1
bufferMin: 0
poolMax: 2
poolMin: 0
+19
View File
@@ -0,0 +1,19 @@
# OpenSandbox
Flux installs the upstream all-in-one OpenSandbox chart pinned to
`helm/opensandbox/0.2.2` (`8f01e935`). The API is cluster-internal and intentionally runs a
single replica until shared server state and HA behaviour have been validated.
`ci-pod` uses `runc`; `ci-vm` uses the separately managed
`kata-clh-runtime-rs` RuntimeClass. Both Pools start at zero and create capacity
on demand. They currently use the upstream interpreter image to validate the
Lifecycle API and Pool allocation independently of the CI scheduler cutover.
The dynamic runner worker, runner image, guest-local SPIRE Agent and Docker
sidecar are introduced only after this layer is Ready. In particular, do not
mount the host SPIFFE CSI socket into `ci-vm`: Unix sockets do not cross the
Kata VM boundary.
Smoke test both backends through the same API by creating sandboxes with
`extensions.poolRef` set to `ci-pod` and `ci-vm`, then confirm their
BatchSandboxes, Pods and VMMs disappear after deletion.
@@ -0,0 +1,31 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: opensandbox
namespace: opensandbox-system
spec:
chart:
spec:
chart: ./kubernetes/charts/opensandbox
interval: 1h
reconcileStrategy: Revision
sourceRef:
kind: GitRepository
name: opensandbox
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: opensandbox
targetNamespace: opensandbox-system
timeout: 15m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: opensandbox-values
@@ -0,0 +1,7 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- namespaces.yaml
- repository.yaml
- values.yaml
- helmrelease.yaml
@@ -0,0 +1,10 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: opensandbox-system
---
apiVersion: v1
kind: Namespace
metadata:
name: opensandbox
@@ -0,0 +1,11 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: opensandbox
namespace: opensandbox-system
spec:
interval: 1h
ref:
tag: helm/opensandbox/0.2.2
timeout: 60s
url: https://github.com/alibaba/OpenSandbox.git
+62
View File
@@ -0,0 +1,62 @@
apiVersion: v1
kind: ConfigMap
metadata:
name: opensandbox-values
namespace: opensandbox-system
data:
values.yaml: |
opensandbox-controller:
controller:
logLevel: info
replicaCount: 1
metrics:
enabled: true
secure: false
port: 8080
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 500m
memory: 256Mi
opensandbox-server:
server:
replicaCount: 1
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: "1"
memory: 1Gi
configToml: |
[server]
host = "0.0.0.0"
port = 80
api_key = ""
[log]
level = "INFO"
[runtime]
type = "kubernetes"
execd_image = "sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/execd:v1.0.22"
[kubernetes]
kubeconfig_path = ""
namespace = "opensandbox"
informer_enabled = true
informer_resync_seconds = 300
informer_watch_timeout_seconds = 60
snapshot_create_timeout_seconds = 900
workload_provider = "batchsandbox"
batchsandbox_template_file = "/etc/opensandbox/example.batchsandbox-template.yaml"
[egress]
image = "sandbox-registry.cn-zhangjiakou.cr.aliyuncs.com/opensandbox/egress:v1.1.6"
mode = "dns+nft"
opensandbox-node-agent:
enabled: false
@@ -0,0 +1,7 @@
---
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- values.yaml
- release.yaml
- smoke-identity.yaml
@@ -0,0 +1,35 @@
---
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: sandbox-spire
namespace: spire-mgmt
spec:
chart:
spec:
chart: spire
interval: 1h
sourceRef:
kind: HelmRepository
name: spiffe-hardened
version: 0.30.2
dependsOn:
- name: spire-crds
namespace: spire-mgmt
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: sandbox-spire
targetNamespace: spire-mgmt
timeout: 15m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: sandbox-spire-values
@@ -0,0 +1,28 @@
---
apiVersion: v1
kind: Namespace
metadata:
name: spire-smoke
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: spire-smoke
namespace: spire-smoke
---
apiVersion: spire.spiffe.io/v1alpha1
kind: ClusterSPIFFEID
metadata:
name: sandbox-spire-smoke
spec:
className: spire-mgmt-spire
namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: spire-smoke
podSelector:
matchLabels:
app.kubernetes.io/name: spire-smoke
spiffeIDTemplate: spiffe://{{ .TrustDomain }}/sandbox/smoke
workloadSelectorTemplates:
- k8s:ns:spire-smoke
- k8s:sa:spire-smoke
+86
View File
@@ -0,0 +1,86 @@
---
apiVersion: v1
kind: ConfigMap
metadata:
name: sandbox-spire-values
namespace: spire-mgmt
data:
values.yaml: |
global:
k8s:
clusterDomain: cluster.local
spire:
bundleConfigMap: spire-bundle
clusterName: sandbox
trustDomain: ddupan.top
namespaces:
create: false
system:
name: spire-system
server:
name: spire-server
recommendations:
enabled: true
namespaceLayout: true
namespacePSS: true
priorityClassName: true
strictMode: true
securityContexts: true
prometheus: false
spire-server:
enabled: false
spire-agent:
enabled: true
serviceAccount:
name: spire-agent
server:
address: spire-server.ad.ddupan.top
port: 8081
nodeAttestor:
k8sPSAT:
enabled: true
workloadAttestors:
k8s:
enabled: true
unix:
enabled: false
telemetry:
prometheus:
enabled: true
podMonitor:
enabled: false
resources:
requests:
cpu: 25m
memory: 64Mi
limits:
cpu: 250m
memory: 192Mi
spiffe-csi-driver:
enabled: true
resources:
requests:
cpu: 10m
memory: 32Mi
limits:
cpu: 100m
memory: 96Mi
spiffe-oidc-discovery-provider:
enabled: false
upstream:
enabled: false
tornjak-frontend:
enabled: false
spire-identity-exchange:
enabled: false
spike-keeper:
enabled: false
spike-nexus:
enabled: false
spike-pilot:
enabled: false
@@ -78,6 +78,39 @@ subjects:
--- ---
apiVersion: rbac.authorization.k8s.io/v1 apiVersion: rbac.authorization.k8s.io/v1
kind: Role kind: Role
metadata:
name: spire-bundle-publisher
namespace: spire-system
rules:
- apiGroups:
- ""
resources:
- configmaps
verbs:
- create
- delete
- get
- list
- patch
- update
- watch
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: spire-bundle-publisher
namespace: spire-system
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: spire-bundle-publisher
subjects:
- kind: ServiceAccount
name: spire-controller-manager
namespace: spire-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata: metadata:
name: spire-controller-manager-leader-election name: spire-controller-manager-leader-election
namespace: spire-server namespace: spire-server
@@ -24,7 +24,18 @@ rules:
resources: resources:
- tokenreviews - tokenreviews
verbs: verbs:
- get
- list
- watch
- create - create
- apiGroups:
- ""
resources:
- nodes
- pods
verbs:
- get
- list
--- ---
apiVersion: rbac.authorization.k8s.io/v1 apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding kind: ClusterRoleBinding
+13 -1
View File
@@ -11,7 +11,8 @@ Authelia 提供;SPIRE 不替代人类 OIDC,也不承担目标服务的资源
Flux 安装 SPIFFE hardened charts: Flux 安装 SPIFFE hardened charts:
- `spire-crds` `0.6.1`; - `spire-crds` `0.6.1`;
- `spire` `0.30.2`(SPIRE `1.15.3`); - 内部 fork 的 `spire` `0.30.2-ddupan.1`(SPIRE `1.15.3`),固定 Git tag
`spire-0.30.2-ddupan.1`;
- SPIRE Server、Agent、Controller Manager、SPIFFE CSI Driver; - SPIRE Server、Agent、Controller Manager、SPIFFE CSI Driver;
- OIDC Discovery Provider。 - OIDC Discovery Provider。
@@ -19,6 +20,17 @@ Flux 安装 SPIFFE hardened charts:
API 或 Broker API。Trust domain 是 `ddupan.top`,Kubernetes cluster name 是 API 或 Broker API。Trust domain 是 `ddupan.top`,Kubernetes cluster name 是
`homelab`。 `homelab`。
同一 Server 也接受 cluster name 为 `sandbox` 的 external PSAT attestation。SPIRE gRPC
只通过内网 `spire-server.ad.ddupan.top:8081` 暴露;external PSAT、external
controller-manager 与 bundle publisher 使用由 sandbox Ansible bootstrap 的独立、受限
kubeconfig。Sandbox 不运行第二套 Server 或 OIDC Provider。
内部 fork 仅在上游 `spire-0.30.2` 基础上暴露
`use_pod_uid_for_agent_id`。现有 `sandbox` profile 保持 node UID 模式,供 DaemonSet
Agent 使用;独立的 `sandbox-kata` profile 复用同一 kubeconfig,但启用 Pod UID 模式,
供每个 Kata guest 内的临时 Agent 使用。不得把现有 `sandbox` profile 切换为 Pod UID,
否则会改变常驻 Agent 的 parent ID。
## PostgreSQL bootstrap ## PostgreSQL bootstrap
SPIRE registration datastore 使用共享 CloudNativePG: SPIRE registration datastore 使用共享 CloudNativePG:
+1 -1
View File
@@ -41,7 +41,7 @@ SPIRE Server(trust domain: ddupan.top)
| 项目 | 当前值 | | 项目 | 当前值 |
|---|---| |---|---|
| SPIRE chart | `0.30.2` | | SPIRE chart | `0.30.2-ddupan.1`(内部 fork,基于 `0.30.2`) |
| SPIRE | `1.15.3` | | SPIRE | `1.15.3` |
| SPIRE CRDs chart | `0.6.1` | | SPIRE CRDs chart | `0.6.1` |
| trust domain | `ddupan.top` | | trust domain | `ddupan.top` |
+11
View File
@@ -0,0 +1,11 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
metadata:
name: spiffe-hardened-fork
namespace: spire-mgmt
spec:
interval: 1h
ref:
tag: spire-0.30.2-ddupan.1
timeout: 60s
url: http://gitea-http.gitea.svc.cluster.local:3000/panxiao81/helm-charts-hardened.git
+4 -4
View File
@@ -6,12 +6,12 @@ metadata:
spec: spec:
chart: chart:
spec: spec:
chart: spire chart: ./charts/spire
interval: 1h interval: 1h
reconcileStrategy: Revision
sourceRef: sourceRef:
kind: HelmRepository kind: GitRepository
name: spiffe-hardened name: spiffe-hardened-fork
version: 0.30.2
dependsOn: dependsOn:
- name: spire-crds - name: spire-crds
namespace: spire-mgmt namespace: spire-mgmt
+1
View File
@@ -12,6 +12,7 @@ configMapGenerator:
resources: resources:
- namespaces.yaml - namespaces.yaml
- helmrepository.yaml - helmrepository.yaml
- gitrepository-fork.yaml
- helmrelease-crds.yaml - helmrelease-crds.yaml
- helmrelease.yaml - helmrelease.yaml
- httproute.yaml - httproute.yaml
+46
View File
@@ -30,6 +30,47 @@ spire-server:
kind: statefulset kind: statefulset
replicaCount: 1 replicaCount: 1
auditLogEnabled: true auditLogEnabled: true
service:
type: LoadBalancer
port: 8081
loadBalancerIP: 192.168.10.127
kubeConfigs:
sandbox:
externalSecret:
name: spire-external-kubeconfigs
key: sandbox
sandbox-controller:
externalSecret:
name: spire-external-kubeconfigs
key: sandbox-controller
nodeAttestor:
externalK8sPSAT:
enabled: true
clusters:
sandbox:
kubeConfigName: sandbox
serviceAccountAllowList:
- spire-system:spire-agent
sandbox-kata:
kubeConfigName: sandbox
serviceAccountAllowList:
- spire-smoke:spire-smoke
usePodUIDForAgentID: true
externalControllerManagers:
enabled: true
clusters:
sandbox:
kubeConfigName: sandbox-controller
bundlePublisher:
externalK8sConfigMap:
enabled: true
clusters:
sandbox:
kubeConfigName: sandbox-controller
namespace: spire-system
configMapName: spire-bundle
configMapKey: bundle.spiffe
format: spiffe
persistence: persistence:
# PostgreSQL stores registrations, but the disk KeyManager still needs durable # PostgreSQL stores registrations, but the disk KeyManager still needs durable
# storage for the trust-domain signing keys. # storage for the trust-domain signing keys.
@@ -65,6 +106,11 @@ spire-server:
enabled: false enabled: false
spire-agent: spire-agent:
server:
# Keep the Agent endpoint aligned with spire-server.service.port. The
# chart defaults this to 443, which only remained unnoticed while the
# Agent's pre-upgrade gRPC connection stayed alive.
port: 8081
nodeAttestor: nodeAttestor:
k8sPSAT: k8sPSAT:
enabled: true enabled: true