Author SHA1 Message Date
panxiao81 818192f453 Merge pull request:修复 NATS Account 配额格式
yaml / yaml (push) Successful in 11s
2026-09-16 12:34:12 +00:00
panxiao81 387953c80a 修复 NATS Account 配额格式
yaml / yaml (pull_request) Successful in 10s
2026-09-16 12:33:16 +00:00
panxiao81 298db6a745 Merge pull request:修复 NATS Bao ACME 证书密钥类型
yaml / yaml (push) Successful in 10s
2026-09-16 12:26:19 +00:00
panxiao81 f36a1cbf11 修复 NATS Bao ACME 证书密钥类型
yaml / yaml (pull_request) Successful in 10s
2026-09-16 12:25:28 +00:00
panxiao81 b953db199e Merge pull request:由 monitoring 提供 Prometheus Operator CRD
yaml / yaml (push) Successful in 11s
2026-09-16 12:23:40 +00:00
panxiao81 b7b7975c9f 由 monitoring 提供 Prometheus Operator CRD
yaml / yaml (pull_request) Successful in 10s
2026-09-16 12:22:00 +00:00
panxiao81 9604ff1004 修复 NATS 监控 CRD 兼容性
yaml / yaml (pull_request) Successful in 10s
2026-09-16 12:18:34 +00:00
panxiao81 819b521039 Merge pull request:修复 cert-manager Gateway API ACME solver
yaml / yaml (push) Successful in 11s
2026-09-16 12:17:07 +00:00
panxiao81 40703782ea 修复 cert-manager Gateway API ACME solver
yaml / yaml (pull_request) Successful in 11s
2026-09-16 12:16:23 +00:00
panxiao81 aae19850cf Merge pull request:部署 NATS JetStream 消息基础设施
yaml / yaml (push) Successful in 16s
ansible / collection-test (push) Successful in 1m9s
ansible / lint (push) Successful in 18m16s
首期使用静态 Account 凭据;SPIRE Auth Callout 后续见 #56。YAML 与 collection tests 已通过,Ansible lint 卡在无关的 Galaxy 依赖下载。
2026-09-16 12:13:06 +00:00
panxiao81 c11e1d5e6f 部署 NATS JetStream 消息基础设施
yaml / yaml (pull_request) Successful in 51s
ansible / collection-test (pull_request) Successful in 2m0s
ansible / lint (pull_request) Successful in 19m43s
2026-09-16 12:00:24 +00:00
18 changed files with 295 additions and 2 deletions
+18
View File
@@ -0,0 +1,18 @@
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: nats
namespace: flux-system
spec:
dependsOn:
- name: cert-manager
- name: external-secrets
- name: openebs
interval: 10m
path: ./platform/nats
prune: false
sourceRef:
kind: GitRepository
name: flux-system
timeout: 10m
wait: true
+1
View File
@@ -10,6 +10,7 @@ resources:
- apps/gitea-actions.yaml
- apps/http-echo.yaml
- apps/openebs.yaml
- apps/nats.yaml
- apps/spire.yaml
- apps/observability.yaml
- apps/zot.yaml
+1
View File
@@ -12,6 +12,7 @@ homelab_dns:
- { zone: ad.ddupan.top, name: pve3, type: A, values: [192.168.10.9] }
- { zone: ad.ddupan.top, name: retrolab, type: A, values: [10.60.0.10] }
- { zone: ad.ddupan.top, name: netbox, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: nats, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: s3, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: spire-oidc, type: A, values: [192.168.10.127] }
- { zone: ad.ddupan.top, name: zot, type: A, values: [192.168.10.127] }
+4
View File
@@ -119,6 +119,10 @@ issuerRef:
kind: ClusterIssuer
```
`values.yaml` 必须保持 `config.gatewayAPI.enabled: true`。`bao-acme` 的 HTTP-01
solver 通过共享 Gateway 创建临时 HTTPRoute;关闭该项不会让 ClusterIssuer 变为
NotReady,而是会让每个 Challenge 卡在 `gateway api is not enabled`。
Issuance is capped by `default_directory_policy = role:bao-server`
(`../../infrastructure/openbao/terraform/pki.tf`), which permits `ad.ddupan.top` subdomains only. Clients
need the internal CA in their trust store — already true for the PVE nodes, the DC and
+7
View File
@@ -44,6 +44,13 @@ cainjector:
limits:
memory: 256Mi
# bao-acme solves HTTP-01 through the shared Gateway. The ClusterIssuer can be
# accepted while this is disabled, but every Challenge then stays pending with
# "gateway api is not enabled". Gateway API CRDs are installed by Envoy Gateway.
config:
gatewayAPI:
enabled: true
# ⚠ DNS-01 self-check: cert-manager polls authoritative NS for the _acme-challenge
# TXT record before telling the CA to validate. By default it asks the cluster's
# resolver, which for ad.ddupan.top is CoreDNS -> the Samba AD DC (k3s/coredns-custom.yaml).
+48
View File
@@ -0,0 +1,48 @@
# NATS
共享的轻量消息基础设施。首期为 Gitea microVM runner 提供 JetStream work queue,
但 Account、subject 与部署位置均不与 CI controller 绑定,其他服务可按独立 Account
复用。
## 当前拓扑
- 单节点 NATS;当前 homelab 没有资源运行有意义的三副本 JetStream quorum。
- JetStream file store 使用 `localpv-zfs-ceph`,PVC 2 GiB。
- 服务通过 k3s ServiceLB 在 `nats.ad.ddupan.top:4222` 暴露给内网;集群内客户端
使用 `nats.nats.svc.cluster.local:4222`。访问控制由 TLS、Account 与用户权限负责,
不额外维护易漂移的源 IP 白名单。
- TLS 证书由 `bao-acme` 签发。PVE 节点已信任内部 CA。
- `bao-server` PKI role 只接受 RSA CSR,因此 Certificate 使用 RSA 2048;不要改成
ECDSA,ACME challenge 会成功但 finalize 会以 `role requires keys of type rsa` 失败。
- `SYS` Account 用于管理;`CI` Account 启用 JetStream,存储上限 1 GiB。
Account 内的 JetStream 配额会原样进入 `nats.conf`,必须使用 NATS 的 `MB`/`GB`
格式;PVC 等 Kubernetes resource quantity 才使用 `Mi`/`Gi`。
首期使用静态用户,密码只存在 OpenBao `kv/k8s/nats`:
```text
sys_password
ci_producer_password
ci_worker_password
```
`ci-producer` 只能发布 `ci.runner.>` 并调用必要的 JetStream API;`ci-worker`
只能调用 JetStream pull/ACK API。二者都不能读取另一个 Account 的 subject。
后续 SPIRE/Auth Callout 动态认证见 homelab-infra issue #56。该迁移只替换连接
凭据,不改变 Account、stream、subject 或 consumer。
## CI stream 约定
controller 首次启动时幂等创建 `CI_RUNNER` stream:`ci.runner.*`、
`WorkQueuePolicy`、file storage、24h/10000 条/256 MiB 上限。每类 runner 使用独立
subject 和 durable pull consumer;同类型的多个 worker 共享 durable consumer。
ACK 后消息立即删除,不保存 CI 历史。
## 验证
```bash
kubectl -n nats get helmrelease,pod,pvc,certificate,externalsecret
kubectl -n nats logs statefulset/nats -c nats
```
+21
View File
@@ -0,0 +1,21 @@
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: nats-ad-ddupan-top
namespace: nats
spec:
secretName: nats-server-tls
issuerRef:
name: bao-acme
kind: ClusterIssuer
group: cert-manager.io
commonName: nats.ad.ddupan.top
dnsNames:
- nats.ad.ddupan.top
duration: 720h
renewBefore: 168h
privateKey:
# OpenBao's bao-server role intentionally accepts RSA keys only.
algorithm: RSA
size: 2048
rotationPolicy: Always
+16
View File
@@ -0,0 +1,16 @@
apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
name: nats-auth
namespace: nats
spec:
refreshInterval: 1h
secretStoreRef:
kind: ClusterSecretStore
name: openbao
target:
creationPolicy: Owner
name: nats-auth
dataFrom:
- extract:
key: k8s/nats
+31
View File
@@ -0,0 +1,31 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: nats
namespace: nats
spec:
chart:
spec:
chart: nats
interval: 1h
sourceRef:
kind: HelmRepository
name: nats
version: 2.14.2
driftDetection:
mode: enabled
install:
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: nats
targetNamespace: nats
timeout: 10m
upgrade:
strategy:
name: RetryOnFailure
retryInterval: 5m
valuesFrom:
- kind: ConfigMap
name: nats-values
+8
View File
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: nats
namespace: nats
spec:
interval: 1h
url: https://nats-io.github.io/k8s/helm/charts/
+17
View File
@@ -0,0 +1,17 @@
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
generatorOptions:
disableNameSuffixHash: true
labels:
reconcile.fluxcd.io/watch: Enabled
configMapGenerator:
- name: nats-values
namespace: nats
files:
- values.yaml=values.yaml
resources:
- namespace.yaml
- helmrepository.yaml
- external-secret.yaml
- certificate.yaml
- helmrelease.yaml
+4
View File
@@ -0,0 +1,4 @@
apiVersion: v1
kind: Namespace
metadata:
name: nats
+76
View File
@@ -0,0 +1,76 @@
config:
jetstream:
enabled: true
fileStore:
enabled: true
maxSize: 2Gi
pvc:
enabled: true
size: 2Gi
storageClassName: localpv-zfs-ceph
memoryStore:
enabled: true
maxSize: 64Mi
nats:
tls:
enabled: true
secretName: nats-server-tls
merge:
system_account: SYS
accounts:
SYS:
users:
- user: sys
password: "<< $NATS_SYS_PASSWORD >>"
CI:
jetstream:
# These values are rendered directly into nats.conf and therefore use
# NATS size syntax, not Kubernetes resource.Quantity syntax.
max_memory: 32MB
max_file: 1GB
max_streams: 16
max_consumers: 64
max_bytes_required: true
users:
- user: ci-producer
password: "<< $NATS_CI_PRODUCER_PASSWORD >>"
permissions:
publish:
allow: [ci.runner.>, $JS.API.>]
subscribe:
allow: [_INBOX.>]
- user: ci-worker
password: "<< $NATS_CI_WORKER_PASSWORD >>"
permissions:
publish:
allow: [$JS.API.>, $JS.ACK.>]
subscribe:
allow: [_INBOX.>]
container:
env:
NATS_SYS_PASSWORD:
valueFrom:
secretKeyRef: {name: nats-auth, key: sys_password}
NATS_CI_PRODUCER_PASSWORD:
valueFrom:
secretKeyRef: {name: nats-auth, key: ci_producer_password}
NATS_CI_WORKER_PASSWORD:
valueFrom:
secretKeyRef: {name: nats-auth, key: ci_worker_password}
resources:
requests: {cpu: 25m, memory: 64Mi}
limits: {memory: 192Mi}
natsBox:
enabled: false
promExporter:
enabled: true
podMonitor:
enabled: true
service:
merge:
spec:
type: LoadBalancer
+4 -2
View File
@@ -56,8 +56,10 @@ hosts and `logs/vlogs-ingress.yaml` to push their logs.
| Manage | **victoria-metrics-operator** | VMSingle/VMAgent/VMAlert/VMAlertmanager/VMRule **and** VLSingle as CRDs |
| Expose | **Tailscale ingress** (private) + **Authelia OIDC** | admin tool: private + SSO |
Grafana's Prometheus-operator converter is on, so any chart shipping a
`ServiceMonitor`/`PodMonitor`/`PrometheusRule` is scraped automatically.
VictoriaMetrics Operator 的 Prometheus converter 已启用;官方
`prometheus-operator-crds` chart 由 `operator/` 一并管理。因此应用 chart 可以原生
声明 `ServiceMonitor`、`PodMonitor` 或 `PrometheusRule`,再由 converter 转换为对应
VM 资源,不需要每个应用额外维护一份 `VM*Scrape`。
## Architecture
@@ -3,6 +3,7 @@ kind: Kustomization
resources:
- namespace.yaml
- helmrepository.yaml
- prometheus-helmrepository.yaml
- grafana-helmrepository.yaml
- operator
- metrics
@@ -10,4 +10,5 @@ configMapGenerator:
files:
- values.yaml=values.yaml
resources:
- prometheus-crds-helmrelease.yaml
- helmrelease.yaml
@@ -0,0 +1,29 @@
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: prometheus-operator-crds
namespace: monitoring
spec:
chart:
spec:
chart: prometheus-operator-crds
interval: 1h
sourceRef:
kind: HelmRepository
name: prometheus-community
namespace: monitoring
version: 32.0.0
install:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
interval: 30m
releaseName: prometheus-operator-crds
targetNamespace: monitoring
timeout: 10m
upgrade:
crds: CreateReplace
strategy:
name: RetryOnFailure
retryInterval: 5m
@@ -0,0 +1,8 @@
apiVersion: source.toolkit.fluxcd.io/v1
kind: HelmRepository
metadata:
name: prometheus-community
namespace: monitoring
spec:
interval: 1h
url: https://prometheus-community.github.io/helm-charts