Compare commits
15
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
cd0d3a70ae
|
||
|
|
984c0aee73 | ||
|
|
813b4341d1
|
||
|
|
5a7b38ad26
|
||
|
|
8ac6283573
|
||
|
|
33fb3ec972 | ||
|
|
86953cd91a
|
||
|
|
ddd717209e
|
||
|
|
2756803ba4
|
||
|
|
76752c8443
|
||
|
|
36138f835a
|
||
|
|
de8b7f9b04
|
||
|
|
e6b9980b3e
|
||
|
|
641db531a8
|
||
|
|
7b841d7dba
|
@@ -0,0 +1,76 @@
|
||||
---
|
||||
name: homelab-knowledge
|
||||
description: Query and maintain the shared homelab-wiki when working on homelab services, infrastructure, architecture, operations, or current service status. Use it to gather existing context before work and to keep durable knowledge synchronized after relevant changes; do not use it for unrelated software work or as a substitute for commit and PR history.
|
||||
---
|
||||
|
||||
# Homelab Knowledge
|
||||
|
||||
Use `homelab-wiki` as the shared long-lived knowledge base for people and agents. Search it directly with `rg`; do not introduce a search index, vector database, or generated copy of the wiki.
|
||||
|
||||
## Locate the wiki
|
||||
|
||||
Resolve the checkout in this order:
|
||||
|
||||
1. `$HOMELAB_WIKI_PATH`, when set.
|
||||
2. A sibling directory named `homelab-wiki` next to the current repository.
|
||||
3. `/home/panxiao81/homelab-wiki` when it exists.
|
||||
|
||||
If no checkout is available, report that constraint. Do not silently skip the knowledge step, clone a repository, or create a replacement wiki without the user's authorization.
|
||||
|
||||
Before using the wiki, read its `AGENTS.md` completely. For edits, also read `README.md` and `CONTRIBUTING.md` completely and follow any more specific instructions associated with the target page.
|
||||
|
||||
## Gather context
|
||||
|
||||
At the beginning of a homelab task:
|
||||
|
||||
1. Derive search terms from the component name, service aliases, hostnames, Kubernetes resources, configuration keys, error text, and task intent.
|
||||
2. Use `rg -n -i` in the wiki to find candidate pages. Prefer several precise searches over reading the whole repository.
|
||||
3. Follow the wiki's task index, service index, architecture constraints, source records, and verification conflicts when they are relevant.
|
||||
4. Read the closest authoritative pages and their material links before making decisions. Also read the corresponding source repository README or runbook when changing an implementation.
|
||||
5. Distinguish documented design, declared configuration, deployment history, live verification, and work currently in progress. Do not present one as another.
|
||||
|
||||
For questions about current project or service status, first obtain the maintainer's current-work and ticket context as required by the wiki, unless the conversation already provides that authorization and scope. Reading documentation does not authorize live-system inspection.
|
||||
|
||||
Answer read-only questions from the evidence found. Include paths or links that let the user verify important claims, and state when evidence may be stale or conflicting.
|
||||
|
||||
## Maintain knowledge after changes
|
||||
|
||||
For any code, configuration, infrastructure, or operational change, perform a documentation-impact check before declaring the task complete.
|
||||
|
||||
Update the wiki in the same task when the change affects durable knowledge such as:
|
||||
|
||||
- service purpose, lifecycle, entry point, authentication, permissions, dependencies, or first-use path;
|
||||
- architecture boundaries or accepted constraints;
|
||||
- deployment ownership or persistent operating behavior;
|
||||
- troubleshooting, recovery, verification, or maintenance procedures;
|
||||
- the addition, replacement, or retirement of a service.
|
||||
|
||||
Keep one-time progress, implementation narration, and release-by-release history in commits, PRs, or tickets. Do not copy them into the wiki unless they change a durable stage summary. Implementation-specific parameters may remain in the source repository README or runbook when the wiki convention says to link rather than duplicate them.
|
||||
|
||||
When editing:
|
||||
|
||||
1. Inspect both the source-repository diff and the wiki working tree before writing. Preserve unrelated user changes in both repositories.
|
||||
2. Update the page closest to the fact first, then only the navigation, indexes, constraints, or verification records that the wiki rules require.
|
||||
3. Preserve evidence metadata. Never advance `last_verified` without performing the stated live verification; ordinary review may update only fields permitted by the wiki.
|
||||
4. Link related source commits, PRs, or paths when available. Clearly mark uncommitted sources and unfinished cross-repository synchronization.
|
||||
5. Record conflicts rather than resolving them by assumption. Ask before live inspection or before choosing among materially conflicting current-state claims.
|
||||
6. Keep credentials, tokens, private keys, Terraform state, secret values, and sensitive command output out of documentation. Never read or copy known sensitive files merely to improve the wiki.
|
||||
|
||||
Wiki edits are a separate repository change. Do not commit, push, open a PR, or modify a live system unless the user has authorized that action.
|
||||
|
||||
## Verify and report
|
||||
|
||||
After editing the wiki, run from its root:
|
||||
|
||||
```bash
|
||||
python3 scripts/check_docs.py
|
||||
git diff --check
|
||||
```
|
||||
|
||||
If the checker itself changed, also run:
|
||||
|
||||
```bash
|
||||
python3 -m unittest discover -s tests -v
|
||||
```
|
||||
|
||||
In the final response, report source-repository changes and wiki changes separately, including validation performed and anything still awaiting verification or cross-repository linkage. If no wiki update was needed, state the concrete reason; do not merely say that documentation was unaffected.
|
||||
@@ -20,7 +20,6 @@ import (
|
||||
"sigs.k8s.io/controller-runtime/pkg/webhook"
|
||||
|
||||
executionv1alpha1 "git.ddupan.top/panxiao81/ayatori/api/execution/v1alpha1"
|
||||
"git.ddupan.top/panxiao81/ayatori/internal/controller"
|
||||
// +kubebuilder:scaffold:imports
|
||||
)
|
||||
|
||||
@@ -166,13 +165,6 @@ func main() {
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
if err := (&controller.JobReconciler{
|
||||
Client: mgr.GetClient(),
|
||||
}).SetupWithManager(mgr); err != nil {
|
||||
setupLog.Error(err, "Failed to create controller", "controller", "Job")
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
// +kubebuilder:scaffold:builder
|
||||
|
||||
if err := mgr.AddHealthzCheck("healthz", healthz.Ping); err != nil {
|
||||
|
||||
+6
-53
@@ -1,58 +1,11 @@
|
||||
---
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRole
|
||||
metadata:
|
||||
labels:
|
||||
app.kubernetes.io/name: ayatori
|
||||
app.kubernetes.io/managed-by: kustomize
|
||||
name: manager-role
|
||||
rules:
|
||||
- apiGroups:
|
||||
- ""
|
||||
resources:
|
||||
- namespaces
|
||||
- serviceaccounts
|
||||
verbs:
|
||||
- get
|
||||
- list
|
||||
- watch
|
||||
- apiGroups:
|
||||
- batch
|
||||
resources:
|
||||
- jobs
|
||||
verbs:
|
||||
- create
|
||||
- delete
|
||||
- get
|
||||
- list
|
||||
- watch
|
||||
- apiGroups:
|
||||
- execution.ayatori.ddupan.top
|
||||
resources:
|
||||
- jobclasses
|
||||
- kubernetesexecutionparameters
|
||||
verbs:
|
||||
- get
|
||||
- list
|
||||
- watch
|
||||
- apiGroups:
|
||||
- execution.ayatori.ddupan.top
|
||||
resources:
|
||||
- jobs
|
||||
verbs:
|
||||
- get
|
||||
- list
|
||||
- patch
|
||||
- update
|
||||
- watch
|
||||
- apiGroups:
|
||||
- execution.ayatori.ddupan.top
|
||||
resources:
|
||||
- jobs/finalizers
|
||||
verbs:
|
||||
- update
|
||||
- apiGroups:
|
||||
- execution.ayatori.ddupan.top
|
||||
resources:
|
||||
- jobs/status
|
||||
verbs:
|
||||
- get
|
||||
- patch
|
||||
- update
|
||||
- apiGroups: [""]
|
||||
resources: ["pods"]
|
||||
verbs: ["get", "list", "watch"]
|
||||
|
||||
@@ -21,15 +21,17 @@ Database 是 Ayatori 首批实际产品领域之一。第一个迁移切片只
|
||||
|
||||
代码被移动到 Ayatori 的 `internal/database/domain/instance`,测试 import 和文档链接相应更新;
|
||||
首个后续切片按已批准合同增加 Instance extension observation:观测与当前 target 绑定,进入重新
|
||||
验证或删除时失效,且支持判定不授权 Tenant provisioning。其余 Ready/observation 行为仍应先更新
|
||||
合同与测试再实现,不能把旧运行链路接回该模型。
|
||||
验证或删除时失效,且支持判定不授权 Tenant provisioning。后续 Ready 切片实现管理能力判定、
|
||||
registry 准备决策与完整回读、Ready 重验及本轮 evidence 前置检查;沿用已批准合同,不能把旧运行
|
||||
链路接回该模型。各层验证边界见 [Instance 领域规格](domain-instance.md)。
|
||||
|
||||
## 边界
|
||||
|
||||
- 领域层不依赖 Kubernetes types、数据库 driver 或凭据 provider。
|
||||
- CredentialReference 只携带管理 Secret 的名称与字段映射,不包含 Secret 内容或 OpenBao path。
|
||||
- Instance checkpoint 不是外部事实;实际能力必须由 application/adapter 观察后交给领域对象判断。
|
||||
- 当前代码不授权 Tenant provisioning,也不表示 Database API 已经可用。
|
||||
- 当前代码只检查 Instance 供应前置条件,不授予 Tenant 所有权或外部写入权限,也不表示
|
||||
Database API 已经可用。
|
||||
|
||||
## 设计入口
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
状态:Draft,含已确认决策。日期:2026-09-13。
|
||||
|
||||
上层合并边界见 [ADR-0008](../decisions/0008-merge-Ayatori Database controller.md)。本文只展开 Instance,不包含 Tenant 的供应
|
||||
上层合并边界见 [ADR-0008](../decisions/0008-merge-postgresql-tenant-operator.md)。本文只展开 Instance,不包含 Tenant 的供应
|
||||
实现,也不新增 CRD 字段。设计签名用于评审职责与行为,不是待复制的 Go 接口代码。
|
||||
|
||||
## 1. 对象职责与生命周期
|
||||
@@ -228,3 +228,26 @@ Ready --registry 需修复/保存--> InitializingRegistry
|
||||
均已确认。其他决策及未决项见总体草案,不增加后台清扫器或状态字段。
|
||||
|
||||
批准本对象结构不等于批准这些未决行为,也不意味着立刻实现完整供应链路。
|
||||
|
||||
## 8. 领域实现与验证边界(2026-09-21)
|
||||
|
||||
`internal/database/domain/instance` 按上述方法合同实现 Ready 纯判定。输入分别表达连接、
|
||||
metadata、role、database、grant、extension 管理能力,以及 registry 的未观察、缺失、需迁移、
|
||||
可用、不兼容和不可访问状态。检查零值或未知值按证据不足处理;操作失败只使用封闭的安全
|
||||
类别,不接收驱动错误。`Snapshot.Failure` 是 Condition 映射的领域输入,不新增 CRD/status 字段。
|
||||
|
||||
管理能力分别指目标连接可用、服务器 metadata 可读,以及执行规格 §7 所要求的角色、数据库、
|
||||
授权和扩展管理操作的能力;不是仅凭版本查询或扩展可用列表判定权限。具体 SQL 权限探测矩阵、
|
||||
最小权限角色和扩展权限例外仍须在 PostgreSQL adapter 切片定义并用真实后端验证。
|
||||
registry 不兼容独立保留为领域失败类别,不将其误报为权限不足;公开 Condition Reason 的映射
|
||||
留待 API/application 切片按原合同评审。
|
||||
|
||||
领域测试验证完整回读、缺少检查项、状态重建、重复判定、目标不匹配、配置变化、依赖失败、
|
||||
registry 丢失/不兼容、操作结果不确定和删除限制。只有本轮完整能力判定通过后,Instance
|
||||
前置条件检查才通过;这不授予 Tenant 所有权,也不替代实际写入前的并发校验。
|
||||
|
||||
本切片不新增 controller、adapter、Secret 读取或外部生命周期操作。checkpoint 保存失败、
|
||||
resourceVersion 冲突、watch 与 finalizer 事件链需由后续 application/envtest 验证;SQL 探测、
|
||||
registry 初始化/迁移、超时后的真实状态回读和并发幂等由 PostgreSQL 集成测试验证;Secret
|
||||
变化后的连接刷新由 Kubernetes API 加真实 PostgreSQL 的集成测试验证。纯领域测试不能证明
|
||||
Database 已可运行或这些集成合同已完成。
|
||||
|
||||
@@ -30,8 +30,12 @@ PostgreSQL Tenant Operator 合并为 Ayatori 的 Database 领域模块。保留
|
||||
|
||||
当前没有可用发布版本、没有被该 operator 托管的 PostgreSQL 实例或 Tenant,也没有需要在线
|
||||
转换的已部署 CR。因此此次合并不承担旧实现兼容性:旧运行链路可以直接撤销,不保留直接读取
|
||||
OpenBao 管理凭据的路径,不兼容旧 status checkpoint、samples 或落后于规范的 CRD。API 字段若
|
||||
妨碍清晰领域模型、恢复行为或测试,可以在 `v1alpha1` 阶段修改并重新生成。
|
||||
OpenBao 管理凭据的路径,也不兼容落后于规范的旧 CRD、samples 或实现细节。
|
||||
|
||||
没有部署兼容负担不等于重新设计已经批准的产品合同。源项目的系统规格、API 语义、Instance 与
|
||||
Tenant 领域模型、状态机、ownership registry、OpenBao/ExternalSecret 凭据交付、Retain/Delete、
|
||||
恢复与测试设计整体作为 Ayatori Database 模块的规范基线。除 API group、项目归属和装配结构外,
|
||||
迁移不得静默改变这些行为;确需改变时必须先单独修订规格并记录决定。
|
||||
|
||||
目标结构遵守 Ayatori 的模块化单体边界:
|
||||
|
||||
@@ -53,14 +57,15 @@ Database API 直接重构为 Ayatori 统一结构:API group 使用
|
||||
domain 与 adapter 放入 Ayatori 对应 Database 模块。原 `database.ddupan.top/v1alpha1` 不保留
|
||||
别名、conversion 或兼容入口。
|
||||
|
||||
迁移前逐项核对批准规格、领域模型与当前 Go types;冲突时以批准规格和代码质量为基线,并在
|
||||
Ayatori 中记录有意改变。无需实现在线 CRD conversion 或数据迁移。
|
||||
迁移前逐项核对批准规格、领域模型与当前 Go types;冲突时以批准规格为准。代码质量通过重写
|
||||
旧运行链路、清晰 application/adapter 边界和测试实现,不通过改变已批准行为获得。无需实现在线
|
||||
CRD conversion 或数据迁移。
|
||||
|
||||
## 迁移方式
|
||||
|
||||
1. 以包含已合并 Instance 领域基础和 CI #14 的最新 `main` commit 作为 source reference;记录
|
||||
commit,并先提取规范、领域模型和纯单元测试中仍然成立的部分。目标是保留知识与验证,不是
|
||||
逐文件复制旧实现。
|
||||
commit,将完整批准规格与设计文档迁入 Ayatori Database 文档,并迁移领域模型和纯单元测试。
|
||||
设计合同直接复用;旧运行代码不逐文件复制。
|
||||
2. 保留源仓库暂停中的脏工作树,不移动、提交或复制两个 extension observation 文件。以后可以
|
||||
先在源仓库形成独立 commit,或在 Ayatori 根据批准合同重新实现,但不得把未提交内容描述为来源。
|
||||
3. 在 Ayatori multi-group 项目中用 Kubebuilder 注册 Database API,按批准规格迁移 types,重新
|
||||
@@ -82,5 +87,6 @@ Ayatori 中记录有意改变。无需实现在线 CRD conversion 或数据迁
|
||||
- 已批准的 DBaaS 设计与测试投资得到保留。
|
||||
- 单一 manager/release 不意味着领域耦合;Database 仍保持独立 package、adapter 和测试边界。
|
||||
- 可以从已合并的领域基础开始迁移;旧运行链路和未提交 extension observation 不进入首个切片。
|
||||
- 无部署兼容负担允许优先修正 API 和架构,不为尚未使用的旧代码保留技术债。
|
||||
- 无部署兼容负担允许彻底重写旧运行链路,不为尚未使用的实现技术债保留兼容层;已批准设计合同
|
||||
仍然有效。
|
||||
- Database 使用 Ayatori 统一 API group 与目录结构,不为未投入使用的旧 group 保留入口。
|
||||
|
||||
+3
-2
@@ -51,9 +51,10 @@
|
||||
- 用户集群只暴露 worker node,控制面完全由平台托管。
|
||||
- 本节记录候选实现边界,不构成路线图承诺。
|
||||
|
||||
## 首个业务里程碑
|
||||
## 后续 Compute 验收场景
|
||||
|
||||
完成 Laptop Rebuild Readiness:
|
||||
Database 等首批资源优先落地。Compute 开始实施后,以 Laptop Rebuild Readiness 验证节点
|
||||
生命周期与恢复能力;该场景不作为首批 Database、LoadBalancer 或 Bucket 的交付前置条件:
|
||||
|
||||
1. 临时节点加入。
|
||||
2. laptop 上的 workload 被重建、迁移或形成可执行人工任务。
|
||||
|
||||
@@ -1,132 +0,0 @@
|
||||
package kubernetes
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
|
||||
executionv1alpha1 "git.ddupan.top/panxiao81/ayatori/api/execution/v1alpha1"
|
||||
batchv1 "k8s.io/api/batch/v1"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
|
||||
"k8s.io/apimachinery/pkg/runtime/schema"
|
||||
)
|
||||
|
||||
const (
|
||||
ControllerName = "execution.ayatori.ddupan.top/kubernetes"
|
||||
ReferenceType = "Job"
|
||||
JobUIDLabel = "execution.ayatori.ddupan.top/job-uid"
|
||||
)
|
||||
|
||||
var ayatoriJobGVK = schema.GroupVersionKind{
|
||||
Group: executionv1alpha1.GroupVersion.Group,
|
||||
Version: executionv1alpha1.GroupVersion.Version,
|
||||
Kind: "Job",
|
||||
}
|
||||
|
||||
// BuildJob translates the stable execution API into the Kubernetes adapter's
|
||||
// backend object. It intentionally does not accept or expose a PodSpec.
|
||||
func BuildJob(
|
||||
job *executionv1alpha1.Job,
|
||||
parameters *executionv1alpha1.KubernetesExecutionParameters,
|
||||
resources executionv1alpha1.ExecutionResourceRequirements,
|
||||
) *batchv1.Job {
|
||||
backoffLimit := int32(0)
|
||||
controller := true
|
||||
blockOwnerDeletion := true
|
||||
|
||||
//nolint:modernize // ObjectMeta is promoted through embedded TypeMeta; embedlit produces invalid Go here.
|
||||
return &batchv1.Job{
|
||||
ObjectMeta: metav1.ObjectMeta{
|
||||
Name: job.Name,
|
||||
Namespace: job.Namespace,
|
||||
Labels: map[string]string{
|
||||
JobUIDLabel: string(job.UID),
|
||||
},
|
||||
OwnerReferences: []metav1.OwnerReference{{
|
||||
APIVersion: ayatoriJobGVK.GroupVersion().String(),
|
||||
Kind: ayatoriJobGVK.Kind,
|
||||
Name: job.Name,
|
||||
UID: job.UID,
|
||||
Controller: &controller,
|
||||
BlockOwnerDeletion: &blockOwnerDeletion,
|
||||
}},
|
||||
},
|
||||
Spec: batchv1.JobSpec{
|
||||
BackoffLimit: &backoffLimit,
|
||||
Template: corev1.PodTemplateSpec{
|
||||
ObjectMeta: metav1.ObjectMeta{Labels: map[string]string{JobUIDLabel: string(job.UID)}},
|
||||
Spec: corev1.PodSpec{
|
||||
RestartPolicy: corev1.RestartPolicyNever,
|
||||
ServiceAccountName: parameters.Spec.ServiceAccountName,
|
||||
RuntimeClassName: optionalString(parameters.Spec.RuntimeClassName),
|
||||
NodeSelector: parameters.Spec.Scheduling.NodeSelector,
|
||||
Tolerations: parameters.Spec.Scheduling.Tolerations,
|
||||
SecurityContext: parameters.Spec.PodSecurityContext,
|
||||
ImagePullSecrets: job.Spec.Task.ImagePullSecrets,
|
||||
Containers: []corev1.Container{{
|
||||
Name: "task",
|
||||
Image: job.Spec.Task.Image,
|
||||
ImagePullPolicy: parameters.Spec.ImagePullPolicy,
|
||||
Command: job.Spec.Task.Command,
|
||||
Args: job.Spec.Task.Args,
|
||||
WorkingDir: job.Spec.Task.WorkingDir,
|
||||
Env: environment(job.Spec.Task.Env),
|
||||
Resources: resourceRequirements(resources),
|
||||
}},
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
func ValidateOwnership(owner *executionv1alpha1.Job, backend *batchv1.Job) error {
|
||||
if backend.Labels[JobUIDLabel] != string(owner.UID) {
|
||||
return fmt.Errorf("backend Job %s/%s is not owned by Ayatori Job UID %s", backend.Namespace, backend.Name, owner.UID)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func environment(values []executionv1alpha1.EnvVar) []corev1.EnvVar {
|
||||
result := make([]corev1.EnvVar, 0, len(values))
|
||||
for _, value := range values {
|
||||
env := corev1.EnvVar{Name: value.Name}
|
||||
if value.Value != nil {
|
||||
env.Value = *value.Value
|
||||
}
|
||||
if value.ValueFrom != nil {
|
||||
env.ValueFrom = &corev1.EnvVarSource{
|
||||
SecretKeyRef: value.ValueFrom.SecretKeyRef,
|
||||
ConfigMapKeyRef: value.ValueFrom.ConfigMapKeyRef,
|
||||
}
|
||||
}
|
||||
result = append(result, env)
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
func resourceRequirements(resources executionv1alpha1.ExecutionResourceRequirements) corev1.ResourceRequirements {
|
||||
return corev1.ResourceRequirements{
|
||||
Requests: resourceList(resources.Requests),
|
||||
Limits: resourceList(resources.Limits),
|
||||
}
|
||||
}
|
||||
|
||||
func resourceList(values executionv1alpha1.ResourceValues) corev1.ResourceList {
|
||||
result := corev1.ResourceList{}
|
||||
if values.CPU != nil {
|
||||
result[corev1.ResourceCPU] = values.CPU.DeepCopy()
|
||||
}
|
||||
if values.Memory != nil {
|
||||
result[corev1.ResourceMemory] = values.Memory.DeepCopy()
|
||||
}
|
||||
if len(result) == 0 {
|
||||
return nil
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
func optionalString(value string) *string {
|
||||
if value == "" {
|
||||
return nil
|
||||
}
|
||||
return &value
|
||||
}
|
||||
@@ -1,71 +0,0 @@
|
||||
package kubernetes
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
executionv1alpha1 "git.ddupan.top/panxiao81/ayatori/api/execution/v1alpha1"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
"k8s.io/apimachinery/pkg/api/resource"
|
||||
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
|
||||
"k8s.io/apimachinery/pkg/types"
|
||||
)
|
||||
|
||||
func TestBuildJob(t *testing.T) {
|
||||
literal := "world"
|
||||
cpuRequest := resource.MustParse("100m")
|
||||
memoryLimit := resource.MustParse("128Mi")
|
||||
//nolint:modernize // ObjectMeta is promoted through embedded TypeMeta; embedlit produces invalid Go here.
|
||||
job := &executionv1alpha1.Job{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: "hello", Namespace: "ci", UID: types.UID("job-uid")},
|
||||
Spec: executionv1alpha1.JobSpec{Task: executionv1alpha1.TaskSpec{
|
||||
Image: "alpine:3.22", Command: []string{"echo"}, Args: []string{"hello"},
|
||||
Env: []executionv1alpha1.EnvVar{
|
||||
{Name: "TARGET", Value: &literal},
|
||||
{Name: "TOKEN", ValueFrom: &executionv1alpha1.EnvVarSource{
|
||||
//nolint:modernize // LocalObjectReference is an embedded Kubernetes API field.
|
||||
SecretKeyRef: &corev1.SecretKeySelector{LocalObjectReference: corev1.LocalObjectReference{Name: "token"}, Key: "value"},
|
||||
}},
|
||||
},
|
||||
}},
|
||||
}
|
||||
parameters := &executionv1alpha1.KubernetesExecutionParameters{Spec: executionv1alpha1.KubernetesExecutionParametersSpec{
|
||||
ServiceAccountName: "runner", RuntimeClassName: "runc", ImagePullPolicy: corev1.PullIfNotPresent,
|
||||
Scheduling: executionv1alpha1.KubernetesSchedulingParameters{NodeSelector: map[string]string{"role": "execution"}},
|
||||
}}
|
||||
resources := executionv1alpha1.ExecutionResourceRequirements{
|
||||
Requests: executionv1alpha1.ResourceValues{CPU: &cpuRequest},
|
||||
Limits: executionv1alpha1.ResourceValues{Memory: &memoryLimit},
|
||||
}
|
||||
|
||||
backend := BuildJob(job, parameters, resources)
|
||||
pod := backend.Spec.Template.Spec
|
||||
if backend.Spec.BackoffLimit == nil || *backend.Spec.BackoffLimit != 0 {
|
||||
t.Fatalf("backoffLimit = %v, want 0", backend.Spec.BackoffLimit)
|
||||
}
|
||||
if pod.RestartPolicy != corev1.RestartPolicyNever || pod.ServiceAccountName != "runner" {
|
||||
t.Fatalf("unexpected pod execution policy: %#v", pod)
|
||||
}
|
||||
if pod.RuntimeClassName == nil || *pod.RuntimeClassName != "runc" {
|
||||
t.Fatalf("runtimeClassName = %v, want runc", pod.RuntimeClassName)
|
||||
}
|
||||
container := pod.Containers[0]
|
||||
if container.Resources.Requests.Cpu().Cmp(cpuRequest) != 0 || container.Resources.Limits.Memory().Cmp(memoryLimit) != 0 {
|
||||
t.Fatalf("resources were not mapped: %#v", container.Resources)
|
||||
}
|
||||
if container.Env[1].ValueFrom == nil || container.Env[1].ValueFrom.SecretKeyRef.Name != "token" {
|
||||
t.Fatalf("secret reference was not preserved: %#v", container.Env[1])
|
||||
}
|
||||
if backend.Labels[JobUIDLabel] != "job-uid" || backend.OwnerReferences[0].UID != job.UID {
|
||||
t.Fatalf("ownership identity was not preserved: %#v", backend.ObjectMeta)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateOwnership(t *testing.T) {
|
||||
job := &executionv1alpha1.Job{}
|
||||
job.UID = types.UID("expected")
|
||||
backend := BuildJob(job, &executionv1alpha1.KubernetesExecutionParameters{}, executionv1alpha1.ExecutionResourceRequirements{})
|
||||
backend.Labels[JobUIDLabel] = "different"
|
||||
if err := ValidateOwnership(job, backend); err == nil {
|
||||
t.Fatal("ValidateOwnership() succeeded for a different Job UID")
|
||||
}
|
||||
}
|
||||
@@ -1,397 +0,0 @@
|
||||
package controller
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"slices"
|
||||
"time"
|
||||
|
||||
executionv1alpha1 "git.ddupan.top/panxiao81/ayatori/api/execution/v1alpha1"
|
||||
kubernetesadapter "git.ddupan.top/panxiao81/ayatori/internal/adapter/kubernetes"
|
||||
batchv1 "k8s.io/api/batch/v1"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
apierrors "k8s.io/apimachinery/pkg/api/errors"
|
||||
"k8s.io/apimachinery/pkg/api/meta"
|
||||
"k8s.io/apimachinery/pkg/api/resource"
|
||||
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
|
||||
"k8s.io/apimachinery/pkg/labels"
|
||||
"k8s.io/apimachinery/pkg/types"
|
||||
ctrl "sigs.k8s.io/controller-runtime"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
"sigs.k8s.io/controller-runtime/pkg/log"
|
||||
)
|
||||
|
||||
const (
|
||||
jobFinalizer = "execution.ayatori.ddupan.top/job-cleanup"
|
||||
reasonResultUnknown = "ResultUnknown"
|
||||
)
|
||||
|
||||
// JobReconciler executes Ayatori Jobs using supported adapters.
|
||||
type JobReconciler struct {
|
||||
client.Client
|
||||
Now func() time.Time
|
||||
}
|
||||
|
||||
// +kubebuilder:rbac:groups=execution.ayatori.ddupan.top,resources=jobs,verbs=get;list;watch;update;patch
|
||||
// +kubebuilder:rbac:groups=execution.ayatori.ddupan.top,resources=jobs/status,verbs=get;update;patch
|
||||
// +kubebuilder:rbac:groups=execution.ayatori.ddupan.top,resources=jobs/finalizers,verbs=update
|
||||
// +kubebuilder:rbac:groups=execution.ayatori.ddupan.top,resources=jobclasses;kubernetesexecutionparameters,verbs=get;list;watch
|
||||
// +kubebuilder:rbac:groups=batch,resources=jobs,verbs=get;list;watch;create;delete
|
||||
// +kubebuilder:rbac:groups="",resources=namespaces;serviceaccounts,verbs=get;list;watch
|
||||
|
||||
func (r *JobReconciler) Reconcile(ctx context.Context, request ctrl.Request) (ctrl.Result, error) {
|
||||
logger := log.FromContext(ctx)
|
||||
job := &executionv1alpha1.Job{}
|
||||
if err := r.Get(ctx, request.NamespacedName, job); err != nil {
|
||||
return ctrl.Result{}, client.IgnoreNotFound(err)
|
||||
}
|
||||
|
||||
if !job.DeletionTimestamp.IsZero() {
|
||||
return ctrl.Result{}, r.finalize(ctx, job)
|
||||
}
|
||||
if isTerminal(job) {
|
||||
return ctrl.Result{}, nil
|
||||
}
|
||||
if job.Spec.DesiredState == executionv1alpha1.JobDesiredStateCancelled {
|
||||
return ctrl.Result{}, r.cancel(ctx, job)
|
||||
}
|
||||
|
||||
if !containsString(job.Finalizers, jobFinalizer) {
|
||||
job.Finalizers = append(job.Finalizers, jobFinalizer)
|
||||
if err := r.Update(ctx, job); err != nil {
|
||||
return ctrl.Result{}, err
|
||||
}
|
||||
return ctrl.Result{}, nil
|
||||
}
|
||||
if job.Status.Execution != nil {
|
||||
return ctrl.Result{}, r.observeExisting(ctx, job)
|
||||
}
|
||||
|
||||
class, parameters, resources, waiting, err := r.resolve(ctx, job)
|
||||
if err != nil {
|
||||
return ctrl.Result{}, err
|
||||
}
|
||||
if waiting {
|
||||
return ctrl.Result{RequeueAfter: 30 * time.Second}, nil
|
||||
}
|
||||
|
||||
backend := &batchv1.Job{}
|
||||
key := types.NamespacedName{Namespace: job.Namespace, Name: job.Name}
|
||||
err = r.Get(ctx, key, backend)
|
||||
if apierrors.IsNotFound(err) {
|
||||
backend = kubernetesadapter.BuildJob(job, parameters, resources)
|
||||
if err := r.Create(ctx, backend); err != nil {
|
||||
return ctrl.Result{}, err
|
||||
}
|
||||
logger.Info("Created Kubernetes backend Job", "backend", key)
|
||||
return ctrl.Result{}, r.markScheduled(ctx, job, class, parameters, resources, backend)
|
||||
}
|
||||
if err != nil {
|
||||
return ctrl.Result{}, err
|
||||
}
|
||||
if err := kubernetesadapter.ValidateOwnership(job, backend); err != nil {
|
||||
return ctrl.Result{}, r.setCondition(ctx, job, metav1.Condition{
|
||||
Type: executionv1alpha1.JobConditionScheduled, Status: metav1.ConditionFalse,
|
||||
Reason: "BackendConflict", Message: err.Error(),
|
||||
})
|
||||
}
|
||||
if !conditionTrue(job.Status.Conditions, executionv1alpha1.JobConditionScheduled) {
|
||||
return ctrl.Result{}, r.markScheduled(ctx, job, class, parameters, resources, backend)
|
||||
}
|
||||
return ctrl.Result{}, r.observe(ctx, job, backend)
|
||||
}
|
||||
|
||||
func (r *JobReconciler) observeExisting(ctx context.Context, job *executionv1alpha1.Job) error {
|
||||
if job.Status.Execution.Adapter != "kubernetes" {
|
||||
return r.setCondition(ctx, job, metav1.Condition{
|
||||
Type: executionv1alpha1.JobConditionSucceeded, Status: metav1.ConditionUnknown,
|
||||
Reason: reasonResultUnknown, Message: fmt.Sprintf("adapter %q is not available", job.Status.Execution.Adapter),
|
||||
})
|
||||
}
|
||||
backend := &batchv1.Job{}
|
||||
key := types.NamespacedName{Namespace: job.Namespace, Name: job.Name}
|
||||
if err := r.Get(ctx, key, backend); err != nil {
|
||||
if apierrors.IsNotFound(err) {
|
||||
return r.setCondition(ctx, job, metav1.Condition{
|
||||
Type: executionv1alpha1.JobConditionSucceeded, Status: metav1.ConditionUnknown,
|
||||
Reason: reasonResultUnknown, Message: "Kubernetes backend Job is missing",
|
||||
})
|
||||
}
|
||||
return err
|
||||
}
|
||||
if err := kubernetesadapter.ValidateOwnership(job, backend); err != nil {
|
||||
return r.setCondition(ctx, job, metav1.Condition{
|
||||
Type: executionv1alpha1.JobConditionSucceeded, Status: metav1.ConditionUnknown,
|
||||
Reason: reasonResultUnknown, Message: err.Error(),
|
||||
})
|
||||
}
|
||||
return r.observe(ctx, job, backend)
|
||||
}
|
||||
|
||||
func (r *JobReconciler) resolve(
|
||||
ctx context.Context,
|
||||
job *executionv1alpha1.Job,
|
||||
) (*executionv1alpha1.JobClass, *executionv1alpha1.KubernetesExecutionParameters, executionv1alpha1.ExecutionResourceRequirements, bool, error) {
|
||||
if job.Spec.JobClassName == "" {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, true, r.reject(ctx, job, "NoDefaultJobClass", "spec.jobClassName is required in the first implementation slice")
|
||||
}
|
||||
|
||||
class := &executionv1alpha1.JobClass{}
|
||||
if err := r.Get(ctx, types.NamespacedName{Name: job.Spec.JobClassName}, class); err != nil {
|
||||
if apierrors.IsNotFound(err) {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, true, r.reject(ctx, job, "JobClassNotFound", fmt.Sprintf("JobClass %q does not exist", job.Spec.JobClassName))
|
||||
}
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, false, err
|
||||
}
|
||||
if class.Spec.ControllerName != kubernetesadapter.ControllerName {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, true, r.reject(ctx, job, "UnsupportedController", fmt.Sprintf("controller %q is not supported", class.Spec.ControllerName))
|
||||
}
|
||||
ref := class.Spec.ParametersRef
|
||||
if ref.Group != executionv1alpha1.GroupVersion.Group || ref.Kind != "KubernetesExecutionParameters" {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, true, r.reject(ctx, job, "InvalidParametersReference", "JobClass must reference KubernetesExecutionParameters")
|
||||
}
|
||||
if allowed, err := r.namespaceAllowed(ctx, job.Namespace, class.Spec.AllowedNamespaces); err != nil {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, false, err
|
||||
} else if !allowed {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, true, r.reject(ctx, job, "NamespaceNotAllowed", fmt.Sprintf("namespace %q is not allowed by JobClass %q", job.Namespace, class.Name))
|
||||
}
|
||||
|
||||
parameters := &executionv1alpha1.KubernetesExecutionParameters{}
|
||||
if err := r.Get(ctx, types.NamespacedName{Name: ref.Name}, parameters); err != nil {
|
||||
if apierrors.IsNotFound(err) {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, true, r.reject(ctx, job, "ParametersNotFound", fmt.Sprintf("KubernetesExecutionParameters %q does not exist", ref.Name))
|
||||
}
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, false, err
|
||||
}
|
||||
serviceAccount := &corev1.ServiceAccount{}
|
||||
if err := r.Get(ctx, types.NamespacedName{Namespace: job.Namespace, Name: parameters.Spec.ServiceAccountName}, serviceAccount); err != nil {
|
||||
if apierrors.IsNotFound(err) {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, true, r.reject(ctx, job, "ServiceAccountNotFound", fmt.Sprintf("ServiceAccount %q does not exist", parameters.Spec.ServiceAccountName))
|
||||
}
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, false, err
|
||||
}
|
||||
|
||||
resources := applyResourceDefaults(job.Spec.Resources, class.Spec.Resources.Defaults)
|
||||
if err := validateResources(resources); err != nil {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, true, r.reject(ctx, job, "InvalidResources", err.Error())
|
||||
}
|
||||
if err := r.accept(ctx, job, class, parameters, resources); err != nil {
|
||||
return nil, nil, executionv1alpha1.ExecutionResourceRequirements{}, false, err
|
||||
}
|
||||
return class, parameters, resources, false, nil
|
||||
}
|
||||
|
||||
func (r *JobReconciler) namespaceAllowed(ctx context.Context, namespace string, selector *metav1.LabelSelector) (bool, error) {
|
||||
if selector == nil {
|
||||
return true, nil
|
||||
}
|
||||
ns := &corev1.Namespace{}
|
||||
if err := r.Get(ctx, types.NamespacedName{Name: namespace}, ns); err != nil {
|
||||
return false, err
|
||||
}
|
||||
compiled, err := metav1.LabelSelectorAsSelector(selector)
|
||||
if err != nil {
|
||||
return false, err
|
||||
}
|
||||
return compiled.Matches(labels.Set(ns.Labels)), nil
|
||||
}
|
||||
|
||||
func (r *JobReconciler) accept(ctx context.Context, job *executionv1alpha1.Job, class *executionv1alpha1.JobClass, parameters *executionv1alpha1.KubernetesExecutionParameters, resources executionv1alpha1.ExecutionResourceRequirements) error {
|
||||
job.Status.ResolvedJobClass = &executionv1alpha1.ResolvedJobClassReference{
|
||||
Name: class.Name, UID: class.UID, ControllerName: class.Spec.ControllerName,
|
||||
ParametersRef: executionv1alpha1.ParametersReference{
|
||||
Group: class.Spec.ParametersRef.Group, Kind: class.Spec.ParametersRef.Kind,
|
||||
Name: parameters.Name, UID: parameters.UID,
|
||||
},
|
||||
}
|
||||
job.Status.EffectiveResources = resources
|
||||
return r.setCondition(ctx, job, metav1.Condition{
|
||||
Type: executionv1alpha1.JobConditionAccepted, Status: metav1.ConditionTrue,
|
||||
Reason: "Accepted", Message: fmt.Sprintf("JobClass %q accepted", class.Name),
|
||||
})
|
||||
}
|
||||
|
||||
func (r *JobReconciler) reject(ctx context.Context, job *executionv1alpha1.Job, reason, message string) error {
|
||||
return r.setCondition(ctx, job, metav1.Condition{
|
||||
Type: executionv1alpha1.JobConditionAccepted, Status: metav1.ConditionFalse,
|
||||
Reason: reason, Message: message,
|
||||
})
|
||||
}
|
||||
|
||||
func (r *JobReconciler) markScheduled(ctx context.Context, job *executionv1alpha1.Job, class *executionv1alpha1.JobClass, parameters *executionv1alpha1.KubernetesExecutionParameters, resources executionv1alpha1.ExecutionResourceRequirements, backend *batchv1.Job) error {
|
||||
job.Status.ResolvedJobClass = &executionv1alpha1.ResolvedJobClassReference{
|
||||
Name: class.Name, UID: class.UID, ControllerName: class.Spec.ControllerName,
|
||||
ParametersRef: executionv1alpha1.ParametersReference{Group: class.Spec.ParametersRef.Group, Kind: class.Spec.ParametersRef.Kind, Name: parameters.Name, UID: parameters.UID},
|
||||
}
|
||||
job.Status.EffectiveResources = resources
|
||||
job.Status.Execution = &executionv1alpha1.ExecutionStatus{
|
||||
Adapter: "kubernetes",
|
||||
References: []executionv1alpha1.ExecutionReference{{Type: kubernetesadapter.ReferenceType, ID: string(backend.UID)}},
|
||||
}
|
||||
meta.SetStatusCondition(&job.Status.Conditions, condition(job, executionv1alpha1.JobConditionAccepted, metav1.ConditionTrue, "Accepted", "Job accepted"))
|
||||
meta.SetStatusCondition(&job.Status.Conditions, condition(job, executionv1alpha1.JobConditionScheduled, metav1.ConditionTrue, "BackendCreated", "Kubernetes Job created"))
|
||||
meta.SetStatusCondition(&job.Status.Conditions, condition(job, executionv1alpha1.JobConditionSucceeded, metav1.ConditionUnknown, "Pending", "Waiting for task to start"))
|
||||
job.Status.ObservedGeneration = job.Generation
|
||||
return r.Status().Update(ctx, job)
|
||||
}
|
||||
|
||||
func (r *JobReconciler) observe(ctx context.Context, job *executionv1alpha1.Job, backend *batchv1.Job) error {
|
||||
if job.Status.StartTime == nil && backend.Status.StartTime != nil {
|
||||
job.Status.StartTime = backend.Status.StartTime.DeepCopy()
|
||||
}
|
||||
for _, backendCondition := range backend.Status.Conditions {
|
||||
switch {
|
||||
case backendCondition.Type == batchv1.JobComplete && backendCondition.Status == corev1.ConditionTrue:
|
||||
completion := backend.Status.CompletionTime
|
||||
if completion == nil {
|
||||
now := metav1.NewTime(r.now())
|
||||
completion = &now
|
||||
}
|
||||
job.Status.CompletionTime = completion.DeepCopy()
|
||||
job.Status.Result = &executionv1alpha1.JobResult{Reason: "Completed"}
|
||||
return r.setCondition(ctx, job, metav1.Condition{Type: executionv1alpha1.JobConditionSucceeded, Status: metav1.ConditionTrue, Reason: "Completed", Message: backendCondition.Message})
|
||||
case backendCondition.Type == batchv1.JobFailed && backendCondition.Status == corev1.ConditionTrue:
|
||||
completion := metav1.NewTime(r.now())
|
||||
job.Status.CompletionTime = &completion
|
||||
job.Status.Result = &executionv1alpha1.JobResult{Reason: "ProcessFailed"}
|
||||
return r.setCondition(ctx, job, metav1.Condition{Type: executionv1alpha1.JobConditionSucceeded, Status: metav1.ConditionFalse, Reason: "ProcessFailed", Message: backendCondition.Message})
|
||||
}
|
||||
}
|
||||
reason := "Pending"
|
||||
message := "Waiting for task to start"
|
||||
if backend.Status.StartTime != nil || backend.Status.Active > 0 {
|
||||
reason = "Running"
|
||||
message = "Task is running"
|
||||
}
|
||||
return r.setCondition(ctx, job, metav1.Condition{Type: executionv1alpha1.JobConditionSucceeded, Status: metav1.ConditionUnknown, Reason: reason, Message: message})
|
||||
}
|
||||
|
||||
func (r *JobReconciler) cancel(ctx context.Context, job *executionv1alpha1.Job) error {
|
||||
backend := &batchv1.Job{}
|
||||
key := types.NamespacedName{Namespace: job.Namespace, Name: job.Name}
|
||||
err := r.Get(ctx, key, backend)
|
||||
if err == nil {
|
||||
if err := kubernetesadapter.ValidateOwnership(job, backend); err != nil {
|
||||
return err
|
||||
}
|
||||
for _, backendCondition := range backend.Status.Conditions {
|
||||
if (backendCondition.Type == batchv1.JobComplete || backendCondition.Type == batchv1.JobFailed) &&
|
||||
backendCondition.Status == corev1.ConditionTrue {
|
||||
return r.observe(ctx, job, backend)
|
||||
}
|
||||
}
|
||||
if err := r.Delete(ctx, backend, client.PropagationPolicy(metav1.DeletePropagationBackground)); err != nil && !apierrors.IsNotFound(err) {
|
||||
return err
|
||||
}
|
||||
return nil
|
||||
}
|
||||
if !apierrors.IsNotFound(err) {
|
||||
return err
|
||||
}
|
||||
now := metav1.NewTime(r.now())
|
||||
job.Status.CompletionTime = &now
|
||||
job.Status.Result = &executionv1alpha1.JobResult{Reason: "Cancelled"}
|
||||
return r.setCondition(ctx, job, metav1.Condition{Type: executionv1alpha1.JobConditionSucceeded, Status: metav1.ConditionFalse, Reason: "Cancelled", Message: "Execution cancelled"})
|
||||
}
|
||||
|
||||
func (r *JobReconciler) finalize(ctx context.Context, job *executionv1alpha1.Job) error {
|
||||
if !containsString(job.Finalizers, jobFinalizer) {
|
||||
return nil
|
||||
}
|
||||
backend := &batchv1.Job{}
|
||||
key := types.NamespacedName{Namespace: job.Namespace, Name: job.Name}
|
||||
if err := r.Get(ctx, key, backend); err == nil {
|
||||
if err := kubernetesadapter.ValidateOwnership(job, backend); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := r.Delete(ctx, backend, client.PropagationPolicy(metav1.DeletePropagationBackground)); err != nil && !apierrors.IsNotFound(err) {
|
||||
return err
|
||||
}
|
||||
return nil
|
||||
} else if !apierrors.IsNotFound(err) {
|
||||
return err
|
||||
}
|
||||
job.Finalizers = removeString(job.Finalizers, jobFinalizer)
|
||||
return r.Update(ctx, job)
|
||||
}
|
||||
|
||||
func (r *JobReconciler) setCondition(ctx context.Context, job *executionv1alpha1.Job, next metav1.Condition) error {
|
||||
meta.SetStatusCondition(&job.Status.Conditions, condition(job, next.Type, next.Status, next.Reason, next.Message))
|
||||
job.Status.ObservedGeneration = job.Generation
|
||||
return r.Status().Update(ctx, job)
|
||||
}
|
||||
|
||||
func condition(job *executionv1alpha1.Job, conditionType string, status metav1.ConditionStatus, reason, message string) metav1.Condition {
|
||||
return metav1.Condition{Type: conditionType, Status: status, Reason: reason, Message: message, ObservedGeneration: job.Generation}
|
||||
}
|
||||
|
||||
func conditionTrue(conditions []metav1.Condition, conditionType string) bool {
|
||||
current := meta.FindStatusCondition(conditions, conditionType)
|
||||
return current != nil && current.Status == metav1.ConditionTrue
|
||||
}
|
||||
|
||||
func isTerminal(job *executionv1alpha1.Job) bool {
|
||||
current := meta.FindStatusCondition(job.Status.Conditions, executionv1alpha1.JobConditionSucceeded)
|
||||
return current != nil && (current.Status == metav1.ConditionTrue || current.Status == metav1.ConditionFalse)
|
||||
}
|
||||
|
||||
func applyResourceDefaults(requested, defaults executionv1alpha1.ExecutionResourceRequirements) executionv1alpha1.ExecutionResourceRequirements {
|
||||
result := requested.DeepCopy()
|
||||
if result.Requests.CPU == nil && defaults.Requests.CPU != nil {
|
||||
result.Requests.CPU = copyQuantity(defaults.Requests.CPU)
|
||||
}
|
||||
if result.Requests.Memory == nil && defaults.Requests.Memory != nil {
|
||||
result.Requests.Memory = copyQuantity(defaults.Requests.Memory)
|
||||
}
|
||||
if result.Limits.CPU == nil && defaults.Limits.CPU != nil {
|
||||
result.Limits.CPU = copyQuantity(defaults.Limits.CPU)
|
||||
}
|
||||
if result.Limits.Memory == nil && defaults.Limits.Memory != nil {
|
||||
result.Limits.Memory = copyQuantity(defaults.Limits.Memory)
|
||||
}
|
||||
return *result
|
||||
}
|
||||
|
||||
func copyQuantity(value *resource.Quantity) *resource.Quantity {
|
||||
copy := value.DeepCopy()
|
||||
return ©
|
||||
}
|
||||
|
||||
func validateResources(resources executionv1alpha1.ExecutionResourceRequirements) error {
|
||||
if resources.Requests.CPU != nil && resources.Limits.CPU != nil && resources.Requests.CPU.Cmp(*resources.Limits.CPU) > 0 {
|
||||
return fmt.Errorf("CPU request must not exceed limit")
|
||||
}
|
||||
if resources.Requests.Memory != nil && resources.Limits.Memory != nil && resources.Requests.Memory.Cmp(*resources.Limits.Memory) > 0 {
|
||||
return fmt.Errorf("memory request must not exceed limit")
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func containsString(values []string, target string) bool {
|
||||
return slices.Contains(values, target)
|
||||
}
|
||||
|
||||
func removeString(values []string, target string) []string {
|
||||
result := values[:0]
|
||||
for _, value := range values {
|
||||
if value != target {
|
||||
result = append(result, value)
|
||||
}
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
func (r *JobReconciler) now() time.Time {
|
||||
if r.Now != nil {
|
||||
return r.Now()
|
||||
}
|
||||
return time.Now()
|
||||
}
|
||||
|
||||
func (r *JobReconciler) SetupWithManager(manager ctrl.Manager) error {
|
||||
return ctrl.NewControllerManagedBy(manager).
|
||||
For(&executionv1alpha1.Job{}).
|
||||
Owns(&batchv1.Job{}).
|
||||
Named("execution-job").
|
||||
Complete(r)
|
||||
}
|
||||
@@ -1,202 +0,0 @@
|
||||
package controller
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
executionv1alpha1 "git.ddupan.top/panxiao81/ayatori/api/execution/v1alpha1"
|
||||
kubernetesadapter "git.ddupan.top/panxiao81/ayatori/internal/adapter/kubernetes"
|
||||
batchv1 "k8s.io/api/batch/v1"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
"k8s.io/apimachinery/pkg/api/meta"
|
||||
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
|
||||
"k8s.io/apimachinery/pkg/runtime"
|
||||
"k8s.io/apimachinery/pkg/types"
|
||||
ctrl "sigs.k8s.io/controller-runtime"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
"sigs.k8s.io/controller-runtime/pkg/envtest"
|
||||
metricsserver "sigs.k8s.io/controller-runtime/pkg/metrics/server"
|
||||
)
|
||||
|
||||
const (
|
||||
integrationNamespace = "controller-integration"
|
||||
integrationClass = "integration"
|
||||
)
|
||||
|
||||
//nolint:modernize // Kubernetes API structs expose ObjectMeta through embedded TypeMeta fields.
|
||||
func TestJobControllerIntegration(t *testing.T) {
|
||||
if os.Getenv("KUBEBUILDER_ASSETS") == "" {
|
||||
t.Skip("KUBEBUILDER_ASSETS is unset; run make test to execute controller integration tests")
|
||||
}
|
||||
|
||||
scheme := runtime.NewScheme()
|
||||
for _, addToScheme := range []func(*runtime.Scheme) error{
|
||||
corev1.AddToScheme,
|
||||
batchv1.AddToScheme,
|
||||
executionv1alpha1.AddToScheme,
|
||||
} {
|
||||
if err := addToScheme(scheme); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
crdPath, err := filepath.Abs("../../config/crd/bases")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
environment := &envtest.Environment{CRDDirectoryPaths: []string{crdPath}}
|
||||
config, err := environment.Start()
|
||||
if err != nil {
|
||||
t.Fatalf("start envtest: %v", err)
|
||||
}
|
||||
t.Cleanup(func() {
|
||||
if err := environment.Stop(); err != nil {
|
||||
t.Errorf("stop envtest: %v", err)
|
||||
}
|
||||
})
|
||||
|
||||
manager, err := ctrl.NewManager(config, ctrl.Options{
|
||||
Scheme: scheme,
|
||||
Metrics: metricsserver.Options{BindAddress: "0"},
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := (&JobReconciler{Client: manager.GetClient()}).SetupWithManager(manager); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
managerContext, cancelManager := context.WithCancel(context.Background())
|
||||
t.Cleanup(cancelManager)
|
||||
managerErrors := make(chan error, 1)
|
||||
go func() {
|
||||
managerErrors <- manager.Start(managerContext)
|
||||
}()
|
||||
if !manager.GetCache().WaitForCacheSync(managerContext) {
|
||||
t.Fatal("manager cache did not synchronize")
|
||||
}
|
||||
|
||||
directClient, err := client.New(config, client.Options{Scheme: scheme})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
ctx := context.Background()
|
||||
objects := []client.Object{
|
||||
&corev1.Namespace{ObjectMeta: metav1.ObjectMeta{Name: integrationNamespace, Labels: map[string]string{testLabelKey: testLabelEnabled}}},
|
||||
&corev1.ServiceAccount{ObjectMeta: metav1.ObjectMeta{Name: testSAName, Namespace: integrationNamespace}},
|
||||
&executionv1alpha1.KubernetesExecutionParameters{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: integrationClass},
|
||||
Spec: executionv1alpha1.KubernetesExecutionParametersSpec{
|
||||
ServiceAccountName: testSAName,
|
||||
ImagePullPolicy: corev1.PullIfNotPresent,
|
||||
},
|
||||
},
|
||||
&executionv1alpha1.JobClass{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: integrationClass},
|
||||
Spec: executionv1alpha1.JobClassSpec{
|
||||
ControllerName: kubernetesadapter.ControllerName,
|
||||
ParametersRef: executionv1alpha1.ParametersReference{
|
||||
Group: executionv1alpha1.GroupVersion.Group,
|
||||
Kind: "KubernetesExecutionParameters",
|
||||
Name: integrationClass,
|
||||
},
|
||||
AllowedNamespaces: &metav1.LabelSelector{MatchLabels: map[string]string{testLabelKey: testLabelEnabled}},
|
||||
},
|
||||
},
|
||||
}
|
||||
for _, object := range objects {
|
||||
if err := directClient.Create(ctx, object); err != nil {
|
||||
t.Fatalf("create %T: %v", object, err)
|
||||
}
|
||||
}
|
||||
|
||||
job := &executionv1alpha1.Job{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: testJobName, Namespace: integrationNamespace},
|
||||
Spec: executionv1alpha1.JobSpec{
|
||||
JobClassName: integrationClass,
|
||||
DesiredState: executionv1alpha1.JobDesiredStateRunning,
|
||||
Task: executionv1alpha1.TaskSpec{Image: "alpine:3.22", Command: []string{"true"}},
|
||||
},
|
||||
}
|
||||
if err := directClient.Create(ctx, job); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
backend := &batchv1.Job{}
|
||||
eventually(t, 10*time.Second, func() (bool, error) {
|
||||
err := directClient.Get(ctx, types.NamespacedName{Namespace: job.Namespace, Name: job.Name}, backend)
|
||||
return err == nil, client.IgnoreNotFound(err)
|
||||
})
|
||||
if backend.Labels[kubernetesadapter.JobUIDLabel] != string(job.UID) {
|
||||
t.Fatalf("backend identity label = %q, want %q", backend.Labels[kubernetesadapter.JobUIDLabel], job.UID)
|
||||
}
|
||||
|
||||
eventually(t, 10*time.Second, func() (bool, error) {
|
||||
if err := directClient.Get(ctx, types.NamespacedName{Namespace: job.Namespace, Name: job.Name}, job); err != nil {
|
||||
return false, err
|
||||
}
|
||||
return conditionStatus(job, executionv1alpha1.JobConditionScheduled) == metav1.ConditionTrue, nil
|
||||
})
|
||||
|
||||
completed := metav1.Now()
|
||||
if err := directClient.Get(ctx, types.NamespacedName{Namespace: job.Namespace, Name: job.Name}, backend); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
backend.Status.StartTime = &completed
|
||||
backend.Status.CompletionTime = &completed
|
||||
backend.Status.Conditions = []batchv1.JobCondition{
|
||||
{Type: batchv1.JobSuccessCriteriaMet, Status: corev1.ConditionTrue, Reason: "CompletionsReached"},
|
||||
{Type: batchv1.JobComplete, Status: corev1.ConditionTrue, Reason: "Completed"},
|
||||
}
|
||||
if err := directClient.Status().Update(ctx, backend); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
eventually(t, 10*time.Second, func() (bool, error) {
|
||||
if err := directClient.Get(ctx, types.NamespacedName{Namespace: job.Namespace, Name: job.Name}, job); err != nil {
|
||||
return false, err
|
||||
}
|
||||
return conditionStatus(job, executionv1alpha1.JobConditionSucceeded) == metav1.ConditionTrue, nil
|
||||
})
|
||||
if job.Status.StartTime == nil || job.Status.CompletionTime == nil || job.Status.Execution == nil {
|
||||
t.Fatalf("controller did not persist execution status: %#v", job.Status)
|
||||
}
|
||||
|
||||
cancelManager()
|
||||
select {
|
||||
case err := <-managerErrors:
|
||||
if err != nil {
|
||||
t.Fatalf("manager stopped with error: %v", err)
|
||||
}
|
||||
case <-time.After(5 * time.Second):
|
||||
t.Fatal("manager did not stop")
|
||||
}
|
||||
}
|
||||
|
||||
func eventually(t *testing.T, timeout time.Duration, check func() (bool, error)) {
|
||||
t.Helper()
|
||||
deadline := time.Now().Add(timeout)
|
||||
for time.Now().Before(deadline) {
|
||||
ready, err := check()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if ready {
|
||||
return
|
||||
}
|
||||
time.Sleep(100 * time.Millisecond)
|
||||
}
|
||||
t.Fatal(fmt.Errorf("condition was not met within %s", timeout))
|
||||
}
|
||||
|
||||
func conditionStatus(job *executionv1alpha1.Job, conditionType string) metav1.ConditionStatus {
|
||||
condition := meta.FindStatusCondition(job.Status.Conditions, conditionType)
|
||||
if condition == nil {
|
||||
return metav1.ConditionUnknown
|
||||
}
|
||||
return condition.Status
|
||||
}
|
||||
@@ -1,245 +0,0 @@
|
||||
package controller
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
executionv1alpha1 "git.ddupan.top/panxiao81/ayatori/api/execution/v1alpha1"
|
||||
kubernetesadapter "git.ddupan.top/panxiao81/ayatori/internal/adapter/kubernetes"
|
||||
batchv1 "k8s.io/api/batch/v1"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
"k8s.io/apimachinery/pkg/api/meta"
|
||||
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
|
||||
"k8s.io/apimachinery/pkg/runtime"
|
||||
"k8s.io/apimachinery/pkg/types"
|
||||
ctrl "sigs.k8s.io/controller-runtime"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client/fake"
|
||||
)
|
||||
|
||||
const (
|
||||
defaultClassName = "default"
|
||||
testJobName = "hello"
|
||||
testSAName = "runner"
|
||||
testLabelKey = "execution"
|
||||
testLabelEnabled = "enabled"
|
||||
)
|
||||
|
||||
//nolint:modernize // controller-runtime and Kubernetes API structs expose promoted embedded fields.
|
||||
func TestJobReconcilerKubernetesLifecycle(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
now := time.Unix(1_700_000_000, 0)
|
||||
reconciler, kubeClient := testReconciler(t, now, validObjects()...)
|
||||
request := ctrl.Request{}
|
||||
request.NamespacedName = types.NamespacedName{Namespace: "ci", Name: testJobName}
|
||||
|
||||
if _, err := reconciler.Reconcile(ctx, request); err != nil {
|
||||
t.Fatalf("add finalizer: %v", err)
|
||||
}
|
||||
if _, err := reconciler.Reconcile(ctx, request); err != nil {
|
||||
t.Fatalf("create backend: %v", err)
|
||||
}
|
||||
|
||||
backend := &batchv1.Job{}
|
||||
if err := kubeClient.Get(ctx, request.NamespacedName, backend); err != nil {
|
||||
t.Fatalf("backend Job was not created: %v", err)
|
||||
}
|
||||
if backend.Labels[kubernetesadapter.JobUIDLabel] != "ayatori-job-uid" {
|
||||
t.Fatalf("backend UID label = %q", backend.Labels[kubernetesadapter.JobUIDLabel])
|
||||
}
|
||||
|
||||
job := getJob(t, ctx, kubeClient, request.NamespacedName)
|
||||
if !conditionIs(job, executionv1alpha1.JobConditionAccepted, metav1.ConditionTrue) ||
|
||||
!conditionIs(job, executionv1alpha1.JobConditionScheduled, metav1.ConditionTrue) {
|
||||
t.Fatalf("Job was not accepted and scheduled: %#v", job.Status.Conditions)
|
||||
}
|
||||
|
||||
started := metav1.NewTime(now.Add(time.Minute))
|
||||
backend.Status.StartTime = &started
|
||||
backend.Status.Active = 1
|
||||
if err := kubeClient.Status().Update(ctx, backend); err != nil {
|
||||
t.Fatalf("set backend running: %v", err)
|
||||
}
|
||||
if _, err := reconciler.Reconcile(ctx, request); err != nil {
|
||||
t.Fatalf("observe running backend: %v", err)
|
||||
}
|
||||
job = getJob(t, ctx, kubeClient, request.NamespacedName)
|
||||
if job.Status.StartTime == nil || !conditionIs(job, executionv1alpha1.JobConditionSucceeded, metav1.ConditionUnknown) {
|
||||
t.Fatalf("running state was not observed: %#v", job.Status)
|
||||
}
|
||||
|
||||
completed := metav1.NewTime(now.Add(2 * time.Minute))
|
||||
backend = &batchv1.Job{}
|
||||
if err := kubeClient.Get(ctx, request.NamespacedName, backend); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
backend.Status.Active = 0
|
||||
backend.Status.CompletionTime = &completed
|
||||
backend.Status.Conditions = []batchv1.JobCondition{{Type: batchv1.JobComplete, Status: corev1.ConditionTrue, Reason: "Completed"}}
|
||||
if err := kubeClient.Status().Update(ctx, backend); err != nil {
|
||||
t.Fatalf("set backend complete: %v", err)
|
||||
}
|
||||
if _, err := reconciler.Reconcile(ctx, request); err != nil {
|
||||
t.Fatalf("observe completed backend: %v", err)
|
||||
}
|
||||
job = getJob(t, ctx, kubeClient, request.NamespacedName)
|
||||
if !conditionIs(job, executionv1alpha1.JobConditionSucceeded, metav1.ConditionTrue) || job.Status.CompletionTime == nil {
|
||||
t.Fatalf("terminal state was not observed: %#v", job.Status)
|
||||
}
|
||||
}
|
||||
|
||||
//nolint:modernize // controller-runtime Request exposes NamespacedName as a promoted embedded field.
|
||||
func TestJobReconcilerRejectsMissingClass(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
job := validObjects()[3].(*executionv1alpha1.Job).DeepCopy()
|
||||
job.Spec.JobClassName = "missing"
|
||||
reconciler, kubeClient := testReconciler(t, time.Now(), validObjects()[0], validObjects()[1], job)
|
||||
request := ctrl.Request{}
|
||||
request.NamespacedName = types.NamespacedName{Namespace: job.Namespace, Name: job.Name}
|
||||
|
||||
if _, err := reconciler.Reconcile(ctx, request); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
result, err := reconciler.Reconcile(ctx, request)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if result.RequeueAfter == 0 {
|
||||
t.Fatal("missing JobClass did not schedule a retry")
|
||||
}
|
||||
stored := getJob(t, ctx, kubeClient, request.NamespacedName)
|
||||
accepted := meta.FindStatusCondition(stored.Status.Conditions, executionv1alpha1.JobConditionAccepted)
|
||||
if accepted == nil || accepted.Status != metav1.ConditionFalse || accepted.Reason != "JobClassNotFound" {
|
||||
t.Fatalf("unexpected Accepted condition: %#v", accepted)
|
||||
}
|
||||
}
|
||||
|
||||
//nolint:modernize // controller-runtime Request exposes NamespacedName as a promoted embedded field.
|
||||
func TestJobReconcilerObservesExistingExecutionWithoutJobClass(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
now := time.Unix(1_700_000_000, 0)
|
||||
job := validObjects()[3].(*executionv1alpha1.Job).DeepCopy()
|
||||
job.Finalizers = []string{jobFinalizer}
|
||||
job.Status.Execution = &executionv1alpha1.ExecutionStatus{Adapter: "kubernetes"}
|
||||
backend := kubernetesadapter.BuildJob(job, validObjects()[4].(*executionv1alpha1.KubernetesExecutionParameters), executionv1alpha1.ExecutionResourceRequirements{})
|
||||
backend.Status.StartTime = &metav1.Time{Time: now}
|
||||
backend.Status.Active = 1
|
||||
reconciler, kubeClient := testReconciler(t, now, job, backend)
|
||||
request := ctrl.Request{NamespacedName: types.NamespacedName{Namespace: job.Namespace, Name: job.Name}}
|
||||
|
||||
if _, err := reconciler.Reconcile(ctx, request); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
stored := getJob(t, ctx, kubeClient, request.NamespacedName)
|
||||
if stored.Status.StartTime == nil || !conditionIs(stored, executionv1alpha1.JobConditionSucceeded, metav1.ConditionUnknown) {
|
||||
t.Fatalf("existing execution was not observed without its JobClass: %#v", stored.Status)
|
||||
}
|
||||
}
|
||||
|
||||
//nolint:modernize // controller-runtime Request exposes NamespacedName as a promoted embedded field.
|
||||
func TestJobReconcilerCancelsBeforeScheduling(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
job := validObjects()[3].(*executionv1alpha1.Job).DeepCopy()
|
||||
job.Spec.DesiredState = executionv1alpha1.JobDesiredStateCancelled
|
||||
reconciler, kubeClient := testReconciler(t, time.Unix(1_700_000_000, 0), job)
|
||||
request := ctrl.Request{NamespacedName: types.NamespacedName{Namespace: job.Namespace, Name: job.Name}}
|
||||
|
||||
if _, err := reconciler.Reconcile(ctx, request); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
stored := getJob(t, ctx, kubeClient, request.NamespacedName)
|
||||
condition := meta.FindStatusCondition(stored.Status.Conditions, executionv1alpha1.JobConditionSucceeded)
|
||||
if condition == nil || condition.Status != metav1.ConditionFalse || condition.Reason != "Cancelled" {
|
||||
t.Fatalf("unexpected cancellation condition: %#v", condition)
|
||||
}
|
||||
if stored.Status.CompletionTime == nil {
|
||||
t.Fatal("cancelled Job has no completionTime")
|
||||
}
|
||||
}
|
||||
|
||||
//nolint:modernize // controller-runtime Request exposes NamespacedName as a promoted embedded field.
|
||||
func TestJobReconcilerKeepsConfirmedSuccessDuringCancellation(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
now := time.Unix(1_700_000_000, 0)
|
||||
job := validObjects()[3].(*executionv1alpha1.Job).DeepCopy()
|
||||
job.Spec.DesiredState = executionv1alpha1.JobDesiredStateCancelled
|
||||
job.Finalizers = []string{jobFinalizer}
|
||||
backend := kubernetesadapter.BuildJob(job, validObjects()[4].(*executionv1alpha1.KubernetesExecutionParameters), executionv1alpha1.ExecutionResourceRequirements{})
|
||||
backend.Status.CompletionTime = &metav1.Time{Time: now}
|
||||
backend.Status.Conditions = []batchv1.JobCondition{{Type: batchv1.JobComplete, Status: corev1.ConditionTrue}}
|
||||
reconciler, kubeClient := testReconciler(t, now, job, backend)
|
||||
request := ctrl.Request{NamespacedName: types.NamespacedName{Namespace: job.Namespace, Name: job.Name}}
|
||||
|
||||
if _, err := reconciler.Reconcile(ctx, request); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
stored := getJob(t, ctx, kubeClient, request.NamespacedName)
|
||||
if !conditionIs(stored, executionv1alpha1.JobConditionSucceeded, metav1.ConditionTrue) {
|
||||
t.Fatalf("confirmed success was overwritten by cancellation: %#v", stored.Status.Conditions)
|
||||
}
|
||||
if err := kubeClient.Get(ctx, request.NamespacedName, &batchv1.Job{}); err != nil {
|
||||
t.Fatalf("successful backend was deleted: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
//nolint:modernize // Kubernetes API structs expose ObjectMeta through embedded TypeMeta fields.
|
||||
func validObjects() []client.Object {
|
||||
return []client.Object{
|
||||
&corev1.Namespace{ObjectMeta: metav1.ObjectMeta{Name: "ci", Labels: map[string]string{testLabelKey: testLabelEnabled}}},
|
||||
&corev1.ServiceAccount{ObjectMeta: metav1.ObjectMeta{Name: testSAName, Namespace: "ci"}},
|
||||
&executionv1alpha1.JobClass{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: defaultClassName, UID: types.UID("class-uid")},
|
||||
Spec: executionv1alpha1.JobClassSpec{
|
||||
ControllerName: kubernetesadapter.ControllerName,
|
||||
ParametersRef: executionv1alpha1.ParametersReference{Group: executionv1alpha1.GroupVersion.Group, Kind: "KubernetesExecutionParameters", Name: defaultClassName},
|
||||
AllowedNamespaces: &metav1.LabelSelector{MatchLabels: map[string]string{testLabelKey: testLabelEnabled}},
|
||||
},
|
||||
},
|
||||
&executionv1alpha1.Job{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: testJobName, Namespace: "ci", UID: types.UID("ayatori-job-uid")},
|
||||
Spec: executionv1alpha1.JobSpec{
|
||||
JobClassName: defaultClassName, DesiredState: executionv1alpha1.JobDesiredStateRunning,
|
||||
Task: executionv1alpha1.TaskSpec{Image: "alpine:3.22", Command: []string{"true"}},
|
||||
},
|
||||
},
|
||||
&executionv1alpha1.KubernetesExecutionParameters{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: defaultClassName, UID: types.UID("parameters-uid")},
|
||||
Spec: executionv1alpha1.KubernetesExecutionParametersSpec{ServiceAccountName: testSAName, ImagePullPolicy: corev1.PullIfNotPresent},
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
func testReconciler(t *testing.T, now time.Time, objects ...client.Object) (*JobReconciler, client.Client) {
|
||||
t.Helper()
|
||||
scheme := runtime.NewScheme()
|
||||
if err := corev1.AddToScheme(scheme); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := batchv1.AddToScheme(scheme); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := executionv1alpha1.AddToScheme(scheme); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
kubeClient := fake.NewClientBuilder().
|
||||
WithScheme(scheme).
|
||||
WithStatusSubresource(&executionv1alpha1.Job{}, &batchv1.Job{}).
|
||||
WithObjects(objects...).
|
||||
Build()
|
||||
return &JobReconciler{Client: kubeClient, Now: func() time.Time { return now }}, kubeClient
|
||||
}
|
||||
|
||||
func getJob(t *testing.T, ctx context.Context, kubeClient client.Client, key types.NamespacedName) *executionv1alpha1.Job {
|
||||
t.Helper()
|
||||
job := &executionv1alpha1.Job{}
|
||||
if err := kubeClient.Get(ctx, key, job); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return job
|
||||
}
|
||||
|
||||
func conditionIs(job *executionv1alpha1.Job, conditionType string, status metav1.ConditionStatus) bool {
|
||||
condition := meta.FindStatusCondition(job.Status.Conditions, conditionType)
|
||||
return condition != nil && condition.Status == status
|
||||
}
|
||||
@@ -38,22 +38,22 @@ const (
|
||||
)
|
||||
|
||||
// Snapshot contains persisted observations only, without credentials or live evidence.
|
||||
// Failure detail mapping will be added with capability assessment, not intent transitions.
|
||||
type Snapshot struct {
|
||||
Phase Phase
|
||||
ObservedRevision int64
|
||||
Readiness Readiness
|
||||
ReportedVersion string
|
||||
Failure Failure
|
||||
}
|
||||
|
||||
// Instance protects registration state and pure lifecycle transitions.
|
||||
// Reconstitution does not establish live capability evidence, even for a Ready snapshot.
|
||||
// This initial slice deliberately exposes no operation that authorizes provisioning.
|
||||
type Instance struct {
|
||||
target ObservationTarget
|
||||
snapshot Snapshot
|
||||
deleting bool
|
||||
extensions ExtensionSupport
|
||||
evidence *CapabilityObservation
|
||||
}
|
||||
|
||||
func Reconstitute(target ObservationTarget, snapshot Snapshot, deleting bool) (*Instance, error) {
|
||||
@@ -85,6 +85,8 @@ func (i *Instance) BeginValidation() error {
|
||||
i.snapshot.Phase = PhaseValidating
|
||||
i.snapshot.Readiness = Unknown
|
||||
i.extensions = ExtensionSupport{}
|
||||
i.evidence = nil
|
||||
i.snapshot.Failure = NoFailure
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -100,6 +102,8 @@ func (i *Instance) BeginDeletion() error {
|
||||
i.snapshot.Phase = PhaseDeleting
|
||||
i.snapshot.Readiness = Unknown
|
||||
i.extensions = ExtensionSupport{}
|
||||
i.evidence = nil
|
||||
i.snapshot.Failure = NoFailure
|
||||
return nil
|
||||
}
|
||||
|
||||
|
||||
@@ -29,10 +29,10 @@ func TestInstanceAcceptsExtensionObservationForCurrentTarget(t *testing.T) {
|
||||
target := value.Target()
|
||||
snapshot := value.Snapshot()
|
||||
|
||||
if err := value.ObserveExtensions(target, instance.ObserveExtensionSupport([]string{"pg_trgm"})); err != nil {
|
||||
if err := value.ObserveExtensions(target, instance.ObserveExtensionSupport([]string{testTrigram})); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got := value.CheckExtensions(instance.NewExtensionSet([]string{"pg_trgm"})); got.Decision != instance.ExtensionsAccepted {
|
||||
if got := value.CheckExtensions(instance.NewExtensionSet([]string{testTrigram})); got.Decision != instance.ExtensionsAccepted {
|
||||
t.Fatalf("CheckExtensions() = %v, want accepted", got)
|
||||
}
|
||||
if value.Snapshot() != snapshot {
|
||||
@@ -52,20 +52,20 @@ func TestInstanceRejectsExtensionObservationForDifferentTarget(t *testing.T) {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if err := value.ObserveExtensions(different, instance.ObserveExtensionSupport([]string{"pg_trgm"})); err == nil {
|
||||
if err := value.ObserveExtensions(different, instance.ObserveExtensionSupport([]string{testTrigram})); err == nil {
|
||||
t.Fatal("observation for a different target was accepted")
|
||||
}
|
||||
if got := value.CheckExtensions(instance.NewExtensionSet([]string{"pg_trgm"})); got.Decision != instance.ExtensionSupportUnobserved {
|
||||
if got := value.CheckExtensions(instance.NewExtensionSet([]string{testTrigram})); got.Decision != instance.ExtensionSupportUnobserved {
|
||||
t.Fatalf("rejected observation changed support: %v", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestInstanceClearsExtensionObservationAcrossLifecycleBoundaries(t *testing.T) {
|
||||
requested := instance.NewExtensionSet([]string{"pg_trgm"})
|
||||
requested := instance.NewExtensionSet([]string{testTrigram})
|
||||
|
||||
t.Run("validation", func(t *testing.T) {
|
||||
value := lifecycleInstance(t, instance.Snapshot{Phase: instance.PhaseReady}, false)
|
||||
if err := value.ObserveExtensions(value.Target(), instance.ObserveExtensionSupport([]string{"pg_trgm"})); err != nil {
|
||||
if err := value.ObserveExtensions(value.Target(), instance.ObserveExtensionSupport([]string{testTrigram})); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := value.BeginValidation(); err != nil {
|
||||
@@ -78,7 +78,7 @@ func TestInstanceClearsExtensionObservationAcrossLifecycleBoundaries(t *testing.
|
||||
|
||||
t.Run("deletion", func(t *testing.T) {
|
||||
value := lifecycleInstance(t, instance.Snapshot{Phase: instance.PhaseReady}, true)
|
||||
if err := value.ObserveExtensions(value.Target(), instance.ObserveExtensionSupport([]string{"pg_trgm"})); err == nil {
|
||||
if err := value.ObserveExtensions(value.Target(), instance.ObserveExtensionSupport([]string{testTrigram})); err == nil {
|
||||
t.Fatal("deleting instance accepted a new observation")
|
||||
}
|
||||
if err := value.BeginDeletion(); err != nil {
|
||||
@@ -92,8 +92,8 @@ func TestInstanceClearsExtensionObservationAcrossLifecycleBoundaries(t *testing.
|
||||
|
||||
func TestInstanceCanExplicitlyInvalidateExtensionObservation(t *testing.T) {
|
||||
value := lifecycleInstance(t, instance.Snapshot{Phase: instance.PhaseValidating}, false)
|
||||
requested := instance.NewExtensionSet([]string{"pg_trgm"})
|
||||
if err := value.ObserveExtensions(value.Target(), instance.ObserveExtensionSupport([]string{"pg_trgm"})); err != nil {
|
||||
requested := instance.NewExtensionSet([]string{testTrigram})
|
||||
if err := value.ObserveExtensions(value.Target(), instance.ObserveExtensionSupport([]string{testTrigram})); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := value.ObserveExtensions(value.Target(), instance.ExtensionSupport{}); err != nil {
|
||||
|
||||
@@ -0,0 +1,263 @@
|
||||
/*
|
||||
Copyright 2026.
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
*/
|
||||
|
||||
package instance
|
||||
|
||||
import "errors"
|
||||
|
||||
// Failure 只表示安全类别;驱动错误、凭据和 Condition 文案留在应用边界。
|
||||
type Failure uint8
|
||||
|
||||
const (
|
||||
NoFailure Failure = iota
|
||||
ObservationIncomplete
|
||||
DependencyUnavailable
|
||||
AuthenticationFailed
|
||||
InsufficientPrivileges
|
||||
RegistryIncompatible
|
||||
RegistryNotUsable
|
||||
)
|
||||
|
||||
// CheckResult 的零值表示未观察,不能视为成功。
|
||||
type CheckResult uint8
|
||||
|
||||
const (
|
||||
CheckUnobserved CheckResult = iota
|
||||
CheckPassed
|
||||
CheckUnavailable
|
||||
CheckAuthenticationFailed
|
||||
CheckInsufficientPrivileges
|
||||
)
|
||||
|
||||
// ManagementChecks 分别记录所需能力;SQL 探测和同轮次关联由 adapter/application 保证。
|
||||
// Extensions 不代表任意扩展均可安装;具体请求仍需支持检查、执行及回读。
|
||||
type ManagementChecks struct {
|
||||
Connection CheckResult
|
||||
Metadata CheckResult
|
||||
Roles CheckResult
|
||||
Databases CheckResult
|
||||
Grants CheckResult
|
||||
Extensions CheckResult
|
||||
}
|
||||
|
||||
func (c ManagementChecks) failure() Failure {
|
||||
for _, check := range []CheckResult{c.Connection, c.Metadata, c.Roles, c.Databases, c.Grants, c.Extensions} {
|
||||
switch check {
|
||||
case CheckPassed:
|
||||
case CheckUnavailable:
|
||||
return DependencyUnavailable
|
||||
case CheckAuthenticationFailed:
|
||||
return AuthenticationFailed
|
||||
case CheckInsufficientPrivileges:
|
||||
return InsufficientPrivileges
|
||||
default:
|
||||
return ObservationIncomplete
|
||||
}
|
||||
}
|
||||
return NoFailure
|
||||
}
|
||||
|
||||
type RegistryState uint8
|
||||
|
||||
const (
|
||||
RegistryUnobserved RegistryState = iota
|
||||
RegistryAbsent
|
||||
RegistryNeedsMigration
|
||||
RegistryUsable
|
||||
RegistryUnsupported
|
||||
RegistryUnavailable
|
||||
)
|
||||
|
||||
// CapabilityObservation 是值对象,不包含连接、凭据或可变集合。
|
||||
type CapabilityObservation struct {
|
||||
target ObservationTarget
|
||||
version string
|
||||
checks ManagementChecks
|
||||
registry RegistryState
|
||||
}
|
||||
|
||||
func NewCapabilityObservation(target ObservationTarget, version string,
|
||||
checks ManagementChecks, registry RegistryState,
|
||||
) (CapabilityObservation, error) {
|
||||
if err := target.Validate(); err != nil {
|
||||
return CapabilityObservation{}, err
|
||||
}
|
||||
return CapabilityObservation{target: target, version: version, checks: checks, registry: registry}, nil
|
||||
}
|
||||
|
||||
func (o CapabilityObservation) managementFailure() Failure {
|
||||
if failure := o.checks.failure(); failure != NoFailure {
|
||||
return failure
|
||||
}
|
||||
if o.version == "" {
|
||||
return ObservationIncomplete
|
||||
}
|
||||
return NoFailure
|
||||
}
|
||||
|
||||
func (o CapabilityObservation) registryFailure() Failure {
|
||||
switch o.registry {
|
||||
case RegistryUsable:
|
||||
return NoFailure
|
||||
case RegistryAbsent, RegistryNeedsMigration:
|
||||
return RegistryNotUsable
|
||||
case RegistryUnsupported:
|
||||
return RegistryIncompatible
|
||||
case RegistryUnavailable:
|
||||
return DependencyUnavailable
|
||||
default:
|
||||
return ObservationIncomplete
|
||||
}
|
||||
}
|
||||
|
||||
type PreparationDecision uint8
|
||||
|
||||
const (
|
||||
PreparationDenied PreparationDecision = iota
|
||||
PreparationAllowed
|
||||
AlreadyUsable
|
||||
)
|
||||
|
||||
func (i *Instance) acceptObservation(o CapabilityObservation, phase Phase) error {
|
||||
if !i.target.Matches(o.target) {
|
||||
return errors.New("capability observation target does not match instance")
|
||||
}
|
||||
if i.deleting || i.snapshot.Phase != phase {
|
||||
return errors.New("capability observation is not allowed in current lifecycle")
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (i *Instance) fail(failure Failure) {
|
||||
i.evidence = nil
|
||||
i.extensions = ExtensionSupport{}
|
||||
i.snapshot.Readiness = NotReady
|
||||
i.snapshot.Failure = failure
|
||||
i.snapshot.ObservedRevision = i.target.Revision().Value()
|
||||
}
|
||||
|
||||
// AssessManagement 只推进意图,不执行 registry 写入,也不完成 observedRevision。
|
||||
func (i *Instance) AssessManagement(o CapabilityObservation) error {
|
||||
if err := i.acceptObservation(o, PhaseValidating); err != nil {
|
||||
return err
|
||||
}
|
||||
if failure := o.managementFailure(); failure != NoFailure {
|
||||
i.fail(failure)
|
||||
return nil
|
||||
}
|
||||
if failure := o.registryFailure(); failure != NoFailure && failure != RegistryNotUsable {
|
||||
i.fail(failure)
|
||||
return nil
|
||||
}
|
||||
i.snapshot.Phase = PhaseInitializingRegistry
|
||||
i.snapshot.Readiness = Unknown
|
||||
i.snapshot.Failure = NoFailure
|
||||
i.evidence = nil
|
||||
return nil
|
||||
}
|
||||
|
||||
// PlanRegistryPreparation 不证明 checkpoint 已落盘;应用层必须先保存意图再执行写入。
|
||||
func (i *Instance) PlanRegistryPreparation(o CapabilityObservation) (PreparationDecision, error) {
|
||||
if err := i.acceptObservation(o, PhaseInitializingRegistry); err != nil {
|
||||
return PreparationDenied, err
|
||||
}
|
||||
if failure := o.managementFailure(); failure != NoFailure {
|
||||
i.fail(failure)
|
||||
return PreparationDenied, nil
|
||||
}
|
||||
switch o.registry {
|
||||
case RegistryUsable:
|
||||
return AlreadyUsable, nil
|
||||
case RegistryAbsent, RegistryNeedsMigration:
|
||||
return PreparationAllowed, nil
|
||||
default:
|
||||
i.fail(o.registryFailure())
|
||||
return PreparationDenied, nil
|
||||
}
|
||||
}
|
||||
|
||||
// RegistryPreparationResult 只能是安全失败或完整回读,不能表达裸操作成功。
|
||||
type RegistryPreparationResult struct {
|
||||
observation CapabilityObservation
|
||||
failure Failure
|
||||
}
|
||||
|
||||
func RegistryReadBack(o CapabilityObservation) RegistryPreparationResult {
|
||||
return RegistryPreparationResult{observation: o}
|
||||
}
|
||||
|
||||
func RegistryPreparationFailed(target ObservationTarget, failure Failure) (RegistryPreparationResult, error) {
|
||||
if err := target.Validate(); err != nil {
|
||||
return RegistryPreparationResult{}, err
|
||||
}
|
||||
if failure < ObservationIncomplete || failure > RegistryNotUsable {
|
||||
return RegistryPreparationResult{}, errors.New("registry preparation requires a known failure category")
|
||||
}
|
||||
return RegistryPreparationResult{observation: CapabilityObservation{target: target}, failure: failure}, nil
|
||||
}
|
||||
|
||||
func (i *Instance) AssessRegistryResult(result RegistryPreparationResult) error {
|
||||
o := result.observation
|
||||
if err := i.acceptObservation(o, PhaseInitializingRegistry); err != nil {
|
||||
return err
|
||||
}
|
||||
if result.failure != NoFailure {
|
||||
i.fail(result.failure)
|
||||
return nil
|
||||
}
|
||||
i.assessComplete(o)
|
||||
return nil
|
||||
}
|
||||
|
||||
func (i *Instance) assessComplete(o CapabilityObservation) {
|
||||
if failure := o.managementFailure(); failure != NoFailure {
|
||||
i.fail(failure)
|
||||
return
|
||||
}
|
||||
if failure := o.registryFailure(); failure != NoFailure {
|
||||
i.fail(failure)
|
||||
return
|
||||
}
|
||||
i.snapshot = Snapshot{Phase: PhaseReady, ObservedRevision: i.target.Revision().Value(),
|
||||
Readiness: Ready, ReportedVersion: o.version}
|
||||
i.evidence = &o
|
||||
}
|
||||
|
||||
// AssessReadiness 每轮接收完整事实,失败立即撤销本轮供应能力。
|
||||
func (i *Instance) AssessReadiness(o CapabilityObservation) error {
|
||||
if err := i.acceptObservation(o, PhaseReady); err != nil {
|
||||
return err
|
||||
}
|
||||
if i.snapshot.ObservedRevision != i.target.Revision().Value() {
|
||||
return i.BeginValidation()
|
||||
}
|
||||
if o.managementFailure() != NoFailure || o.registry == RegistryUnavailable {
|
||||
i.snapshot.Phase = PhaseValidating
|
||||
} else if o.registryFailure() != NoFailure {
|
||||
i.snapshot.Phase = PhaseInitializingRegistry
|
||||
}
|
||||
i.assessComplete(o)
|
||||
return nil
|
||||
}
|
||||
|
||||
// RequireProvisioningReady 仅检查 Instance 前置条件,不授予 Tenant 所有权或外部写入许可。
|
||||
func (i *Instance) RequireProvisioningReady() error {
|
||||
if i.deleting || i.evidence == nil || i.snapshot.Phase != PhaseReady || i.snapshot.Readiness != Ready ||
|
||||
i.snapshot.ObservedRevision != i.target.Revision().Value() {
|
||||
return errors.New("instance is not ready for provisioning")
|
||||
}
|
||||
return nil
|
||||
}
|
||||
@@ -0,0 +1,352 @@
|
||||
/*
|
||||
Copyright 2026.
|
||||
|
||||
Licensed under the Apache License, Version 2.0 (the "License");
|
||||
you may not use this file except in compliance with the License.
|
||||
You may obtain a copy of the License at
|
||||
|
||||
http://www.apache.org/licenses/LICENSE-2.0
|
||||
|
||||
Unless required by applicable law or agreed to in writing, software
|
||||
distributed under the License is distributed on an "AS IS" BASIS,
|
||||
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
See the License for the specific language governing permissions and
|
||||
limitations under the License.
|
||||
*/
|
||||
|
||||
package instance_test
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"git.ddupan.top/panxiao81/ayatori/internal/database/domain/instance"
|
||||
)
|
||||
|
||||
const testServerVersion = "17.6"
|
||||
|
||||
func completeChecks() instance.ManagementChecks {
|
||||
return instance.ManagementChecks{
|
||||
Connection: instance.CheckPassed, Metadata: instance.CheckPassed,
|
||||
Roles: instance.CheckPassed, Databases: instance.CheckPassed,
|
||||
Grants: instance.CheckPassed, Extensions: instance.CheckPassed,
|
||||
}
|
||||
}
|
||||
|
||||
func capability(t *testing.T, value *instance.Instance, checks instance.ManagementChecks,
|
||||
registry instance.RegistryState,
|
||||
) instance.CapabilityObservation {
|
||||
t.Helper()
|
||||
o, err := instance.NewCapabilityObservation(value.Target(), testServerVersion, checks, registry)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return o
|
||||
}
|
||||
|
||||
func readyInstance(t *testing.T) *instance.Instance {
|
||||
t.Helper()
|
||||
i := lifecycleInstance(t, instance.Snapshot{Phase: instance.PhaseInitializingRegistry}, false)
|
||||
if err := i.AssessRegistryResult(instance.RegistryReadBack(capability(t, i, completeChecks(), instance.RegistryUsable))); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := i.RequireProvisioningReady(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return i
|
||||
}
|
||||
|
||||
func TestReadinessRequiresCompleteReadBack(t *testing.T) {
|
||||
i := lifecycleInstance(t, instance.Snapshot{}, false)
|
||||
if err := i.BeginValidation(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
absent := capability(t, i, completeChecks(), instance.RegistryAbsent)
|
||||
if err := i.AssessManagement(absent); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if s := i.Snapshot(); s.Phase != instance.PhaseInitializingRegistry || s.ObservedRevision != 0 || s.Readiness != instance.Unknown {
|
||||
t.Fatalf("management observation prematurely concluded readiness: %+v", s)
|
||||
}
|
||||
for range 2 {
|
||||
decision, err := i.PlanRegistryPreparation(absent)
|
||||
if err != nil || decision != instance.PreparationAllowed {
|
||||
t.Fatalf("preparation: %v, %v", decision, err)
|
||||
}
|
||||
if i.RequireProvisioningReady() == nil {
|
||||
t.Fatal("preparation authorized provisioning")
|
||||
}
|
||||
}
|
||||
if err := i.AssessRegistryResult(instance.RegistryReadBack(absent)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if i.Snapshot().Failure != instance.RegistryNotUsable || i.RequireProvisioningReady() == nil {
|
||||
t.Fatal("absent registry accepted as ready")
|
||||
}
|
||||
usable := capability(t, i, completeChecks(), instance.RegistryUsable)
|
||||
decision, err := i.PlanRegistryPreparation(usable)
|
||||
if err != nil || decision != instance.AlreadyUsable {
|
||||
t.Fatalf("retry after external preparation: %v, %v", decision, err)
|
||||
}
|
||||
if err := i.AssessRegistryResult(instance.RegistryReadBack(usable)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if s := i.Snapshot(); s.Readiness != instance.Ready || s.ReportedVersion != testServerVersion || s.ObservedRevision != i.Target().Revision().Value() {
|
||||
t.Fatalf("complete observation not accepted: %+v", s)
|
||||
}
|
||||
if err := i.RequireProvisioningReady(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
// 重启只恢复 checkpoint;依赖稍后恢复时必须重新取得完整事实。
|
||||
func TestReadinessRecoveryAndInvalidation(t *testing.T) {
|
||||
i := readyInstance(t)
|
||||
restored, err := instance.Reconstitute(i.Target(), i.Snapshot(), false)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if restored.RequireProvisioningReady() == nil {
|
||||
t.Fatal("persisted Ready fabricated fresh evidence")
|
||||
}
|
||||
for range 2 {
|
||||
if err := restored.AssessReadiness(capability(t, restored, completeChecks(), instance.RegistryUsable)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := restored.RequireProvisioningReady(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := restored.BeginValidation(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if restored.RequireProvisioningReady() == nil {
|
||||
t.Fatal("validation retained evidence")
|
||||
}
|
||||
deleted, err := instance.Reconstitute(i.Target(), i.Snapshot(), true)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if deleted.RequireProvisioningReady() == nil {
|
||||
t.Fatal("deletion allowed provisioning")
|
||||
}
|
||||
if err := deleted.BeginDeletion(); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if deleted.RequireProvisioningReady() == nil {
|
||||
t.Fatal("deleting checkpoint allowed provisioning")
|
||||
}
|
||||
}
|
||||
|
||||
func TestEachManagementCheckIsRequired(t *testing.T) {
|
||||
for field := range 6 {
|
||||
for _, result := range []instance.CheckResult{instance.CheckUnobserved, instance.CheckUnavailable,
|
||||
instance.CheckAuthenticationFailed, instance.CheckInsufficientPrivileges, 255} {
|
||||
checks := completeChecks()
|
||||
fields := []*instance.CheckResult{&checks.Connection, &checks.Metadata, &checks.Roles,
|
||||
&checks.Databases, &checks.Grants, &checks.Extensions}
|
||||
*fields[field] = result
|
||||
i := readyInstance(t)
|
||||
if err := i.AssessReadiness(capability(t, i, checks, instance.RegistryUsable)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if s := i.Snapshot(); s.Phase != instance.PhaseValidating || s.Readiness != instance.NotReady ||
|
||||
s.Failure == instance.NoFailure || i.RequireProvisioningReady() == nil {
|
||||
t.Fatalf("check %d result %d accepted: %+v", field, result, s)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistryDecisionsAndReadinessLoss(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
state instance.RegistryState
|
||||
decision instance.PreparationDecision
|
||||
failure instance.Failure
|
||||
}{
|
||||
{instance.RegistryUsable, instance.AlreadyUsable, instance.NoFailure},
|
||||
{instance.RegistryAbsent, instance.PreparationAllowed, instance.RegistryNotUsable},
|
||||
{instance.RegistryNeedsMigration, instance.PreparationAllowed, instance.RegistryNotUsable},
|
||||
{instance.RegistryUnsupported, instance.PreparationDenied, instance.RegistryIncompatible},
|
||||
{instance.RegistryUnavailable, instance.PreparationDenied, instance.DependencyUnavailable},
|
||||
{instance.RegistryUnobserved, instance.PreparationDenied, instance.ObservationIncomplete},
|
||||
{255, instance.PreparationDenied, instance.ObservationIncomplete},
|
||||
} {
|
||||
i := lifecycleInstance(t, instance.Snapshot{Phase: instance.PhaseInitializingRegistry}, false)
|
||||
o := capability(t, i, completeChecks(), tc.state)
|
||||
decision, err := i.PlanRegistryPreparation(o)
|
||||
if err != nil || decision != tc.decision {
|
||||
t.Fatalf("registry %d: %v, %v", tc.state, decision, err)
|
||||
}
|
||||
i = readyInstance(t)
|
||||
if err := i.AssessReadiness(o); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if i.Snapshot().Failure != tc.failure {
|
||||
t.Fatalf("registry %d: %+v", tc.state, i.Snapshot())
|
||||
}
|
||||
if tc.state != instance.RegistryUsable {
|
||||
wantPhase := instance.PhaseInitializingRegistry
|
||||
if tc.state == instance.RegistryUnavailable {
|
||||
wantPhase = instance.PhaseValidating
|
||||
}
|
||||
if i.Snapshot().Phase != wantPhase || i.RequireProvisioningReady() == nil {
|
||||
t.Fatal("registry drift retained readiness")
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestInitializationRejectsIncompleteOrFailedManagement(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
checks instance.ManagementChecks
|
||||
registry instance.RegistryState
|
||||
failure instance.Failure
|
||||
}{
|
||||
{instance.ManagementChecks{}, instance.RegistryUsable, instance.ObservationIncomplete},
|
||||
{completeChecks(), instance.RegistryUnsupported, instance.RegistryIncompatible},
|
||||
{completeChecks(), instance.RegistryUnavailable, instance.DependencyUnavailable},
|
||||
} {
|
||||
i := lifecycleInstance(t, instance.Snapshot{Phase: instance.PhaseValidating}, false)
|
||||
if err := i.AssessManagement(capability(t, i, tc.checks, tc.registry)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if s := i.Snapshot(); s.Phase != instance.PhaseValidating || s.Failure != tc.failure ||
|
||||
s.ObservedRevision != i.Target().Revision().Value() || s.Readiness != instance.NotReady {
|
||||
t.Fatalf("invalid management accepted: %+v", s)
|
||||
}
|
||||
}
|
||||
i := lifecycleInstance(t, instance.Snapshot{Phase: instance.PhaseInitializingRegistry}, false)
|
||||
decision, err := i.PlanRegistryPreparation(capability(t, i, instance.ManagementChecks{}, instance.RegistryAbsent))
|
||||
if err != nil || decision != instance.PreparationDenied || i.Snapshot().Failure != instance.ObservationIncomplete {
|
||||
t.Fatal("incomplete management allowed registry writes")
|
||||
}
|
||||
if err := i.AssessRegistryResult(instance.RegistryReadBack(capability(t, i, completeChecks(), instance.RegistryUsable))); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := i.RequireProvisioningReady(); err != nil {
|
||||
t.Fatal("dependency recovery did not restore readiness", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestReadinessMethodsRejectWrongPhaseAndDeletion(t *testing.T) {
|
||||
for _, deleting := range []bool{false, true} {
|
||||
for _, phase := range []instance.Phase{instance.PhasePending, instance.PhaseValidating,
|
||||
instance.PhaseInitializingRegistry, instance.PhaseReady, instance.PhaseDeleting} {
|
||||
for _, operation := range []struct {
|
||||
phase instance.Phase
|
||||
apply func(*instance.Instance, instance.CapabilityObservation) error
|
||||
}{
|
||||
{instance.PhaseValidating, (*instance.Instance).AssessManagement},
|
||||
{instance.PhaseReady, (*instance.Instance).AssessReadiness},
|
||||
{instance.PhaseInitializingRegistry, func(i *instance.Instance, o instance.CapabilityObservation) error {
|
||||
_, err := i.PlanRegistryPreparation(o)
|
||||
return err
|
||||
}},
|
||||
{instance.PhaseInitializingRegistry, func(i *instance.Instance, o instance.CapabilityObservation) error {
|
||||
return i.AssessRegistryResult(instance.RegistryReadBack(o))
|
||||
}},
|
||||
} {
|
||||
if !deleting && operation.phase == phase {
|
||||
continue
|
||||
}
|
||||
i := lifecycleInstance(t, instance.Snapshot{Phase: phase}, deleting)
|
||||
before := i.Snapshot()
|
||||
if err := operation.apply(i, capability(t, i, completeChecks(), instance.RegistryUsable)); err == nil {
|
||||
t.Fatalf("phase %s deleting=%t accepted operation for %s", phase, deleting, operation.phase)
|
||||
}
|
||||
if i.Snapshot() != before {
|
||||
t.Fatal("rejected operation mutated snapshot")
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestOldGenerationObservationDoesNotReplaceEvidence(t *testing.T) {
|
||||
i := readyInstance(t)
|
||||
target := i.Target()
|
||||
revision, err := instance.NewRevision(target.Revision().Value() + 1)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
other, err := instance.NewObservationTarget(target.Identity(), revision, target.Definition())
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
o, err := instance.NewCapabilityObservation(other, testServerVersion, completeChecks(), instance.RegistryUsable)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
before := i.Snapshot()
|
||||
if err := i.AssessReadiness(o); err == nil || i.Snapshot() != before {
|
||||
t.Fatal("mismatched generation observation was accepted")
|
||||
}
|
||||
if err := i.RequireProvisioningReady(); err != nil {
|
||||
t.Fatal("rejected unrelated input changed previously accepted evidence", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestPreparationFailureCannotEstablishReadiness(t *testing.T) {
|
||||
i := lifecycleInstance(t, instance.Snapshot{Phase: instance.PhaseInitializingRegistry}, false)
|
||||
for _, failure := range []instance.Failure{instance.DependencyUnavailable, instance.AuthenticationFailed,
|
||||
instance.InsufficientPrivileges, instance.RegistryIncompatible} {
|
||||
result, err := instance.RegistryPreparationFailed(i.Target(), failure)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := i.AssessRegistryResult(result); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if s := i.Snapshot(); s.Failure != failure || s.Readiness != instance.NotReady ||
|
||||
s.Phase != instance.PhaseInitializingRegistry || i.RequireProvisioningReady() == nil {
|
||||
t.Fatalf("failed operation accepted: %+v", s)
|
||||
}
|
||||
}
|
||||
for _, failure := range []instance.Failure{instance.NoFailure, 255} {
|
||||
if _, err := instance.RegistryPreparationFailed(i.Target(), failure); err == nil {
|
||||
t.Fatal("invalid failure accepted")
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestCapabilityInputsAndLifecycleGuards(t *testing.T) {
|
||||
i := readyInstance(t)
|
||||
if _, err := instance.NewCapabilityObservation(instance.ObservationTarget{}, testServerVersion,
|
||||
completeChecks(), instance.RegistryUsable); err == nil {
|
||||
t.Fatal("invalid target accepted")
|
||||
}
|
||||
if _, err := instance.RegistryPreparationFailed(instance.ObservationTarget{}, instance.DependencyUnavailable); err == nil {
|
||||
t.Fatal("invalid failure target accepted")
|
||||
}
|
||||
for _, method := range []func(instance.CapabilityObservation) error{
|
||||
i.AssessManagement, i.AssessReadiness,
|
||||
func(o instance.CapabilityObservation) error { _, err := i.PlanRegistryPreparation(o); return err },
|
||||
func(o instance.CapabilityObservation) error {
|
||||
return i.AssessRegistryResult(instance.RegistryReadBack(o))
|
||||
},
|
||||
} {
|
||||
before := i.Snapshot()
|
||||
if err := method(instance.CapabilityObservation{}); err == nil || i.Snapshot() != before {
|
||||
t.Fatal("mismatched observation accepted or mutated state")
|
||||
}
|
||||
}
|
||||
old := i.Snapshot()
|
||||
old.ObservedRevision = 0
|
||||
changed := lifecycleInstance(t, old, false)
|
||||
if err := changed.AssessReadiness(capability(t, changed, completeChecks(), instance.RegistryUsable)); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if s := changed.Snapshot(); s.Phase != instance.PhaseValidating || s.ObservedRevision != 0 || s.Readiness != instance.Unknown {
|
||||
t.Fatalf("changed generation accepted old checkpoint: %+v", s)
|
||||
}
|
||||
o, err := instance.NewCapabilityObservation(i.Target(), "", completeChecks(), instance.RegistryUsable)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if err := i.AssessReadiness(o); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if i.Snapshot().Failure != instance.ObservationIncomplete {
|
||||
t.Fatal("missing version accepted")
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user