Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
221 changes: 208 additions & 13 deletions docs/code-style/test_automation.md

Large diffs are not rendered by default.

61 changes: 42 additions & 19 deletions test/orchestrator/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,10 +18,12 @@ The runner separates product execution from verification:
| `verify` | Verify existing state; do not execute the product playbook |
| `test` | Execute the selected tag once, then verify only if execution succeeds |

Most verification is observational. Tests that create workloads, download
images, drain scheduler nodes, reboot machines, or delete deployment state
are protected by explicit markers. Cleanup is excluded from every implicit
full-lifecycle run and must be selected by name.
Most verification is observational. Tests that reboot machines run only when
`--marker` selects `reboot`, and node remove/add tests only when it selects
`node_lifecycle`. Other markers narrow test selection but do not authorize
workloads or downloads. Destructive tests retain their separate opt-in, and
cleanup is excluded from every implicit full-lifecycle run and must be selected
by name.

## Prerequisites

Expand Down Expand Up @@ -138,7 +140,8 @@ without putting secrets in process arguments:
}
```

LDAP identity fields are required by the selected Slurm/PAM LDAP tests. The
LDAP identity fields are required by the selected Slurm/PAM LDAP tests, which
skip unless `validate_external_ldap: true`. The
external bind secret is additionally required when
`validate_external_ldap: true` and `configure_external_ldap: true`. Product-
domain credentials continue to use `--domain-creds-stdin`.
Expand Down Expand Up @@ -207,15 +210,15 @@ operations and includes destructive prepare, provision, and cleanup cases.

| Tag | Suites | Main contract |
|---|---|---|
| `precheck` | `environment`, `storage`, `dependencies`, `inputs` | OIM identity, selected NFS reachability, upstream artifacts, and required inputs |
| `prepare` | `openchami`, `network`, `openldap` | OpenCHAMI, PostgreSQL, networking, DNS/DHCP, and LDAP readiness |
| `precheck` | `environment`, `oim_readiness`, `storage`, `dependencies`, `inputs` | OIM identity and readiness, selected NFS reachability, upstream artifacts, and required inputs |
| `prepare` | `openchami`, `network`, `openldap` | OpenCHAMI, PostgreSQL, networking, DNS/DHCP, and local OpenLDAP readiness |
| `provision` | `openchami` | Provision reports plus SMD, Boot Service, Metadata Service, and network inventory |
| `pxeboot` | `connectivity`, `cloudinit`, `kubernetes`, `slurm`, `apptainer` | Node boot completion and workload-cluster behavior |
| `pxeboot` | See `SUITES` in `library/vars/domain_vars.py` | Node boot completion, node features, and workload-cluster behavior; reboot and node-removal suites run last |
| `cleanup` | `openchami`, `openldap`, `slurm`, `kubernetes`, `artifacts`, `credentials` | Explicit full-cleanup postconditions |

An untagged FVT flow uses `precheck -> prepare -> provision -> pxeboot`.
Cleanup is always explicit. Untagged verification excludes negative and
disruptive cases.
Cleanup is always explicit. Reboot cases never run without `--marker reboot`,
and node remove/add cases never run without `--marker node_lifecycle`.

### Options

Expand All @@ -237,6 +240,14 @@ current directories.
| Comma | Logical OR | `--marker sanity,functional` |
| Plus | Logical AND | `--marker slurm+non_disruptive` |

Without `--marker`, every collected test runs, including functional,
negative, image-download, and scheduler-drain tests, except reboot and node
remove/add tests. Those are deselected (not listed in the run), even with
`--suite`, until `--marker` selects `reboot` or `node_lifecycle`; run them in a
maintenance window. Lifecycle execution cases
still run only in the `exec` phase. Pass a marker such as `--marker sanity` to
narrow the run.

`sanity+functional` selects only tests carrying both markers; it does not mean
“run sanity, then functional.” Use `sanity,functional` for that union.

Expand All @@ -245,12 +256,16 @@ Registered selectors include:
- Baseline and capability: `sanity`, `functional`, `connectivity`,
`cloudinit`, `kubernetes`, `slurm`, `openldap`, and `apptainer`.
- Controlled mutation: `image_download`, `negative`, and `non_disruptive`.
- Maintenance-window operations: `disruptive`, `reboot`, and
- Maintenance-window operations: `reboot` (required to run reboot tests),
`node_lifecycle` (required to run node remove/add tests), and
`scheduler_state`.
- Cleanup authorization: `destructive`.
- Non-functional contracts: `nft`, `performance`, `idempotency`, and
`security`.

`sanity` marks the baseline positive checks only. Negative tests, which
expect a rejection or failure, are never marked `sanity`.

`deploy` is attached to lifecycle execution cases and is normally managed by
the runner rather than selected manually.

Expand Down Expand Up @@ -281,9 +296,12 @@ active PXE mapping, or `groups` matches a mapped `GROUP_NAME`. The stock
./run_validation.sh fvt_orchestrator prepare verify --suite openldap
```

Set `validate_external_ldap: true` to run external proxy and backend checks.
Set `configure_external_ldap: true` only when those checks may also reconcile
the local `omnia_auth` proxy configuration. An unchanged desired configuration
Prepare only starts the local `omnia_auth` container. The external LDAP proxy
and backend checks run first in the pxeboot `slurm_ldap` suite, immediately
before the LDAP login and job checks. Set `validate_external_ldap: true` to run
them and every LDAP-identity check. Set `configure_external_ldap: true` only
when the proxy check may also reconcile the local `omnia_auth` proxy
configuration. An unchanged desired configuration
is an idempotent no-op; a failed changed configuration is rolled back.

### Provision
Expand All @@ -300,7 +318,8 @@ access token.

### PXE boot

PXE verification defaults to `sanity` when no marker is supplied:
PXE verification runs every test except reboot and node remove/add cases when
no marker is supplied. Use `--marker sanity` for the baseline only:

```bash
./run_validation.sh fvt_orchestrator pxeboot test
Expand All @@ -316,6 +335,7 @@ PXE verification defaults to `sanity` when no marker is supplied:
./run_validation.sh fvt_orchestrator pxeboot verify --suite slurm_hpc_benchmarks
./run_validation.sh fvt_orchestrator pxeboot verify --suite coredns_coredhcp
./run_validation.sh fvt_orchestrator pxeboot verify --suite powervault
./run_validation.sh fvt_orchestrator pxeboot verify --suite vast_storage
```

Focused workload and image examples:
Expand All @@ -337,13 +357,16 @@ Focused workload and image examples:
--suite slurm_apptainer --marker functional+image_download
```

Run disruptive checks only in an approved maintenance window:
Run reboot and node remove/add checks only in an approved maintenance window:

```bash
./run_validation.sh fvt_orchestrator pxeboot verify --marker reboot
./run_validation.sh fvt_orchestrator pxeboot verify \
--suite slurm_recovery --marker reboot
./run_validation.sh fvt_orchestrator pxeboot verify \
--marker disruptive+reboot
--suite slurm_lifecycle --marker node_lifecycle
./run_validation.sh fvt_orchestrator pxeboot verify \
--suite slurm_jobs --marker disruptive+scheduler_state
--suite slurm_jobs --marker scheduler_state
```

The authorization checks are enforced even when pytest is invoked directly.
Expand Down Expand Up @@ -384,7 +407,7 @@ artifact permissions:
A complete NFT run provisions state and finishes with full cleanup. Duration
limits come from `nft_performance_threshold_seconds` in `test_config.yml`.
Run NFT separately from FVT cleanup and review the cleanup policy first. See
[nft/README.md](nft/README.md) for its 11 contracts and execution order.
[nft/README.md](nft/README.md) for its 14 contracts and execution order.

## Batch execution

Expand Down
16 changes: 1 addition & 15 deletions test/orchestrator/_run.py
Original file line number Diff line number Diff line change
Expand Up @@ -46,20 +46,6 @@ def _runner_all_exec_tags(args, lifecycle_tags):
return lifecycle_tags


def _apply_safe_pxeboot_marker_default(args):
"""Default PXE verification to positive, non-disruptive sanity checks."""
normalized = list(args)
if (
len(normalized) >= 3
and normalized[0] == "fvt_orchestrator"
and normalized[1] == "pxeboot"
and normalized[2] in {"test", "verify"}
and "--marker" not in normalized
):
normalized.extend(["--marker", "sanity"])
return normalized


def main():
"""Load domain config and run ValidationRunner."""
script_dir = os.path.dirname(os.path.abspath(__file__))
Expand All @@ -81,7 +67,7 @@ def main():
)
from omnia_auto.functions.validation_runner import ValidationRunner

args = _apply_safe_pxeboot_marker_default(sys.argv[1:])
args = sys.argv[1:]
runner = ValidationRunner(
domain=DOMAIN_NAME,
script_dir=script_dir,
Expand Down
115 changes: 76 additions & 39 deletions test/orchestrator/conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,7 @@
)
from library.vars import TEST_CASES
from omnia_auto import (
TestLogger,
TestReport,
add_session_result,
build_report_name,
Expand Down Expand Up @@ -194,6 +195,19 @@
)


_TC_TITLE_MAP = {test_case["id"]: test_case["title"] for test_case in TEST_CASES.values()}


def _skip_reason(result) -> str:
"""Return the skip reason recorded by pytest, without its prefix."""
longrepr = result.longrepr
if isinstance(longrepr, tuple) and len(longrepr) == 3:
text = str(longrepr[2])
else:
text = str(longrepr or "")
return text.removeprefix("Skipped: ").strip()


def _registered_test_case_id(item) -> str:
"""Return a deterministic TC ID without relying on logger state."""
return _TC_ID_MAP.get(item.name, "")
Expand Down Expand Up @@ -241,9 +255,15 @@ def pytest_configure(config):
"kubernetes": "Kubernetes post-boot checks",
"slurm": "Slurm post-boot checks",
"apptainer": "Apptainer runtime, image, and Slurm integration checks",
"benchmark": "HPC benchmark staging and execution tests",
"additional_cloud_init": "Additional cloud-init post-boot verification checks",
"mount_config": "NFS mount_config post-boot verification checks",
"minimal_os": "Minimal OS validation for OS-only provisioned nodes",
"boot_image": "Provisioned boot image identity and architecture checks",
"powervault_infrastructure": "PowerVault iSCSI and multipath infrastructure checks",
"powervault_mounts": "PowerVault partition, filesystem, and mount checks",
"powervault_binds": "PowerVault bind-mount and targeting checks",
"powervault_cloudinit": "PowerVault cloud-init and setup-log checks",
"vast_installation": "VAST NFS client installation tests",
"vast_mounts": "VAST mount point and mount option validation",
"vast_targeting": "VAST functional group targeting validation",
Expand All @@ -252,8 +272,10 @@ def pytest_configure(config):
"image_download": "Explicitly authorized Apptainer image download checks",
"negative": "Expected-failure and rejection behavior checks",
"non_disruptive": "Checks that do not reboot or drain cluster nodes",
"disruptive": "Explicitly enabled reboot or scheduler-state checks",
"reboot": "Node reboot and post-reboot recovery checks",
"reboot": "Node reboot checks; run only when --marker selects reboot",
"node_lifecycle": (
"Node remove/add checks; run only when --marker selects node_lifecycle"
),
"scheduler_state": "Scheduler drain, queue, and resume checks",
"destructive": "Explicitly selected destructive cleanup checks",
"nft": "Non-functional quality-contract checks",
Expand Down Expand Up @@ -288,6 +310,14 @@ def _item_has_marker(item, marker_name):
return item.get_closest_marker(marker_name) is not None


OPT_IN_MARKERS = ("reboot", "node_lifecycle")


def _opt_in_markers(item):
"""Return the opt-in markers a test carries."""
return {name for name in OPT_IN_MARKERS if _item_has_marker(item, name)}


def pytest_collection_modifyitems(session, config, items):
"""Filter markers, apply safe defaults, and sort by order marker."""
marker_expr = config.getoption("--marker", default="")
Expand All @@ -308,35 +338,26 @@ def pytest_collection_modifyitems(session, config, items):
"negative",
}
)
disruptive_authorized = bool(
explicitly_enabled & {"disruptive", "reboot", "scheduler_state"}
)

# Only apply the normal deploy auto-skip if no marker expression is given.
# Without a marker expression every collected test runs except opt-in
# (reboot, node_lifecycle) cases, which are deselected. Deploy cases still
# run only in the exec phase, which creates the verified state.
if mode == "none":
selected = []
deselected = []
for item in items:
if _opt_in_markers(item):
deselected.append(item)
continue
if _item_has_marker(item, "deploy") and command_type != "exec":
item.add_marker(
pytest.mark.skip(
"Deploy tests run only during the runner exec phase"
)
)
elif _item_has_marker(item, "disruptive"):
item.add_marker(
pytest.mark.skip(
"Select a disruptive, reboot, or scheduler_state marker "
"to authorize this test"
)
)
elif _item_has_marker(item, "functional") and not _item_has_marker(
item, "sanity"
):
item.add_marker(
pytest.mark.skip(
"Select a functional or workload capability marker "
"to authorize this test"
)
)
selected.append(item)
if deselected:
config.hook.pytest_deselected(items=deselected)
items[:] = selected
else:
# Feature filters apply to verification cases. The execution phase
# still needs its one deploy test to create the state being verified.
Expand All @@ -352,8 +373,9 @@ def pytest_collection_modifyitems(session, config, items):
else:
match = _item_has_marker(item, markers[0])

if match and _item_has_marker(item, "disruptive"):
match = disruptive_authorized
opt_in = _opt_in_markers(item)
if match and opt_in:
match = bool(opt_in & explicitly_enabled)
elif match and _item_has_marker(item, "functional"):
match = functional_authorized
if (
Expand Down Expand Up @@ -386,10 +408,28 @@ def _get_order(item):


def pytest_runtest_setup(item):
"""Expose only explicitly selected mutation markers to runtime helpers."""
"""Expose mutation authorization to runtime helpers.

With no marker every gate a test carries is authorized; opt-in cases are
deselected at collection. With a marker expression only the selected
mutation markers are.
"""
marker_expr = item.config.getoption("--marker", default="")
_mode, markers = _parse_marker_expression(marker_expr)
selected = set(markers)
if not selected:
# No marker selected: authorize every gate the test carries. Opt-in
# cases are deselected at collection.
authorized = {
marker
for marker in ("functional", "image_download")
if _item_has_marker(item, marker)
}
if authorized:
os.environ["OMNIA_FVT_AUTHORIZED_MARKERS"] = ",".join(sorted(authorized))
else:
os.environ.pop("OMNIA_FVT_AUTHORIZED_MARKERS", None)
return
sanity_authorized = _item_has_marker(item, "sanity") and (
not selected or "sanity" in selected or "buildstream" in selected
)
Expand All @@ -407,12 +447,7 @@ def pytest_runtest_setup(item):
}
):
authorized.add("functional")
if _item_has_marker(item, "disruptive") and selected & {
"disruptive",
"reboot",
"scheduler_state",
}:
authorized.add("disruptive")
authorized |= _opt_in_markers(item) & selected
if _item_has_marker(item, "image_download") and (
"image_download" in selected or sanity_authorized
):
Expand Down Expand Up @@ -578,13 +613,7 @@ def pytest_runtest_makereport(item, call):
skip_reason = ""

if result.skipped:
if hasattr(result, "wasxfail"):
status = "SKIPPED"
rep_text = str(result.longrepr) if result.longrepr else ""
if "Skipped:" in rep_text:
skip_reason = rep_text.split("Skipped:", 1)[-1].strip()
elif "SKIP" in rep_text:
skip_reason = rep_text.split("SKIP", 1)[-1].strip()
skip_reason = _skip_reason(result)

if status == "SKIPPED" and skip_reason:
details = (details + "\n" if details else "") + f"SKIPPED: {skip_reason}"
Expand All @@ -600,11 +629,19 @@ def pytest_runtest_makereport(item, call):
if not tc_id:
tc_id = get_last_tc_id()

if result.when == "setup":
# Marker and fixture skips happen before the test body creates its
# TestLogger, so record the start and skip lines here.
skip_log = TestLogger(_TC_TITLE_MAP.get(tc_id, item.name), tc_id)
skip_log.skipped(skip_reason or "Skipped before the test started")
details = skip_log.get_output() + (f"\nSKIPPED: {skip_reason}" if skip_reason else "")

add_session_result(
test_name=item.name,
status=status,
duration=getattr(result, "duration", 0),
tc_id=tc_id,
reason=skip_reason,
)

report = get_current_report()
Expand Down
Loading
Loading