Choosing what happens after a failed step
A practical workshop for a non-critical notification must not block the load. Build the smallest manifest, validate it, run it and read the result.
Context — a non-critical notification must not block the load
The task is currently manual and its assumptions are not recorded. Hydra turns it into files that can be reviewed and rerun.
Where things stand
- The current result is fragile. Its assumptions are split between tools, clicks and memory.
- Reruns are uncertain. The write mode or orchestration rule is not visible beside the data.
- Evidence is missing. A colleague cannot compare a declared rule with a concrete before and after state.
The question
How do you make on_failure explicit in workflow.yaml and verify the resulting execution states?
- Keep jobs independently runnable
- Give every workflow step a unique name
- Validate the graph before execution
- Read every step status in the summary
The solution in one line
A small workflow.yaml, one validation command and one run with observable step states.
workflow.yaml jobs/ sales/ inventory/ publish/
Name the work the workflow will coordinate.
Steps
1. prepare the jobs and boundary
- Name the work the workflow will coordinate.
- Compare the file with the explanation in the workbench.
- Record the shown check before moving to the next step.
workflow.yaml jobs/ sales/ inventory/ publish/
2. declare on_failure
- Write the trigger, steps and policy.
- Compare the file with the explanation in the workbench.
- Record the shown check before moving to the next step.
workflow:
version: "1.0"
name: on_failure_workshop
description: "Choosing what happens after a failed step"
trigger:
type: manual
steps:
- name: notify
type: action
action: webhook
params:
url: ${ENV:OPS_WEBHOOK_URL}
on_failure: continue
- name: load_sales
type: job
job: ./jobs/sales
depends_on: [notify] 3. validate the graph
- Catch missing jobs, duplicate names and invalid dependencies.
- Compare the file with the explanation in the workbench.
- Record the shown check before moving to the next step.
$ hdrctl workflow validate workflow.yaml ✅ Workflow valid: on_failure_workshop Steps: 1
4. run the workflow
- Execute and follow step states.
- Compare the file with the explanation in the workbench.
- Record the shown check before moving to the next step.
$ hdrctl workflow run workflow.yaml notify failed · continued load_sales succeeded
5. read the orchestration result
- Distinguish success, skip, retry and continuation.
- Compare the file with the explanation in the workbench.
- Record the shown check before moving to the next step.
notify failed · continued load_sales succeeded
Expected result
- The workflow validates before execution.
- notify failed · continued
- The run summary explains every terminal state.
Reading the results
Read the summary as a graph, not a flat log. Use `continue` only when downstream work remains valid without the failed step.
What the counters do—and do not—prove
They prove how many records entered and left this run. They do not replace checking the target schema, the business meaning of values or the reason a workflow step was skipped.
The rule to carry forward
Use `continue` only when downstream work remains valid without the failed step.
Did we answer the question?
| objective | result | where |
|---|---|---|
| Configuration explicit | yes | the relevant YAML block |
| Safe rerun | yes | the destination or workflow policy |
| Observable result | yes | the command output and counters |
| Hidden manual rule | removed | the rule now lives in versioned text |
Before and after
- Before — a non-critical notification must not block the load requires a person to remember the order, options and checks.
- After — one reviewed manifest and one command produce the same observable result.
- What is really gained — The durable gain is the contract: a colleague can read the configuration, reproduce the run and challenge the assumptions.