The Status Field Is Not the Work
The Status Field Is Not the Work
There is a specific way automated operations go wrong, and it is quieter than a crash. A job reports success. The dashboard is green. Nothing was produced.
The inverse happens just as often. A job reports failure, someone spends an hour on the incident, and the output was on disk the whole time. The error came from a cleanup step that ran after the real work was already done.
Both cases have the same root. A status field is a claim about a process. The artifact is the thing you actually wanted. They are produced by different code paths, so they can disagree, and when they disagree the status is the one more likely to be wrong.
Why the field drifts from the truth
An exit code describes the last thing a process did, not the sum of what it accomplished. A wrapper script that ends in a log rotation returns the status of the log rotation. A pipeline that writes its output in step three and validates in step seven returns the verdict of step seven.
Retries widen the gap. A task that fails twice and succeeds on the third attempt is a success by outcome and a pile of errors by record. Counters that only increment on failure will insist something is broken long after it started working.
Then there is the case where the status is honest and still useless. A job whose only job is to call another service will happily report success for having made the call. Whether the other side did anything is a separate question that nobody asked.
The check that costs one command
The discipline is to name the artifact before you trust the field. Not "did the backup job succeed" but "is there a file, is it the expected size, and does it restore." Not "did the notification send" but "is the message in the channel." Not "did the row get written" but query for the row.
This is one extra command, and it replaces a claim with an observation. In our own operations the rule is written down in plain terms: before describing a system as broken or fine, go look at the thing rather than the field that describes the thing. It exists because reading a status field and reporting from it produced a wrong diagnosis, and the wrong diagnosis spawned three more wrong theories before anyone checked the disk.
Build the verification into the job
If verification is a habit, it gets skipped under pressure. If it is a step, it does not.
The pattern that holds: the job does not get to mark itself complete. A separate step reads the artifact, confirms it exists and is well formed, and only then records the work as done. Publishing a document is not done when the deploy exits zero. It is done when the URL returns the document.
This also gives you an honest failure mode. When the artifact check fails, you know the work did not happen, regardless of what the process claimed. When it passes, you know the work happened, even if the process logged noise on the way out. You stop debugging the reporting layer and start debugging the thing that matters.
The cost is small and the failures it catches are the expensive kind: the silent ones, where everything looks fine for a week and the backup you needed was never written.
Checking the artifact instead of the status field is one of the verification habits in the book. If you are wiring up your own automated operations, One Agent, One Company ($9.97) collects the checks that turned out to matter, along with the incidents that taught us which ones those were.