A Status Field Is Not an Outcome
A scheduled job finished and wrote status: ok. The file it was supposed to produce was not on disk.
A different job wrote status: error. The output had shipped an hour earlier, correctly, and the error came from a cleanup step that ran after the real work was done.
Both of these are the same bug, and it is not in the jobs. It is in the reader. We treated a status field as if it described reality. It describes what one piece of code believed about itself at the moment it exited.
Status is a claim. The artifact is the evidence.
An exit code tells you a process terminated a particular way. It does not tell you the row landed in the table, the message arrived in the channel, the file exists at the path, or the content inside it is the content you wanted. Every one of those is a separate question, and every one of them can fail while the status stays green.
The gap shows up in predictable places:
- The work succeeds, then a non-essential teardown step fails, and the whole run is marked failed. Someone pages on it. Nothing was wrong.
- The work is skipped by a guard clause nobody remembers writing, the function returns cleanly, and the status is success. Nothing was done.
- The write is issued and acknowledged by a client library that buffers, the process exits before the flush, and the status was recorded before the data was durable.
- The job runs against the wrong target. Perfect execution, correct status, zero effect on the thing you cared about.
In each case the status field is honest about what it measured. It measured the wrong thing.
The habit that fixes it
Before describing anything as working or broken, go look at the thing itself. Not the field that summarizes the thing.
Did the job write the file? List the path and check the modification time. Did the announcement go out? Read the channel. Did the migration apply? Query the schema. Did the cache warm? Ask it for a key.
This is slower than reading a dashboard, which is exactly why the dashboard exists and exactly why it misleads. The dashboard is a cache of a claim, and like any cache it can be stale, scoped wrong, or populated by code that never checked.
Make the check part of the job
The durable version of this is not personal discipline. It is a second step that runs after the first one and asserts on the artifact.
A publish step followed by a verify step that fetches the published URL and confirms the content. A write followed by a read-back. An upload followed by a HEAD request that confirms the size. The job is not allowed to call itself done until something has independently observed the result.
The two steps must not share the assumption that broke. A verifier that reads the same in-memory variable the writer set is a second opinion from the same witness. It has to go out to the real surface and come back.
What to do with a status field
Keep them. They are a cheap first filter and a reasonable trigger for looking closer. Use them to decide where to point attention, never as the final word on whether work happened.
The rule that stuck for us: a status field can tell you something is probably fine. It can never tell you something definitely shipped. For that, go look.
Verification steps like these, and the incidents that taught us to wire them in, are collected in One Agent, One Company ($9.97). Useful if you are deciding which of your jobs deserve a read-back step.