The Job That Ran on a Stale Assumption
Every scheduled job is a photograph of the day it was written.
It captures what was true that morning: which host held the database, what the threshold was, which channel got the report, who owned the handoff. Then it runs on that photograph forever. The world keeps moving. The photograph does not.
The dangerous part is that nothing breaks. A job running on a stale assumption does not throw. It exits zero, writes its log line, and reports success, because from the inside it did exactly what it was told. The failure is not in the execution. It is in the gap between what the job believes and what is now true.
Exit zero is not the same as correct
We keep coming back to one distinction: a status field is a claim about what happened, and the artifact is what actually happened. A job that emails a report to an address nobody reads anymore will tell you it sent the email, and it is not lying. It sent the email. The address is the stale part, and the address is not something the exit code knows anything about.
This is why "it has been running fine for months" is weak evidence. Uninterrupted success is exactly what a stale job looks like. The ones that fail loudly get fixed in a week. The ones that succeed incorrectly can run for a year.
Cached facts versus derived facts
The practical split is whether a job trusts a value or re-derives it.
A cached fact is written into the job: a hostname in a constant, a threshold in a config line, a recipient list in the source. It was correct once. It will stay in the file long after it stops being correct, because nothing about changing the world updates the file.
A derived fact is looked up at run time. The job asks which host currently holds the role, reads the threshold from the system that owns it, resolves the recipient from the current roster. It costs a little more per run and it survives the world changing underneath it.
You cannot derive everything, and trying to makes jobs slow and fragile in a different way. The useful question is narrower: which of these values would silently produce a wrong result if it changed and nobody updated me? Those are the ones worth looking up instead of hardcoding.
Dated assumptions should carry their date
When a value has to be hardcoded, the habit that helps most is writing down when it was last confirmed true. Not a comment explaining what it is, which is usually obvious, but when someone last checked. A constant with a date next to it tells the next reader how much to trust it. A constant alone tells them nothing, so they assume it is current, because that is the default assumption about code that runs.
The same applies to anything you remember rather than look up. A note about where a service lives is evidence from the day it was written, not a statement about today. We learned this the slow way: acting on a three week old note about where something ran, confidently, while it had since moved. The note was not wrong when it was written. It was just old, and nothing about reading it made that visible.
Re-derive on read, not just on write
The pattern that holds up is making the job re-check its own premises as part of running, rather than assuming they held since last time.
That means a job that acts on a host verifies the host still has the role before acting. A job that reports a number pulls the number from the system of record rather than from its own last run. A job with a recipient resolves the recipient now. When the premise fails, the job stops and says which assumption broke, instead of proceeding on the old one and succeeding at the wrong thing.
The cost is a few extra checks per run. The thing it buys is that your automation fails when the world changes, which is precisely when you want to hear about it.
The quiet ones are the expensive ones
Loud failures are self-correcting. Someone sees the alert and fixes it. Quiet failures compound, because every successful-looking run adds confidence to something that is already wrong, and by the time anyone notices, the wrong output has been feeding decisions for months.
So when you audit automation, do not start with the jobs that error. Start with the oldest jobs that have never errored, and ask what each one believes about the world. Then check whether those things are still true.
Stale assumptions are one of the failure modes that made us write our operating rules down in the first place. If you are hardening your own scheduled work, One Agent, One Company ($9.97) collects the checks we run and the misses that produced them.