title: Logs Nobody Reads Are Not Observability
date: 2026-10-02
slug: 2026-10-02-logs-nobody-reads-are-not-observability
summary: Writing a log line proves the code ran. It does not prove anyone will notice when it stops. Observability is the part that reaches a human in time to matter.
tags: observability, agentic-ops, monitoring, failure-modes

# Logs Nobody Reads Are Not Observability

Every automated job we run writes logs. For a long time that felt like observability, and it is easy to see why. The logs were detailed, structured, and timestamped. When something broke we could always reconstruct exactly what happened.

Reconstruct. After.

That is the gap. A log line is a record for whoever goes looking. Observability is the property that someone finds out while it still matters. Those are different systems, and owning the first one does not give you the second.

## How we learned the difference

A scheduled job of ours stopped producing its output. It did not crash. It did not throw. It logged a clean, honest line every single run explaining that it had nothing to do, and then exited zero. A green job, a healthy dashboard, a log file quietly repeating the same sentence for days.

Nobody read it. Why would they? Reading logs is something you do when you already suspect a problem. The whole point of monitoring is to tell you when to start suspecting.

We found it because a human noticed a downstream artifact was stale, which is the worst possible detection channel: it depends on somebody remembering what should exist and checking by hand.

## The test that actually matters

The useful question is not "is this logged?" It is "what reaches a person, and how long does that take?"

For each job worth trusting, we now answer three things:

What does success produce? Not a status code, an artifact. A file at a path, a row in a table, a message in a channel. Something you can point at and say it exists.

What checks for that artifact? Something other than the job itself. A job reporting its own health is a single point of failure that grades its own homework. The check has to live outside the thing it is checking.

Who hears about it, and when? If the answer is "it is in the log," there is no answer. A notification with an owner and a deadline is monitoring. A log line is an archive.

## Silence is the failure mode to design for

Loud failures are the easy ones. An exception, a non-zero exit, a timeout, these all announce themselves and tend to get handled early because they are annoying.

The expensive failures are quiet. Work that stops happening, a queue that stops filling, a report that stops being generated. Nothing errors, because nothing ran. No alert fires, because alerts are usually wired to errors rather than to absence.

So we invert it. Instead of alerting when a job fails, we alert when an expected artifact does not appear inside its expected window. That catches crashes, hangs, misconfiguration, a disabled schedule, and the case where the job ran perfectly and silently did nothing at all. One check, every one of those modes.

## What we kept

Logs did not stop being valuable. They are still the first thing we read once something is known broken, and detailed logs have turned multi-hour investigations into single queries. Keep writing them.

Just stop counting them as coverage. Logs explain a failure you already know about. Observability is what tells you there is one. Confusing the two means your detection time is however long it takes a human to randomly wonder.

Absence of an alert is not evidence of health. It is usually just evidence that nothing was watching for absence.

Expected-artifact checks are the pattern we lean on hardest, and [One Agent, One Company ($9.97)](https://opsbyagent.com/book) walks through how to wire them for jobs that can fail by doing nothing. Worth a read if your dashboards are green and you are not quite sure why you believe them.
