Jobs that overlap themselves
When a run outlasts its interval, the next start can proceed, and success heartbeats stay green while both runs write.
When a run outlasts its interval, the next start can proceed, and success heartbeats stay green while both runs write.
A successful retry can erase the first failure from dashboards and alerts. Keep attempt history visible so late greens do not hide standing defects.
A missed heartbeat is only useful if it reaches the owner who can act, not whatever on-call rotation happens to hold the phone.
DST skips, repeated hours, and UTC alert windows create pages that get closed as noise. The schedule definition is usually the real bug.
Quiet batch jobs with perfect success rates often hide the most expensive failures. Treat absence and outcome as the signals that matter.
A guest post from the Modern Serverless team on the workload most likely to be forgotten during an office migration and the absence-watch that catches it.
A practical comparison of heartbeat-based and check-based monitoring for scheduled work, with guidance on which combination to install for which class of job.
An operator's guide to recognizing the silent failures of scheduled jobs, and the small set of practices that prevent the next one from being a customer-facing surprise.