cronalerts all systems nominal

Jobs that overlap themselves

When a run outlasts its interval, the next start can proceed, and success heartbeats stay green while both runs write.

A reconciliation job on a ten-minute cron finished in about four minutes for most of a year. The query grew with the ledger. Median runtime moved through six minutes, then eight. The job still exited zero, and the completion heartbeat still landed inside a fifteen-minute grace, so the absence alert stayed quiet. On a month-end night a vacuum ran in the same window and one run took fourteen minutes. At the ten-minute mark the scheduler started a second run while the first still held an advisory lock on the snapshot table, and the second run waited. At the twenty-minute mark a third run started. When the first run committed and released the lock, the second wrote another snapshot for a window the first had already closed. Both runs completed and sent heartbeats, and billing ingested the two snapshots.

The period had become shorter than the work, and nothing in the path refused the extra start. Cron, on the hosts most teams still use, fires on the clock. It does not look up whether the previous invocation is alive. A Kubernetes CronJob does the same unless concurrencyPolicy is set, and the default is Allow. The schedule chooses the start time. Whether a second start may proceed is a separate setting.

Duration against the period

The collision is rarely the first day the job is slow. Runtime creeps: a new join, a missing index, a downstream call that picked up its own retries, a table that doubled after a backfill. Exit codes stay at zero through all of it. Alerts that watch failures and missed heartbeats stay quiet until a run crosses the period and a second process is on the box.

A job scheduled every ten minutes has ten minutes of wall clock, which has to cover process startup, the slow tail, and whatever else that host is doing. When p95 duration passes about two thirds of the period, one slow night is enough to overlap the next start. A flat four-minute job on a sixty-minute schedule can sit for years. A job whose p95 is eight minutes on a ten-minute schedule will overlap on the next run that goes long.

Load on the cron host rises at the same minute and stays high longer each week. Two database sessions show the same application name, and lock waits clear after a few minutes, so they never become a separate incident. Success rate stays flat while duration climbs toward the period.

Allow, skip, or replace

If a second start is possible, write the policy down. The default is that both runs proceed. That is safe when each run is given a distinct set of rows. Two runs that both select the unprocessed batch will each perform the side effect and then mark the batch done. A second charge, or a second copy of the partner file, shows up as another success, and the retry counter stays at zero.

Give concurrent runs an explicit partition: a key range, or a message only one consumer can acknowledge. A job that emits one file per window, or that scans every pending row, should stay on one runner.

Skip means the new start exits while an older run is still open. Kubernetes calls this Forbid. On one host, flock -n around the script does the same thing. The long run continues, and later starts exit immediately until it finishes. One process can drain pending rows to completion. More processes blocked on the same lock wait, then often repeat rows the first run already handled. For a job that must produce a specific window, the skipped start is a missed window. The 10:10 file is absent because the 10:00 run has not finished.

A common wrapper turns a failed flock -n into exit zero so cron stays silent. Scheduler history stays green. If the long run will send a heartbeat when it finishes, the absence alert stays green too, and nothing records the skip. Record it: previous run still open, when it started, how long it has been running, and that this slot did no work. Send that record to the same destination as a missed heartbeat, with the event marked as a skip.

Replace kills the open run and lets the new one take over. It fits output that should match the latest inputs, where a partial result can be discarded, such as a cache fill or a status document overwritten as a whole. It fits poorly once the run has written somewhere external. Stopping a settlement writer at minute ten can leave a partial file in a partner drop, and the new run may write a second file beside it. If the killed process sends no heartbeat, the monitor shows a miss for the old run and a success for the new one. The miss is the run Replace killed.

The lock has to match the policy. An advisory lock or a Redis lock whose TTL is shorter than a slow but valid run expires while the first run is still in the critical section. The next start takes the lock and enters the same section. A lock with no expiry, held by a process that was killed and never released it, blocks every later start until someone removes the key. The page is a missed heartbeat, so check that key before spending the hour in the scheduler. A flock on a local disk only coordinates that host. I have seen both workers run after a move from one cron box to two, each holding /var/lock on its own machine.

For a pending-work drain, let the timer enqueue a token and let one worker with concurrency one do the work. The next start shows up as queue depth, which you can alert on, and the depth stays bounded when the worker keeps up on average.

Open runs and skip counts

Absence alerts fire when an overlap stalls badly enough that heartbeats stop. When both runs finish, they still send heartbeats, and the absence alert stays quiet. A heartbeat at the start makes several waiting runs look like a normal cadence. A heartbeat only at completion can show two successes close together and a gap after them, and that pattern gets closed as scheduler drift. Emit a start with a run id and a completion with the same id. A start that arrives while an earlier id for that job is still open is the overlap. Page on it while both runs are still succeeding.

One success heartbeat for the schedule means some run finished. A concurrent start is visible only when the payload includes the run id and a flag that another run was already open. Otherwise the heartbeat stream looks healthy during the duplicate write.

A grace wider than the interval hides the same condition. A ten-minute job with a thirty-minute grace can skip a slot, or meet the next start, and still deliver a heartbeat before the grace expires. Set the grace from the lateness you are willing to answer for, and chart duration on its own so the climb shows up before heartbeats stop.

Alert when recent p95 crosses a line you picked for that job, with the job still succeeding. Put skip counts on the same view. Repeated skips under Forbid mean the job no longer fits the slot, even while the long run's late heartbeat satisfies the absence check.

When a miss does page, look for another live run of the same job before you start a new one. A restart on top of a live run turns one late job into two. Find the open run, then decide to wait, to stop it, or to leave it alone.

This week, list production schedules whose period is less than twice their recent p95. For each, write allow, skip, or replace, and one sentence on why. Emit a distinct event for a skipped slot and for a start that arrives while another run is open, and send both to the same place as the absence alert.