How to monitor a cron job
5 min read ยท published 2026-10-09
The short version
A cron job that stops running does not page anyone. The schedule keeps firing, the failed run goes nowhere, and the first one to notice is whoever depends on the output: the backup that is three weeks stale, the import that never landed, the digest that stopped going out. The fix is a heartbeat. The job pings a URL when it finishes, the monitor expects the ping on its schedule, and silence opens the incident. Heartbeats are free on every UpControl plan and never count toward the check limit.
What the monitor does, from the product sheet as of 2026-10:
| Heartbeats | |
|---|---|
| Cost | free, on every plan |
| Counts toward the check limit | never |
| Cadences | every 5 or 30 minutes, hourly, every 6 hours, daily |
| When the alert fires | a missed ping opens an incident at most an hour past its due time |
| Where alerts go | Telegram, Slack, Discord, email, unlimited |
| What the incident page carries | the log lines around the failure and the deploy that landed before it |
Why cron fails silently
Cron does one thing: it starts your command at the appointed time. What happens after that is not its problem. If the script dies halfway through, the exit code goes nowhere. Cron can mail the output, but on most servers local mail is never routed, so the message sits in a root mailbox nobody has opened since the box was provisioned.
The failures worth catching:
- The script errors halfway and exits nonzero, every night, forever.
- The script "succeeds" with exit 0 while doing half its work, a truncated import, an empty backup.
- The box rebooted and the scheduler never came back.
- The job is still running from yesterday, holding a lock, and today's run exits immediately.
None of these announce themselves. All of them stop the ping.
Two ways to watch a job
Log scraping: something reads the job's output and looks for errors. It works, but it needs a reader, a place to keep the logs, and rules for what counts as bad. For a nightly backup on a $5 box, that is a second project.
Heartbeat: the job pings a URL on success. The monitor knows when the ping is due and treats silence as the failure. There is nothing to parse, it works from any language that can make one HTTP request, and it catches all four cases above, including the exit 0 half-run: a job that did half its work either does not reach the ping or does not send it.
Setting up the heartbeat
- Run
npx upcontrolin the repo. It hands the wiring to the coding agent you already use (Claude Code, Cursor, Codex, Gemini CLI, Copilot, Windsurf, Amp, Aider, Cline, opencode) and you ask in plain words: "add a heartbeat to the nightly backup job". No tracking code written by hand. One curl line is all it takes if you would rather write it yourself. - Match the cadence to the schedule: a daily job gets a daily heartbeat, an every-30-minutes job gets the 30-minute one. Available cadences are 5 or 30 minutes, hourly, every 6 hours, daily.
- Pick where the alert goes: Telegram, Slack, Discord or email. Unlimited recipients on every plan.
- Test the failure, not the success. Comment the ping out, let the due time pass. The incident opens at most an hour past the due time, which is the honest latency of this method: a heartbeat tells you a job did not finish, an hour after it should have. If you need to know within seconds, the job needs an error path of its own, and the log alerts below are the better fit.
The why, not just down
A heartbeat tells you the job died. The interesting question is what it was doing when it did. UpControl correlates checks with application logs and product events, so the incident page for the missed ping carries the log lines around the failure and the deploy that went out ahead of it, not only "did not run since 03:00". Alerts fire on new errors in your logs on every plan including free, and paid plans add a 15-minute follow-up that says whether the job recovered or is still down.
The same npx upcontrol wires the logging. The agent adds the SDK to the worker, and what the job logs lands in the same timeline as the heartbeat, the site checks and the deploys.
The honest limits
- No phone or SMS escalation, no on-call rotations. Telegram, Slack, Discord and email is the whole list of alert channels.
- The latency above: a missed heartbeat is caught at most an hour past its due time, not instantly.
- UpControl is a younger product with fewer integrations than the decade-old incumbents. Heartbeats, logs and alert channels cover the scheduled-job case; a large existing toolchain may miss a connector it relied on elsewhere.
- Website checks are a separate budget: 3 on the free plan, every 5 minutes, counted workspace-wide. Heartbeats never touch that count.
- The whole core is AGPL-3.0 and self-hostable from one compose file on a 1GB RAM box, with zero telemetry in the OSS build, if the jobs live somewhere you would rather not send pings from.
One npx command wires the heartbeat and the logs: [start with npx upcontrol](https://upcontrol.io/?utm_source=blog&utm_medium=vs&utm_campaign=engine).