Kubernetes can run scheduled tasks natively through CronJobs — no external scheduler needed, everything stays inside the cluster. But debugging them is harder than a regular crontab: there's no terminal to watch, pod logs disappear after garbage collection, and if a task silently exits with an error, the cluster marks it as a success. This guide covers how to create a CronJob, which spec fields actually matter, six common problems (including the one where Kubernetes permanently stops scheduling after 100 missed runs), and three monitoring approaches — from kubectl to external heartbeat pings.
What a Kubernetes CronJob actually is
Before touching YAML files, it helps to understand the three-layer hierarchy that Kubernetes uses for scheduled work.
A CronJob is the top-level object. It holds your schedule (the same cron expression syntax used in traditional Linux crontabs) and a template describing what to run. On each trigger, the CronJob creates a Job. That Job, in turn, creates one or more Pods — the actual containers where your code runs.
Think of it like this:
- CronJob — the alarm clock. "Run this every night at 3 AM."
- Job — one execution. "The 3 AM run on September 25th."
- Pod — the container that does the work. Runs your script, then exits.
Each completed Job leaves behind a record (and its pod logs) so you can check what happened. Kubernetes keeps the last 3 successful and last 1 failed Job by default, then garbage-collects older ones. That detail matters for debugging: if you're looking at a problem from last week, the evidence might already be gone.
Create your first CronJob
Here's a minimal CronJob that prints the date every five minutes. Not useful in production, but good for proving the machinery works.
Save this as date-cronjob.yaml:
apiVersion: batch/v1
kind: CronJob
metadata:
name: print-date
namespace: default
spec:
schedule: "*/5 * * * *"
jobTemplate:
spec:
template:
spec:
containers:
- name: date-printer
image: busybox:1.36
command: ["sh", "-c", "echo Current date: $(date)"]
restartPolicy: OnFailure
Apply it with kubectl:
kubectl apply -f date-cronjob.yaml
You should see:
cronjob.batch/print-date created
Now check that it registered:
kubectl get cronjobs
NAME SCHEDULE SUSPEND ACTIVE LAST SCHEDULE AGE
print-date */5 * * * * False 0 <none> 12s
Wait five minutes (or change the schedule to * * * * * for every minute while testing), then check for Jobs:
kubectl get jobs --sort-by=.metadata.creationTimestamp
NAME COMPLETIONS DURATION AGE
print-date-28275845 1/1 3s 2m
One completion, three seconds to run. Grab the logs from the pod that Job created:
kubectl logs job/print-date-28275845
Current date: Thu Sep 25 03:05:01 UTC 2026
That's the full cycle: CronJob triggered, Job created, Pod ran, output captured. When you're done testing, clean up with kubectl delete cronjob print-date.
CronJob spec fields that matter
The minimal example above skips several fields you'll want in production. Here's what each one does and when to set it.
| Field | Default | What it controls |
|---|---|---|
schedule |
(required) | Standard cron expression. Kubernetes 1.25+ also supports time zones via timeZone field. |
timeZone |
UTC | IANA time zone name (e.g., America/New_York). Available since Kubernetes 1.27 as stable. |
concurrencyPolicy |
Allow | Allow: overlapping runs OK. Forbid: skip if previous still running. Replace: kill the old one, start new. |
startingDeadlineSeconds |
none | How many seconds late a Job can start. If the controller misses the window, the run is skipped. Set this. More on why below. |
successfulJobsHistoryLimit |
3 | How many completed Jobs to keep. Lower means fewer old pods cluttering the namespace. |
failedJobsHistoryLimit |
1 | How many failed Jobs to keep. Raise this to 3-5 so you have logs to investigate. |
suspend |
false | Set to true to pause scheduling without deleting the CronJob. Handy during maintenance. |
backoffLimit |
6 | Set on the Job spec (inside jobTemplate). How many times a failed pod retries before the Job is marked failed. |
activeDeadlineSeconds |
none | Also on the Job spec. Maximum runtime before the Job is killed. Prevents stuck jobs from running forever. |
A production-ready version of the earlier example would add concurrencyPolicy: Forbid, a startingDeadlineSeconds value, and bump failedJobsHistoryLimit to at least 3. That gives you overlap protection, missed-schedule handling, and enough history to debug problems.
Problems you'll run into (and why they happen)
Kubernetes CronJobs have a handful of failure modes that trip up almost everyone who uses them for the first time. Some are obvious once you know about them. Others are genuinely annoying to track down.
Silent exit-code failures
Your script might fail halfway through but still exit with code 0. Maybe a database backup command runs, but the upload to S3 never completes because the credentials expired. Kubernetes sees exit 0 and marks the Job as successful. No alert, no retry, no indication that anything went wrong. This is the same silent failure pattern that plagues traditional cron jobs, and it's even harder to catch in Kubernetes because you might not be watching the logs.
Fix: always check the actual result inside your script, and exit with a non-zero code on failure.
#!/bin/bash
pg_dump mydb | gzip | aws s3 cp - s3://backups/mydb-$(date +%F).sql.gz
if [ $? -ne 0 ]; then
echo "Backup failed" >&2
exit 1
fi
That's the bare minimum. Without it, Kubernetes has no way to know the job didn't actually do its work.
The 100-missed-schedules deadlock
Here's one that surprises people. If a CronJob misses 100 consecutive scheduled runs (because the controller was down, the cluster was overloaded, or the CronJob was suspended for too long), Kubernetes stops scheduling it entirely. You'll see this in the controller logs:
Cannot determine if job needs to be started: too many missed start times (> 100)
No error on the CronJob object itself. It just silently stops creating Jobs. Ran into this on a staging cluster where someone suspended a CronJob for a month and couldn't figure out why it wouldn't resume.
Fix: set startingDeadlineSeconds. When this field is present, Kubernetes only counts missed schedules within that window instead of counting all the way back to the last successful run. A value of 200 works well for most minute-level schedules. For hourly jobs, 3600 (one hour) is reasonable.
Stuck jobs eating resources
Without activeDeadlineSeconds, a Job that hangs (waiting on a network call, stuck in an infinite loop, deadlocked) will run forever. If concurrencyPolicy is set to Forbid, that stuck job blocks all future runs. If it's set to Allow, you'll get a pile-up of running pods.
Set activeDeadlineSeconds to something generous but finite. For a backup that normally takes 10 minutes, 1800 seconds (30 minutes) gives plenty of headroom without letting it run for days.
Image pull errors
CronJob pods are created fresh each run. If your container registry requires authentication and the pull secret is missing or expired, every single run will fail with ImagePullBackOff. Because the pod never starts, there are no application logs to check. You'll want to look at Kubernetes events instead:
kubectl describe job print-date-28275845
Check the Events section at the bottom of the output for pull errors and authentication failures.
OOMKill: out-of-memory termination
If your container's memory usage exceeds its resource limit, Kubernetes kills the pod. The exit code is 137, and the reason shows as OOMKilled. For CronJobs that process variable-size datasets (daily report generation, log aggregation), this can happen intermittently. Works fine on Monday when there's 50MB of data, fails on Friday when there's 500MB.
Set memory requests and limits based on peak usage, not average. And always check the pod's termination reason when debugging failures:
kubectl get pod print-date-28275845-abc12 -o jsonpath='{.status.containerStatuses[0].state.terminated.reason}'
Time zone confusion
Before Kubernetes 1.27, CronJob schedules always ran in the controller's time zone, usually UTC. If your cluster runs in UTC but you scheduled a job for 0 9 * * * expecting 9 AM Eastern, it would actually fire at 9 AM UTC (5 AM Eastern). Not a bug, just a mismatch that wastes time tracking down.
On Kubernetes 1.27+, use the timeZone field explicitly:
spec:
schedule: "0 9 * * *"
timeZone: "America/New_York"
With older clusters, convert your schedule to UTC and add a comment explaining the conversion so the next person doesn't "fix" it.
Three ways to monitor Kubernetes CronJobs
Knowing the failure modes is one thing. Catching them before they cause real damage is another. There are three main approaches, each with different trade-offs.
kubectl and manual checks
The simplest approach: periodically run kubectl get cronjobs and kubectl get jobs to see what's running and what failed. Check logs when something looks wrong.
# Show all CronJobs with their last schedule time
kubectl get cronjobs -A
# Find failed jobs across all namespaces
kubectl get jobs -A --field-selector status.successful=0
Honest assessment: this works if you have 2-3 CronJobs and you're already in the terminal regularly. It does not scale. You won't catch a failure at 3 AM. You won't notice a CronJob that silently stopped scheduling. And once those old Job objects get garbage-collected, the evidence is gone.
Prometheus and kube-state-metrics
For teams already running Prometheus, kube-state-metrics exposes CronJob and Job metrics that you can build alerts on. The relevant ones:
# Alert when a CronJob hasn't run recently
- alert: CronJobNotScheduled
expr: |
time() - kube_cronjob_next_schedule_time > 3600
for: 10m
labels:
severity: warning
annotations:
summary: "CronJob {{ $labels.cronjob }} missed its schedule"
You can also watch kube_job_status_failed for failures, kube_job_status_active for stuck jobs, and kube_cronjob_status_last_schedule_time for scheduling health.
This approach gives you good visibility and proper alerting. The downside: it takes real work to set up. You need Prometheus running, kube-state-metrics deployed, alerting rules written and tuned, and a notification pipeline (Alertmanager to Slack or PagerDuty). For a team that already has this infrastructure, adding CronJob alerts is straightforward. For a team that doesn't, standing up Prometheus just for CronJob monitoring is overkill.
Heartbeat monitoring (the simplest reliable option)
A dead man's switch flips the model: instead of watching for failures, you watch for the absence of success. Your CronJob pings a URL after each successful run. If the ping stops arriving, you get alerted.
The idea is dead simple. Add a curl call at the end of your script, after the real work completes:
#!/bin/bash
set -e
# Do the actual work
pg_dump mydb | gzip > /tmp/backup.sql.gz
aws s3 cp /tmp/backup.sql.gz s3://backups/mydb-$(date +%F).sql.gz
# Signal success
curl -fsS -m 10 --retry 3 https://monitoring.example.com/ping/YOUR_TOKEN
Because set -e stops the script on any error, the curl only fires if everything above it succeeded. No false "all clear" signals. The monitoring service knows your schedule and expects a ping within a grace window. Miss it, and you hear about it through whatever notification channel you've configured. Slack, email, SMS, whatever your team uses.
This approach works across anything that runs scheduled tasks. Kubernetes CronJobs, traditional crontabs, Windows Task Scheduler, CI/CD pipelines, or even managed cloud schedulers. You can read more about setting up heartbeat monitoring with curl in our dedicated guide.
Production best practices (short list)
After working with CronJobs across a handful of clusters, a few patterns hold up consistently:
- Always set
startingDeadlineSeconds. Prevents the 100-miss deadlock and makes recovery from controller downtime predictable. - Always set
activeDeadlineSeconds. No job should be able to run forever. Pick a value that's generous but finite. - Use
concurrencyPolicy: Forbidunless you have a specific reason for overlapping runs. Most scheduled tasks aren't designed to run concurrently with themselves. - Exit non-zero on failure. This sounds basic, but a surprising number of CronJob scripts use
|| trueor swallow errors. If the job didn't do its work, Kubernetes needs to know. - Bump
failedJobsHistoryLimitto at least 3. The default of 1 means you only have logs from the most recent failure. If two failures happen back to back, you lose the first one. - Set resource requests and limits. Without memory limits, a job that leaks memory can take down a node. Without CPU requests, it might get starved and time out.
- Add a heartbeat ping at the end. The cheapest insurance you can buy. A curl call after a successful run, pointed at any monitoring service, catches every category of failure: crashes, hangs, scheduling issues, OOMKills, and the silent exit-0-but-nothing-worked scenario.
- Use the
timeZonefield on 1.27+. Don't rely on "the cluster is in UTC so we just offset." Someone will change the cluster config or migrate to a different provider, and all your schedules will shift.
How WatchCron monitors your Kubernetes CronJobs
WatchCron uses the heartbeat approach described above. You create a cron job monitor, tell it your schedule and grace period, and get a unique ping URL:
curl -fsS -m 10 --retry 3 https://watchcron.com/ping/your-unique-id
Add that line to the end of your CronJob's script. WatchCron calculates expected ping times from your cron expression, and if a ping doesn't arrive within the grace window, it fires alerts through whatever channels your team uses. Email, Slack, Telegram, Discord, Microsoft Teams on Starter. SMS on Pro. Phone calls and PagerDuty on Business. We keep a log of every ping (and every missed ping), so when you're debugging a failure you have a clear timeline of when the job last ran successfully, when it started missing, and what the gap looked like.
The free plan covers up to 20 monitors, which is enough for most small clusters. For production workloads spread across multiple namespaces or clusters, the paid plans scale to 1,000 monitors with team collaboration built in.
Where to go from here
Kubernetes CronJobs are a natural fit for any workload that already runs on your cluster. No external scheduler needed, no separate infrastructure to maintain. But the failure modes are real, and Kubernetes won't go out of its way to tell you about them. Getting your CronJob spec right (deadline seconds, concurrency, history limits) handles the preventable problems. Heartbeat monitoring handles everything else.
If you're still building out your cron job knowledge, the pillar guide covers traditional crontab syntax and scheduling patterns in depth. For the specific problem of jobs that exit successfully but didn't actually work, the debugging silent failures guide goes deeper into detection strategies. And if you'd rather not manage the scheduler yourself at all, Cloud Cron runs your scheduled HTTP calls from our infrastructure. No cluster, no YAML, no pod lifecycle to worry about.