How Do You Know Your Bot's Recurring Routine Has Gone Silent?
A quiet recurring routine has three possible failure points: no invocation, failed delivery to the target, or work that starts but stalls. An independent deadline check detects a missed slot; scheduler retries and a dead-letter queue record failed delivery; a configured heartbeat timeout detects stalled long-running work. A hypothetical nightly digest shows how to combine the signals without mistaking one for another.
AI-assisted editorial: researched and drafted with AI using the primary sources linked below. Examples are illustrative; they are not measured product results. Editorial policy.
Key takeaways
- Check each expected schedule slot against a durable run or completion record after a defined deadline; a dead-letter queue cannot report an invocation that never occurred.
- EventBridge Scheduler retries failed target delivery and can send exhausted failures to a dead-letter queue; its CloudWatch metrics help diagnose attempts but are best effort.
- For long-running work, a configured heartbeat timeout can expose lost progress; safe retries still need idempotent external actions or checkpoints.
Three places a recurring routine can go quiet
Imagine a bot expected to produce a support digest every night. If there is no digest in the morning, that absence alone does not identify the fault. The schedule may never have invoked the target. It may have attempted delivery and failed. Or the target may have accepted the request while the digest work later stalled. Those cases have different owners, evidence, and recovery steps.
Start with a durable record keyed by the expected slot, such as the routine name and scheduled date. Record invocation, running, completed, and failed states where possible. A separate watchdog can compare the expected slot with that record after a deadline and alert a person when the slot is absent or incomplete. Run this check outside the routine being watched: a bot cannot reliably report its own failure to start.
Amazon EventBridge Scheduler's invocation-attempt metrics can help investigate whether it tried to invoke a target. AWS cautions that these CloudWatch metrics are delivered on a best-effort basis and are not a complete accounting. A missing metric is therefore a clue, not proof that no invocation happened. The independent expected-slot check remains the backstop.
Sources: Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch.
When the schedule fires but target delivery fails
EventBridge Scheduler can retry a failed target invocation according to the schedule's configured retry policy. If delivery still fails, a configured dead-letter queue (DLQ) can receive a record with the target, error code, error message, retry attempts, and reason retries stopped. The retry limits are configuration choices; do not assume a particular number of attempts or duration for your schedule.
That DLQ answers a narrow question: what happened when Scheduler tried to reach the target? It does not mean the downstream job completed successfully, and it cannot report a schedule that never attempted an invocation. If the target accepted the request and later hung, delivery may look successful while the business result is missing.
For a small system, the same idea can be a durable failed-delivery row with slot ID, target, error, and resolution state. Alert on new unresolved failures, inspect the cause, and replay only after the target action can tolerate duplicate attempts. If the row is empty but the deadline watchdog sees no completed slot, investigate the scheduler and worker rather than declaring the routine healthy.
Sources: Configuring a schedule's dead-letter queue in EventBridge Scheduler, Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch.
When work begins and then stalls
A successful handoff says little about work that lasts minutes or hours. A worker might freeze during a tool call, lose its process, or stop moving through a batch. Temporal documents Activity Heartbeats for long-running work: the Activity reports progress, and a configured Heartbeat Timeout can fail an attempt when heartbeats stop. A retry follows only when the retry policy allows it.
Heartbeat timeouts must actually be configured, and the Activity must send heartbeats. A zero or unset Heartbeat Timeout disables that missing-heartbeat check. A separate overall deadline is also useful when a worker keeps sending heartbeats but never completes meaningful work. A heartbeat is evidence of a live worker, not proof of a correct result.
Temporal lets an Activity include progress details in a heartbeat that a later attempt can read. That can help resume a batch, but the detail alone does not make side effects safe. Store checkpoints durably and make external writes idempotent, for example by giving each digest post a stable slot ID, before allowing retries to repeat them.
Sources: Detecting Activity failures, Timeouts and Retry Policies.
A nightly digest, traced from schedule to result
Consider a hypothetical team bot that reads support tickets, drafts a digest, and posts it to a channel. This is an illustrative design, not a report of a live deployment. The team expects one completed record for each night's slot. Its watchdog checks the previous slot after the agreed deadline and sends an actionable alert if the record is missing or unfinished.
If the Scheduler attempted to call the digest target but that delivery failed, retry and DLQ evidence point to an address, permission, throttling, or target-service problem. The operator checks the DLQ and target logs, repairs the cause, and replays the same slot safely. If there is no invocation record or delivery failure, the operator checks the schedule state and independent logs; an empty DLQ alone proves nothing.
If the digest worker started and then stopped making progress, a heartbeat timeout or run deadline marks the attempt as stalled. A retry can read the last durable checkpoint, but posting the digest must still use the slot ID to avoid a duplicate message. The final completed record should identify the actual posted result, so the watchdog can tell completion from mere launch.
Sources: Configuring a schedule's dead-letter queue in EventBridge Scheduler, Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch, Detecting Activity failures.
The smallest useful reliability loop
Give every scheduled run a stable slot ID and a completion deadline. Persist a state transition when the target starts and when the intended result is actually delivered. Check expected slots from a separate watchdog and route misses to someone who can act; otherwise the failure only moves from one silent database row to another.
Configure retry and failed-delivery capture for the scheduler-to-target boundary. For long-running tasks, add an appropriate progress heartbeat and timeout, plus an overall deadline. Make retries safe through idempotent writes and durable checkpoints. The right signal depends on where the work stopped, so inspect the slot record, delivery evidence, and worker progress together.
Finally, test the three cases deliberately: disable a future schedule in a safe test environment, make a target reject a test delivery, and pause a long-running test worker. Verify that each produces a distinct alert and a recoverable record. A green scheduler dashboard by itself does not establish that the digest arrived.
Sources: Configuring a schedule's dead-letter queue in EventBridge Scheduler, Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch, Detecting Activity failures.
Primary sources
Sources checked 2026-09-30. Standards and product documentation can change; follow the linked version when implementing.
- Configuring a schedule's dead-letter queue in EventBridge Scheduler — Amazon Web Services
- Monitoring Amazon EventBridge Scheduler with Amazon CloudWatch — Amazon Web Services
- Detecting Activity failures — Temporal Technologies
- Timeouts and Retry Policies — Temporal Technologies
BotBento is in development. Suggest a correction.