Queue alerts: DLQ, stalled queues and backlog
Kubepier’s queue dashboard tells you when a queue needs attention: messages in the DLQ, a growing DLQ, a burst of failures, a stalled queue (no consumption) or backlog above the threshold. They are the same alerts on RabbitMQ and Azure Service Bus, on the web and on desktop. This page shows when each alert fires and what to do.
Queue alerts and when they fire
| Alert | When it shows | Level |
|---|---|---|
| DLQ | The DLQ has messages: 1 to 99 | Attention (100 or more: critical) |
| DLQ growing | The DLQ went up in the last 5 minutes, counted from the last drop, with at least 60 s of history | Critical |
| DLQ spike | The DLQ gained 200 or more messages between two samples less than 60 s apart, in the last 5 minutes | Critical |
| No consumption | The backlog only went up in the last 10 minutes and the out rate stayed under 0.1 message per minute | Critical |
| High backlog | Ready messages above the queue threshold (1,000 by default) | Attention (10 times the threshold: critical) |
Each queue shows its own alerts on its row, with the level (attention or critical). At the top of the dashboard, the ready-message and DLQ totals also change level according to the thresholds.
Where the alerts show
- On the Panel tab (Dashboard on desktop) of each client’s RabbitMQ and Azure Service Bus menus, on the web and on desktop. The dashboard refreshes every 15 seconds (you can switch to 30 s or 60 s, or pause).
- Alerts are computed on the open dashboard, from the last 60 minutes of history kept only in memory (in the browser on the web; in the app on desktop). Reloading or closing starts the history over, and time-based alerts (DLQ growing, no consumption) need their full window again.
- Kubepier sends no email or notification for queue alerts: they show on the dashboard while it is open.
- On Free, the dashboard shows totals and counts, without rates, charts or alerts.
Messages in the DLQ (1 to 99 attention, 100 or more critical)
The DLQ holds messages the application could not process. A few may be one-off cases; 100 or more is a problem that needs an owner.
- Peek the DLQ (Pro and Team) and read the reason: on Service Bus, DeadLetterReason and DeadLetterErrorDescription (MaxDeliveryCountExceeded, TTLExpiredException or whatever the application set); on RabbitMQ, the x-death header (rejected, expired, maxlen or delivery_limit).
- Look for the error in the log of the pod that consumes the queue, at the time of the messages.
- Fix the cause before touching the DLQ. Then reprocess with the application’s tooling or purge the DLQ, if the messages can be discarded. Kubepier peeks and purges; it does not move or resend messages.
DLQ growing (critical)
The DLQ grew over the last 5 minutes, counting from its last drop, with at least 60 seconds of history. The failure is happening now.
- Check for a recent consumer deploy: if the failure started with it, roll back (kubectl rollout undo) while you investigate.
- Check the consumer pod: CrashLoopBackOff, OOMKilled or an error in the log for every message.
- Check the dependency the consumer calls (database, API, another service): if it is down, every message fails and lands in the DLQ.
- Do not purge the DLQ while it grows: you lose the messages and the evidence.
DLQ spike (critical)
The DLQ gained 200 or more messages between two samples less than 60 seconds apart. The alert stays for 5 minutes after the spike.
- A whole batch failed at once: peek the newest DLQ messages and see whether they come from the same producer or are of the same type.
- A malformed message from a producer, a contract change between services or a dependency that went down for a few seconds are the most common causes.
- If the spike stopped and the DLQ no longer grows, treat it like the messages-in-the-DLQ alert.
No consumption: the stalled queue (critical)
The backlog only went up over the last 10 minutes, without a single drop between samples, and output stayed below 0.1 message per minute. On RabbitMQ, output comes from the broker counters; on Service Bus, which does not report output, a backlog that never drops is the signal.
- Check for consumers: on RabbitMQ, the queue’s consumers column; zero means no process is connected.
- Check the consumer Deployment: replicas at zero, pods in Pending, CrashLoopBackOff or stuck without errors.
- Check the consumer’s credential and permission on the queue: a changed password or a removed SAS policy stop consumption without taking the pod down.
- If the consumer is alive but stuck, restarting the deployment (Pro and Team) usually unblocks it; then look for the cause in the log.
High backlog (attention above the threshold, critical at 10 times)
Ready messages passed the queue threshold: by default, 1,000 for attention and 10,000 for critical. Producers are faster than consumers.
- Compare in and out per minute (RabbitMQ) or the net change (Service Bus), and the 5-minute trend: a high backlog that is going down is recovering.
- Scale the consumers, by hand (Pro and Team) or with an HPA or KEDA, if consumption is limited by the replica count.
- If the time per message went up, the cause is in the consumer or its dependency, not in the queue.
- If that queue normally lives above 1,000, adjust its threshold in the queue drawer (Active message threshold, or Use the default). On the web, only admins set it, and it applies to the organization; on desktop, the threshold stays on your machine.
Alert limits and timings
- Each queue’s backlog threshold changes in the queue drawer (Active message threshold, or Use the default). On the web, only admins set it, on Pro and Team; it applies to the organization and every change goes to the audit log (definir_limiar, remover_limiar). On desktop, the threshold stays on your machine only, per client and queue.
- Chart history lives in memory only: in the browser on the web, in the app on desktop. It covers the last 60 minutes since the screen opened; reloading or closing starts over.
- The service read is cached for 10 seconds per client and service; the dashboard reads at most 1,000 queues and, after 3 errors in a row, retries every 60 seconds.