RabbitMQ: monitor queues, DLQ and purge
The RabbitMQ menu monitors each client’s queues on a dashboard, finds each queue’s DLQ and purges with confirmation. It shows up on clients with a management API URL, user and password saved, and everything goes through the RabbitMQ Management HTTP API.
What it does
| Plan | What it shows and does |
|---|---|
| Free | The queue dashboard and every queue in every vhost the user can see: ready, unacked and total messages, consumers and each queue’s DLQ. Read-only. |
| Pro and Team | Peek the first messages of the queue or its DLQ, with properties, headers and body. Purge the whole queue or only the DLQ. |
RabbitMQ queue dashboard
The Panel tab (Dashboard on desktop) is the one the RabbitMQ menu opens on: every queue of the client on one screen, refreshed every 15 seconds (you can switch to 30 s or 60 s, or pause). Counts and names only: message content never enters the dashboard.
| Plan | What it shows |
|---|---|
| Free | The totals (queues, ready messages, DLQ and queues with a DLQ) and the counts table (ready, unacked, consumers and the paired DLQ), with search, sorting (DLQ then ready, largest first), the "Only with DLQ" filter and "Group variants", which joins sibling queues by name (-retry, -dlq, -error, -memory-error, -resume, -reprocessing). |
| Pro and Team | In and out per minute, from RabbitMQ’s counters (in: published; out: acked plus delivered without ack), the backlog trend over the last 5 minutes, the client chart and each queue’s chart, and the alerts below. |
Clicking a queue opens its details in a drawer, like a pod’s. Peeking the queue or its DLQ opens the messages in the bottom panel, like logs, and purging asks you to type the name, under the rules on this page.
Queue dashboard: alerts and limits
- Each queue’s backlog threshold changes in the queue drawer (Active message threshold, or Use the default). On the web, only admins set it, on Pro and Team; it applies to the organization and every change goes to the audit log (definir_limiar, remover_limiar). On desktop, the threshold stays on your machine only, per client and queue.
- Chart history lives in memory only: in the browser on the web, in the app on desktop. It covers the last 60 minutes since the screen opened; reloading or closing starts over.
- The service read is cached for 10 seconds per client and service, one read at a time; purging clears the cache. The dashboard reads at most 1,000 queues and says when the list was cut.
- After 3 errors in a row, the dashboard retries every 60 seconds and shows a notice.
| Alert | When it shows | Level |
|---|---|---|
| DLQ | The DLQ has messages: 1 to 99 | Attention (100 or more: critical) |
| DLQ growing | The DLQ went up in the last 5 minutes, counted from the last drop, with at least 60 s of history | Critical |
| DLQ spike | The DLQ gained 200 or more messages between two samples less than 60 s apart, in the last 5 minutes | Critical |
| No consumption | The backlog only went up in the last 10 minutes and the out rate stayed under 0.1 message per minute | Critical |
| High backlog | Ready messages above the queue threshold (1,000 by default) | Attention (10 times the threshold: critical) |
How the RabbitMQ DLQ is found
- First, the queue named in x-dead-letter-routing-key, when x-dead-letter-exchange is the default exchange ("") and that queue exists in the same vhost.
- Otherwise, a sibling queue with the same name and the suffix .dlq, -dlq or _error (in that order).
- Other names (such as .dead-letter) are not paired: the queue shows on its own.
- The DLQ is a queue like any other: peeking and purging the DLQ use its name.
Peeking does not consume, but it moves things
RabbitMQ has no read-without-consuming. Kubepier uses the API get with ack_requeue_true: messages go back to the queue, but flagged as redelivered and possibly in a different position. The screen warns you first.
Limits and timeouts
| Limit | Value |
|---|---|
| Messages per peek | 20 by default, 1 to 100 |
| Body of each message | up to 64 KB |
| Purges per client and service | 5 every 10 minutes |
| Timeout of each API call | 15 seconds |
Confirmations
Purging asks you to type the exact queue (or DLQ) name. Without the right name, nothing is deleted.
Audit
- Web: every purge goes to the organization audit log, in the database: who, when, client, vhost, queue, how many messages there were and the error, if any.
- Desktop: every purge goes to the kubepier-audit.log file, in the app data folder: when, who, client, vhost and queue, result and how many were removed.
- Web: every backlog threshold set or removed on the dashboard (definir_limiar, remover_limiar).
- Peeking is not audited. Message bodies never go to the audit log or any log.
Permissions you need on your side
| Feature | RabbitMQ user |
|---|---|
| List queues | management tag (sees the vhosts it has permissions on) or monitoring (sees all) |
| Peek | read permission on the queue, in the vhost |
| Purge | read permission on the queue, in the vhost (RabbitMQ requires read to purge) |
Create a user just for Kubepier, with the management tag and read limited to the queues that matter (configure and write empty: ^$).
What is not accepted
- Peek or purge on Free.
- Purge without typing the queue name, or as a member (on the web; on desktop, from 2.7.0).
- More than 5 purges in 10 minutes on the same client and service.
- Queues in vhosts the user cannot see.
Network and connection route
In the RabbitMQ form, pick the Connection route. Direct: allow the egress IPs on the firewall, for the management API port. Through the cluster: for RabbitMQ on a private network, through a relay in one of the client’s clusters. On the Direct route, a host that resolves only to a private IP answers host_privado right away, without trying to connect. A connection closed by the server (conexao_encerrada) is usually a firewall without the egress IPs, the wrong TLS setting or port, or a connection limit: allow the IPs or switch to Through the cluster.
Common errors
- "no permission on the service" / 401 or 403: the user lacks the tag or the read permission above.
- "too many purges" (429 on the web): wait: 5 purges per client and service every 10 minutes.
- 404: the queue or vhost no longer exists.
- Connection error: check the management API URL (port 15672, or 443 on managed services) and, on the web, the egress IPs.