Azure Service Bus: monitor queues, dead letters and purge
The Azure Service Bus menu monitors each client’s queues and subscriptions on a dashboard, shows dead-letter (DLQ) messages without locking anything and purges with confirmation. It shows up on clients with the namespace connection string saved.
What it does
| Plan | What it shows and does |
|---|---|
| Free | The queue dashboard and the queues and topics with their subscriptions, and the active, dead-letter and scheduled message counts. Read-only. |
| Pro and Team | Peek the first messages of the queue, subscription or DLQ, with properties, dead-letter reason and body. Purge the queue, the subscription or only the DLQ. |
Azure Service Bus queue dashboard
The Panel tab (Dashboard on desktop) is the one the Azure Service Bus menu opens on: every queue and subscription of the client on one screen, refreshed every 15 seconds (you can switch to 30 s or 60 s, or pause). Counts and names only: message content never enters the dashboard.
| Plan | What it shows |
|---|---|
| Free | The totals (queues, ready messages, DLQ and queues with a DLQ) and the counts table (active, DLQ, scheduled and transfer), with search, sorting (DLQ then ready, largest first), the "Only with DLQ" filter and "Group variants", which joins sibling queues by name (-retry, -dlq, -error, -memory-error, -resume, -reprocessing). |
| Pro and Team | The net change of active messages per minute (the Service Bus API does not report in and out separately), the backlog trend over the last 5 minutes, the client chart and each queue’s chart, and the alerts below. |
Clicking a queue opens its details in a drawer, like a pod’s. Peeking the queue or its DLQ opens the messages in the bottom panel, like logs, and purging asks you to type the name, under the rules on this page.
Queue dashboard: alerts and limits
- Each queue’s backlog threshold changes in the queue drawer (Active message threshold, or Use the default). On the web, only admins set it, on Pro and Team; it applies to the organization and every change goes to the audit log (definir_limiar, remover_limiar). On desktop, the threshold stays on your machine only, per client and queue.
- Chart history lives in memory only: in the browser on the web, in the app on desktop. It covers the last 60 minutes since the screen opened; reloading or closing starts over.
- The service read is cached for 10 seconds per client and service, one read at a time; purging clears the cache. The dashboard reads at most 1,000 queues and says when the list was cut.
- After 3 errors in a row, the dashboard retries every 60 seconds and shows a notice.
| Alert | When it shows | Level |
|---|---|---|
| DLQ | The DLQ has messages: 1 to 99 | Attention (100 or more: critical) |
| DLQ growing | The DLQ went up in the last 5 minutes, counted from the last drop, with at least 60 s of history | Critical |
| DLQ spike | The DLQ gained 200 or more messages between two samples less than 60 s apart, in the last 5 minutes | Critical |
| No consumption | The backlog only went up in the last 10 minutes and the out rate stayed under 0.1 message per minute | Critical |
| High backlog | Ready messages above the queue threshold (1,000 by default) | Attention (10 times the threshold: critical) |
View dead-letter messages without locking anything (peek)
Kubepier uses the AMQP peek (the official SDK’s peekMessages): the message is not locked, the delivery count does not change and it works on the DLQ.
How purging the queue and the DLQ works
- Service Bus has no purge operation. Kubepier receives and deletes (receive-and-delete) in batches, up to the number of messages there were when the purge started, never more.
- Scheduled messages and messages locked by a consumer are not reached and stay.
- Received bodies are dropped without being read.
- The result says how many were removed and, if the purge stopped early, how many are left. Run it again to continue.
- On a session-enabled queue or subscription, only the DLQ can be purged: receiving the main queue would mean accepting session by session.
Limits and timeouts
| Limit | Value |
|---|---|
| Messages per peek | 20 by default, 1 to 100 |
| Body of each message | up to 64 KB |
| Purge batch | 250 messages |
| Ceiling of one purge | 50,000 messages or 5 minutes |
| Purges per client and service | 5 every 10 minutes |
| Entities listed | up to 1,000 queues, 1,000 topics and 1,000 subscriptions per topic |
| Empty wait that ends the purge | 2 consecutive waits of 3 seconds |
| Timeout of each management API call | 15 seconds |
Confirmations
Purging asks you to type the name: the queue name, or topic/subscription for a subscription.
Audit
- Web: every purge goes to the organization audit log, in the database: who, when, client, entity, whether it was the DLQ, how many there were, how many were removed, why it stopped and the error, if any.
- Desktop: every purge goes to the kubepier-audit.log file, in the app data folder, with the entity (and /$deadletterqueue for the DLQ), the result and how many were removed.
- Web: every backlog threshold set or removed on the dashboard (definir_limiar, remover_limiar).
- Peeking is not audited.
Permissions you need on your side
| Feature | SAS policy right |
|---|---|
| List queues, topics and counts | Manage (the management API does not accept Listen alone) |
| Peek | Listen |
| Purge | Listen to receive and delete, and Manage to read the starting count |
Since listing requires Manage, the menu needs a policy with Manage in practice (which includes Listen and Send). Create a policy just for Kubepier on the namespace.
What is not accepted
- Peek or purge on Free.
- Purge without typing the name, or as a member (on the web; on desktop, from 2.7.0).
- Purging the main queue of a session-enabled entity (DLQ only).
- A connection string that is not for a Service Bus namespace.
Network and connection route
In the Service Bus form, pick the Connection route. Direct: if the namespace restricts IPs, allow the egress IPs under Networking → Selected networks. Through the cluster: for a namespace with a private endpoint; through the route, Service Bus uses AMQP over WebSocket on port 443. On the Direct route, a host that resolves only to a private IP answers host_privado right away, without trying to connect. A connection closed by the server (conexao_encerrada) is usually a firewall without the egress IPs, the wrong TLS setting or port, or a connection limit: allow the IPs or switch to Through the cluster.
Common errors
- 401 or 403: the SAS policy lacks Manage (listing) or Listen (peek and purge).
- "Queue or subscription not found": the entity was deleted or renamed.
- Purge stopped with messages left: it hit the count or time ceiling; run it again.
- "too many purges" (429 on the web): 5 purges per client and service every 10 minutes.