Protocols and data

MQTT QoS, retained messages and stale data

MQTT QoS 0, 1 and 2, what each acknowledgement proves, retained messages, persistent sessions, last will and message expiry, and how to avoid showing stale telemetry as current.

A dashboard connects to an MQTT broker at 08:00 and shows a meter at 42 kW. The meter lost power at 18:00 the day before. The broker sent the dashboard the last retained message for the meter's topic, and nothing in the MQTT packet said that the value was 14 hours old.

This guide explains what each MQTT delivery setting proves: the three QoS levels, retained messages, persistent sessions, the last will and message expiry. Section numbers refer to the OASIS MQTT 5.0 standard. Where MQTT 3.1.1 differs, the guide says so.

This guide is part of the OPC UA and MQTT series, which starts with OPC UA vs MQTT vs Modbus. For TLS, certificates and topic permissions, read the MQTTS guide.

The three QoS levels

QoS (quality of service) sets the acknowledgement exchange on each hop. Section 4.3 of the standard defines three levels:

QoSNamePackets per message on one hopResult on that hop
0At most oncePUBLISHDelivered once or lost. No acknowledgement.
1At least oncePUBLISH, PUBACKDelivered, but can arrive more than once.
2Exactly oncePUBLISH, PUBREC, PUBREL, PUBCOMPDelivered once, with no duplicates on that hop.

QoS is per hop

A message crosses two hops: publisher to broker, and broker to each subscriber. QoS applies to each hop separately. The broker delivers at the lower of two values: the QoS of the PUBLISH, and the maximum QoS granted to the subscription. A reading published at QoS 2 to a QoS 0 subscription arrives at QoS 0.

A PUBACK proves only that the broker accepted the message. It does not prove that a subscriber received it, or that an application stored it. In MQTT 5 the PUBACK carries a reason code (section 3.4.2.1). Code 0x00 Success and code 0x10 No matching subscribers both mean acceptance. A broker may send 0x10 when it knows that no client subscribes to the topic, but it does not have to, so 0x00 does not prove that a subscriber exists. Codes 0x80 and above reject the message, for example 0x87 Not authorized and 0x97 Quota exceeded. A publisher that counts every PUBACK as a success loses those readings without an error.

MQTT 3.1.1 has no reason codes. A broker that refuses a publish must either acknowledge it as normal or close the connection (section 3.3.5 of 3.1.1). An acknowledged publish can therefore still be a refused one.

Choosing a level for telemetry

DataUsual choice
Frequent readings where one lost sample does not matterQoS 0
Energy readings, counters and events that must arriveQoS 1, with duplicates removed at the receiver
Commands and one-off transactionsQoS 1 or 2, an expiry on the command, and an application-level acknowledgement

QoS 1 is the usual choice for energy data. It needs two packets per message on each hop, and QoS 2 needs four. Some services do not offer QoS 2: AWS IoT Core supports QoS 0 and 1 only.

Give each command an expiry as well as a QoS. A command that waits in a queue during an outage is sent when the client reconnects. An MQTT 5 message expiry interval (section 3.3.2.3.3), or a valid-until time in the payload, stops last night's setpoint from reaching the plant this morning.

Duplicates

At QoS 1, the sender resends every PUBLISH that was not acknowledged when the connection dropped (section 4.4). If the receiver had already processed the first copy, it gets the reading twice.

The protocol fields do not identify the duplicate. The broker sets the DUP flag for its own retries on the outgoing hop, and it does not pass on the DUP flag that it received (section 3.3.1.1). The packet identifier belongs to one hop, and the sender can use it again as soon as the PUBACK arrives (section 2.2.1).

Remove duplicates with an application key: the device identity, the channel and the measurement timestamp, or a sequence number that the source increments for each reading. For example:

JSON
{"device":"meter-12","channel":"p_total","ts":"2026-09-23T14:05:00Z","seq":48213,"value":42.0,"unit":"kW"}

A receiver that stores readings with a unique key on device, channel and ts discards the second copy. Write the timestamp in UTC, as RFC 3339 text or as epoch milliseconds, and keep the source clock synchronised with NTP. The timestamps guide explains measurement time and arrival time.

Retained messages

When a PUBLISH has the retain flag set, the broker stores it as the last value of that topic (section 3.3.1.3). The broker keeps one retained message per topic. A new retained message replaces the old one. A retained message with an empty payload deletes it. If the retained message was published at QoS 0, the broker may discard it at any time.

The broker sends the retained message to each new subscription on a matching topic, with the retain flag set. Messages that arrive later from a live publisher reach the subscriber with the retain flag cleared. A subscriber can use the flag to tell a stored value from a live one. The flag does not tell it how old the stored value is. In MQTT 5, a subscription with Retain As Published = 1 receives the flag as the publisher set it. Bridges use this option, so a subscriber behind a bridge cannot rely on the flag.

In MQTT 5, the subscriber also controls whether it receives retained messages at all (section 3.8.3.1). Retain Handling 0 sends them on every subscribe, 1 sends them only for a new subscription, and 2 never sends them. MQTT 3.1.1 has no such option.

Retain state that a new subscriber needs at once, such as a device's configuration or its online status. Retained telemetry causes the problem in the introduction. Three measures prevent it:

  1. Put the measurement timestamp in every payload, and make the subscriber check it. Mark a reading stale after a fixed number of missed intervals. Three intervals gives 3 minutes for a meter that reports every minute.
  2. Set an MQTT 5 message expiry interval equal to the staleness limit, for example 180 s. When the interval passes, the broker discards the message, including a retained one. The broker also reduces the interval by the time that the message waited, so the subscriber can see how long it has left. MQTT 3.1.1 has no message expiry, so the timestamp check is the only protection.
  3. Publish readings without the retain flag, and retain only a status topic.

Sessions and the last will

A persistent session keeps a client's subscriptions while it is disconnected. The broker queues QoS 1 and QoS 2 messages that arrive for the client, and resends messages that were not acknowledged when the connection dropped. Queuing QoS 0 messages is optional in the standard (section 4.1).

In MQTT 5, the client connects with Clean Start = 0 and a Session Expiry Interval greater than 0, for example 86,400 s for one day. Clean Start = 1 discards the old session. When the Session Expiry Interval is 0 or absent, the session ends when the connection closes (section 3.1.2.11.2). In MQTT 3.1.1, the client connects with Clean Session = 0, and the protocol sets no session expiry. In both versions, the Session Present flag in the CONNACK tells the client whether the broker still had its session.

The broker's queue has a limit. Mosquitto holds up to 1,000 QoS 1 and 2 messages per client above those in flight (max_queued_messages), and drops messages over that limit. By default it does not queue QoS 0 messages for a disconnected client (queue_qos0_messages false). A subscriber that receives 6 readings a minute fills a 1,000-message queue in under 3 hours. Size the queue from the message rate and the longest outage that you expect, or keep the history at the publisher.

The last will is a message that the broker publishes for the client when the connection ends without a normal DISCONNECT: after a network failure, a missed keep-alive or a closed socket (section 3.1.2.5). A status topic uses it like this:

  1. The gateway connects with a will of "offline" on its status topic, with Will Retain = 1.
  2. When the connection is up, the gateway publishes "online" to the same topic with the retain flag.
  3. Before a planned shutdown, the gateway publishes "offline" with the retain flag, then disconnects.

Without Will Retain, the broker sends "offline" to current subscribers only. The retained "online" stays on the topic, and a dashboard that connects later shows the gateway as online. Step 3 is necessary because a DISCONNECT with reason code 0x00 deletes the will without publishing it.

MQTT 5 adds a Will Delay Interval (section 3.1.3.2.2). If the client reconnects before the delay ends, the broker does not publish the will. A delay of 30 s stops a short network drop from showing the gateway as offline. The broker publishes the will when the delay ends or when the session ends, whichever is first. With a Session Expiry Interval of 0, the session ends at disconnect and the delay has no effect. MQTT 3.1.1 has no will delay.

Freshness checklist

For each topic that carries telemetry, record:

  • the QoS for publishing and for each subscription;
  • the PUBACK reason codes that the publisher treats as a failure;
  • whether messages are retained, and why;
  • the message expiry, if any;
  • where the measurement timestamp is in the payload, and its format;
  • the key that the receiver uses to remove duplicates;
  • the session expiry and the broker's queue limit;
  • the status topic and last will of the source, with Will Retain set;
  • how old a reading may be before a dashboard or a calculation treats it as stale.

The stale data guide explains how to mark and handle readings that are too old.

MQTT with Edge

The ZGW-20 Gateway runs a local MQTT broker next to Edge, and Edge publishes readings to the MQTT destinations of energy platforms, over TLS where the platform requires it. Every reading that Edge sends carries its measurement timestamp. For MQTT, Edge treats a publish as complete when the broker confirms it. What the platform does with the message after that is the platform's part of the delivery, so check storage at the platform when you commission the link. Connecting Edge to your platform covers the delivery options. Edge on the ZGW-20 Gateway publishes readings to a platform's MQTT broker, over MQTTS where the platform requires it. Edge publishes at QoS 1, and a delivery is complete only when the broker's PUBACK arrives. A disconnect or a publish error puts the readings in a local outbox on the Gateway. By default, Edge tries the outbox again every minute while the destination is available. A replayed reading keeps its original measurement timestamp. The outbox has no age limit. By default, a low-disk guard pauses new recording when free space falls below 10% of the filesystem, capped at 1 GiB, or below 256 MiB. Readings that are still in memory when the Gateway loses power are not yet in the outbox, and can be lost. The store-and-forward guide explains how to size it.

The PUBACK proves that the platform's broker accepted the reading. It does not prove that the platform stored it. A replay after a reconnect can also deliver a reading twice, so the platform must remove duplicates with the device, channel and timestamp. When you commission the link, find a known reading in the platform's storage. Compare its timestamp and value with the reading in Edge.

Common questions

Should I use QoS 1 or QoS 2 for sensor data?

Use QoS 1 and remove duplicates at the receiver with the device, channel and measurement timestamp. QoS 2 needs four packets per message on each hop, removes duplicates on that hop only, and is not available on some services: AWS IoT Core supports QoS 0 and 1. A publisher that resends readings after an outage can still create duplicates that QoS 2 cannot see.

How do I delete a retained message?

Publish a message with the retain flag set and an empty payload to the same topic. The broker removes the retained message and does not store the empty one. Current subscribers still receive the empty message, so they must ignore an empty payload rather than parse it as a reading.

Why does a dashboard show a device as online after it has failed?

Usually because the will was sent without Will Retain. The broker publishes "offline" to current subscribers only, and the retained "online" stays on the topic for every later subscriber. The other common cause is a will that never fires: a normal DISCONNECT deletes the will, so the client must publish "offline" itself before a planned shutdown.