Protocols and data

IoT sensor data: timestamps, quality and gaps

Define IoT sensor points that survive outages: identity, units, source timestamps, quality, stale limits, counter rollover and the commissioning tests that prove them.

A Gateway loses its uplink at 09:00 and gets it back at 12:00. It then forwards the three hours of readings it held. If the receiving platform keeps each reading's original timestamp, the chart fills the gap. If the platform stamps each reading when it arrives, three hours of energy land in one 15-minute interval. At an average load of 40 kW that is 120 kWh in 15 minutes, and the demand report shows a 480 kW peak that never happened.

Every value in that example was real. Only the time was wrong, and a live dashboard did not show it.

Each observation must answer four questions. What property was observed, and where? When did the condition apply? Is the value valid and fresh enough for this use? Which measurement and processing chain produced it? Billing, protection, safety functions and regulatory reporting add their own equipment, approval, uncertainty and audit requirements on top of what follows.

Observation models in SSN and SensorThings

The W3C Semantic Sensor Network ontology (SOSA/SSN) defines an observation as an act of observing, typically to estimate or determine the value of a property of a feature of interest. It keeps the property, result, sensor, procedure, phenomenon time and result time as separate terms. It also states that phenomenon time is not necessarily the same as result time.

The OGC SensorThings API uses a similar model for sensor web services. An Observation belongs to a Datastream and carries a result, phenomenon time, result time, result quality, valid time and parameters.

You need not adopt either model. Store meaning, observation time, quality and provenance as separate fields, whatever schema you use.

The minimum observation record

Field or relationshipQuestion it answersFailure if omitted
Site, asset and point identityWhich physical thing and channel produced this?Values from different sources are merged or compared incorrectly
Observed propertyIs this active power, cumulative energy, temperature or status?A number has no stable engineering meaning
Value and unitWhat magnitude and unit were reported?W, kW and MW, or °C and °F, are silently confused
Phenomenon or source timeWhen did the measured condition apply?Buffered old data looks current
Result or ingestion timeWhen was the result produced or received here?Transport delay and replay cannot be diagnosed
Quality or stateIs the value good, uncertain, bad, stale, substituted or unavailable?Invalid data looks authoritative
Procedure and configuration identityWhich range, scaling, firmware, calibration and calculation applied?A configuration change looks like a physical change
Sequence or stable event identityHas this observation already been processed?Retries become duplicate records or double-counted energy

Here is one record for the energy register on a three-phase distribution board, reporting on 15-minute clock boundaries. It is the 09:45 reading from the outage above, delivered after the uplink came back:

JSON
{
  "site": "plant-02",
  "asset": "db-3-compressor-house",
  "point": "db-3/energy-active-total",
  "property": "active energy, cumulative register",
  "value": 184213.4,
  "unit": "kWh",
  "observed_at": "2026-09-14T09:45:00Z",
  "received_at": "2026-09-14T12:02:41Z",
  "quality": "good",
  "config_revision": "point-list r7",
  "sequence": 88412
}

The two times differ by 2 hours 17 minutes. That difference tells the consumer the reading was held during an outage and is not current load. The sequence number lets the consumer discard a second copy of the same reading.

Not every source can fill every field. Record which fields a source cannot fill, and do not invent precision. For example, a third-party Zigbee sensor reports a value with no timestamp, so the Gateway stamps it on arrival. That timestamp is an arrival time. Label it as one in the point list.

Designing the point

Define the measurement first

Start with the measurand: the quantity you intend to measure. "Temperature" is incomplete if the decision depends on air temperature at a named location, pipe surface temperature, immersion temperature or an average across rooms.

Write down the range, the required uncertainty over that range, the response time and the installation conditions. Write down the decision, alarm or calculation that uses the value. Then write down what happens if the value is wrong, missing or late. That last entry sets the stale limit and the alarm behaviour later.

A component's accuracy label does not give the performance of the installed chain. NIST's metrological-traceability policy treats traceability as a property of a measurement result. A documented chain of calibrations supports it, and each calibration adds to the measurement uncertainty. A calibrated instrument alone does not make every later result traceable. For electricity, the high-accuracy metering guide separates meter class, complete-chain uncertainty and legal approval. For an RTD, the Pt100 and Pt1000 calculator gives the standard curve and the tolerance-class limit. Add the transmitter, wiring and installation effects to that.

Keep scaling and conversions traceable

A raw register or analogue input can need signed decoding, word ordering, a CT or VT ratio, linear scaling or a non-linear curve before it becomes an engineering value. Store the parameters, or the configuration revision that applied them. A wrong word order can decode to a plausible number, so decode one raw value by hand. The Modbus commissioning guide has a worked decode.

For a calculated point, also keep the input point identities, the formula and its revision, and the units before and after conversion. Record the time-alignment rule and what the formula does with a bad, missing or stale input. Mark any substitution, clamping or manual override on the result.

Do not replace a bad value with zero. Zero is a valid reading for a stopped motor or an empty tank, and it leads to a different operational decision from "unknown".

Keep observation time separate from arrival time

A reading carries the time the condition applied and the time each system received it, and the two separate whenever data is buffered and replayed. Telemetry timestamps, source time and arrival time covers which time to keep, UTC and summer time, and how to test replay on site.

Transport guarantees stop at each hop

MQTT 5.0 defines three quality-of-service levels between one sender and one receiver: at most once, at least once and exactly once. It also has message expiry and session state. The MQTT over TLS guide covers broker identity, client authentication and topic access.

None of this proves that the physical reading was correct or that every sample became a message. It does not prove that the payload has the right point, unit or timestamp, that the next integration hop kept the same guarantee, or that the consumer stored the message once. A retained message is the last one published on its topic, and it can be hours or days old.

Give each observation a stable identity. For a metered point, the point identity plus the observation time is usually enough. Make writes idempotent, so that a retried batch does not count twice. Test disconnect and replay on site. Do not infer the behaviour from a protocol name.

Make quality and freshness visible

A consumer must be able to tell a current value from a late, stale, substituted or missing one, and each source's quality states need mapping to yours on purpose. Stale, missing and invalid telemetry sets out those states, how to choose a stale limit and how to show it beside the value.

Aggregate from raw evidence

Each quantity needs its own aggregation function. Mean power over an interval, maximum demand, last-known state, alarm count and the change in a cumulative register all differ. Record the interval boundaries and time zone. Record whether each interval is labelled by its start or its end, the minimum completeness rule and how late data revises a closed interval. Mark each result as measured, aggregated, estimated or substituted.

Never average a cumulative energy register. Take the difference between the readings at the interval boundaries. Never add instantaneous power readings without multiplying each one by the time it represents.

Counter rollover and counter reset both produce a negative difference. They need different treatment. Take a 32-bit Wh counter, which wraps from 4,294,967,295 to 0. At an average of 40 kW it takes about 12 years to get there, so few systems test the wrap before it happens.

At 23:45 the register reads 4,294,960,296 Wh. At 00:00 it reads 3,000 Wh. As a rollover, the interval energy is (4,294,967,296 − 4,294,960,296) + 3,000 = 10,000 Wh. That is 10 kWh in 15 minutes, an average demand of 40 kW:

Now the meter is replaced. The last reading from the old meter is 1,250,400 Wh and the first from the new meter is 3 Wh. The same wrap formula gives 4,293,716,899 Wh, about 4.3 GWh in 15 minutes.

Accept a wrap only when the energy it implies is possible in the interval. A 100 kW supply can deliver at most 25 kWh in 15 minutes. A larger result is a reset. Close that interval as incomplete, start a new baseline and record the event. On a ZEM-6x electricity monitor such as the ZEM-65, Edge shows the register's rollover point, in kWh, as a read-only setting. The per-phase kWh registers are writable, so an engineer can preset them. Log every preset as a reset.

Keep the cumulative register alongside any interval values. After a telemetry gap, the register difference recovers the total energy across the gap. It cannot recover the power profile inside the gap.

Sampling, reporting and dashboard refresh

"One-minute data" can mean five different things. A sensor sampled once a minute. A meter sampled fast and calculated a one-minute mean. A device published its latest value once a minute. A platform stored a one-minute roll-up. A dashboard refreshed once a minute over a point that changes at another rate. State each rate that matters in the point list.

EpiSensor Zigbee sensors show the distinction. They can report on wall-clock boundaries set in minutes, at a fixed rate in seconds, or on change.

Rate has a cost. A point that reports every 10 s produces 8,640 messages a day. At every 60 s it produces 1,440. A thousand points at one report a minute produce 1.44 million points a day before filtering. On a battery sensor, each report also costs radio airtime and battery life. Check which reporting interval the datasheet's battery-life figure assumes.

UseTypical intervalEvidence to keep
Energy allocation and billingThe settlement or tariff period, commonly 15 or 30 min, plus the cumulative registerBoundaries, completeness, register resets and any estimation
Operational trend1 min for power, 5 to 15 min for room temperatureSource rate, aggregation and display rate
AlarmDetection delay plus stale limit must fit inside the permitted response timeThreshold, persistence, deadband, stale limit and notification delay
Demand managementWell inside the demand window, for example 1 min in a 15 or 30 min window, to leave time to actWindow definition, update rate, threshold logic and command evidence
Condition monitoring1 min trends of current or temperature for slow degradation. Vibration spectra need kHz sampling on a dedicated instrumentSensor bandwidth, sampling method, feature calculation and baseline
Power quality, protection or safetyInstruments and methods qualified for that function, such as IEC 61000-4-30 for power qualityApplicable standard, instrument configuration, event timing and acceptance evidence

For what current and power data can show about equipment condition, read condition monitoring from electrical data.

Commissioning tests that expose bad data

Run each test on the complete path, from the sensor to the receiving platform. Write down the expected result before you run the test.

  1. Trace each displayed point to the physical device, channel and asset label.
  2. Apply at least two known values, one of them not zero, and check every conversion. For a CT channel, compare the current in Edge and in the platform with a calibrated clamp meter reading on the same conductor.
  3. Check polarity, the import and export sign convention, and the phase association. A reversed CT gives negative import power on that phase. Correct it at the CT or with the meter's per-phase CT direction setting, and record which one you changed.
  4. Compare the source and receiving clocks. Then hold data back for a known time and check that both timestamps stay distinct.
  5. Disconnect or invalidate the source. The application must show bad or unavailable, not zero.
  6. Stop updates without clearing the last value. The point must become stale at the agreed limit.
  7. Break the onward connection for a fixed time, for example 30 minutes. Watch the queue grow. Reconnect, then check replay order, timestamps and completeness at the receiver.
  8. Replay a record or retry a delivery. An idempotent consumer must not count it twice.
  9. Restart each layer in turn: sensor, Gateway, broker and platform. Check that identity, configuration, clock and delivery recover.
  10. Compare one closed interval with the source readings, including an interval boundary, a missing value and a counter reset.
  11. Change an approved range or formula revision. The change must appear in the lineage. It must not change the meaning of earlier data.
  12. Compare important measurements with an independent reference under representative load.

The acceptance record names the equipment, configuration, reference, test time, expected result, observed result and the person who approved it.

Failure modes seen on site

SymptomLikely causeHow to detect it
Negative import power on one phaseCT reversed or on the wrong phaseSign of each phase's power against the known load, and a power factor far from the expected value
Values 1,000 times too large or too smallW reported as kW, or a missing scale factorComparison with the meter display or a clamp meter at commissioning
Huge or near-zero values after a register-map changeWrong word order or data typeA hand decode of one raw register
An hour missing in March and doubled in OctoberIntervals keyed on local timeA count of periods per local day: 92, 96 or 100
A gap, then a burst of readings stamped at the same minuteReplayed data stamped on arrivalReceived time minus observed time, per reading
Readings dated years in the past or futureA device or gateway clock not set after power lossObserved time outside a plausible window around received time
A flat line at the last valueThe source stopped and the display holds its last valueStale state and the age of the newest observation
An energy step of several GWh in one intervalA meter reset read as a rolloverA limit on the maximum energy possible in the interval

Monitoring data quality in operation

Commissioning proves one state on one day. Operation must detect when that state changes. For each point, track expected against received observations and the age of the newest valid observation. Count late, out-of-order, duplicate, bad, uncertain, substituted and out-of-range readings. Track clock offset, queue depth and the age of the oldest queued item. Log rejected schema, unit or identity changes, and reconcile interval totals against the cumulative register.

Alert on thresholds that match the use. For example: newest observation older than the stale limit, oldest queued item older than one hour, or clock offset above 5 s for one-minute data.

Define the denominator before you quote availability. "99.9% data availability" means nothing until the expected point set, reporting schedule, exclusions, quality rule and treatment of late data are fixed. For one point at one-minute reporting, 99.9% over 30 days still allows 43 missing readings.

ISO/IEC 25012 gives a general model for the quality of structured data. It has 15 characteristics, split between an inherent view and a system-dependent view, and ISO last reviewed and confirmed it in 2025. It supplies the vocabulary. The project still has to turn the relevant characteristics into measurable requirements.

Sensor data with EpiSensor

EpiSensor devices, such as ZEM electricity monitors and ZHT temperature and humidity sensors, report over Zigbee to a ZGW-20 Gateway. The Gateway runs EpiSensor Edge, which keeps local history (30 days by default), runs calculations and rules, and forwards data to other platforms.

Readings from EpiSensor Zigbee devices reach Edge with a timestamp on each data point, and Edge keeps that timestamp. Edge stamps a third-party Zigbee device's value with the Gateway clock when it arrives. Edge discards a data point with an invalid timestamp and logs a warning.

Edge calculates quality for each sensor over a 15-minute window. For a sensor on a fixed schedule, quality is a percentage of expected points. A sensor that reports every minute expects 15 points in the window, so 13 received gives 87%. For a sensor that reports on change, or less often than every 15 minutes, Edge reports a count of received points instead. Read the quality type before you compare two sensors.

When a supported destination fails, Edge moves the undelivered data to a queue on disk. Some integrations own their retry policy instead, so check the integration's guide. It checks the destination every minute and replays the queue when the destination returns. Replayed readings keep their original timestamps. They do not run local calculations or automation again. For MQTT, Edge counts a delivery as complete when the broker confirms the publish. A batch can arrive twice if the destination received it before Edge recorded completion, so the receiver must discard duplicates. Pending replay data has no age limit, so a long outage keeps growing the queue. Size the Gateway's free storage for the longest outage you plan to ride through.

Inside Edge, each reading carries a sensor ID, an export ID made from the device serial number and sensor ID, a timestamp and a value. The unit, scaling and quality come from the device configuration. Agree in the point list how the receiving platform gets them. Also agree which layer applies scaling, how Edge quality maps to the platform's quality model, and what alarms and control do when an input goes stale.

Once the point list is fixed, continue to IoT data-storage architecture to choose where each store sits.

Common questions

What should an IoT sensor-data record contain?

It needs the site, asset and point identity, the observed property, the value and unit, the time the observation applies to, the time it was received, and a quality state. Billing, sub-metering allocation and demand response settlement also need the scaling, calibration and calculation revision that produced the value.

What is the difference between a source timestamp and an ingestion timestamp?

A source timestamp says when the measured condition applied. An ingestion timestamp says when another system received or stored the record. The two differ during buffering, retries and clock errors. If you replace the source time with the ingestion time, old data looks current.

Does MQTT quality of service guarantee complete sensor data?

No. MQTT quality of service covers delivery between one sender and one receiver. It does not prove that the sensor measured correctly or that every sample became a message. It does not prove that the next application stored the message once, or that the payload has the correct unit and timestamp.

How should an IoT reporting interval be selected?

Start from the fastest change the application must detect and from the settlement or calculation period. Then check the permitted detection delay, device and network capacity, and storage. Sampling, calculation, publication and dashboard refresh are separate rates. Set each one on purpose.

Is a zero sensor value the same as missing or stale data?

No. Zero can be a valid measurement. Missing means no value arrived for an expected interval. Stale means the latest value is older than the freshness limit for that use. Store and show the three states separately.