A Gateway loses its uplink at 09:00 and gets it back at 12:00. It then forwards the three hours of readings it held. If the receiving platform keeps each reading's original timestamp, the chart fills the gap. If the platform stamps each reading when it arrives, three hours of energy land in one 15-minute interval. At an average load of 40 kW that is 120 kWh in 15 minutes, and the demand report shows a 480 kW peak that never happened.
Every value in that example was real. Only the time was wrong, and a live dashboard did not show it.
Each observation must answer four questions. What property was observed, and where? When did the condition apply? Is the value valid and fresh enough for this use? Which measurement and processing chain produced it? Billing, protection, safety functions and regulatory reporting add their own equipment, approval, uncertainty and audit requirements on top of what follows.
Observation models in SSN and SensorThings
The W3C Semantic Sensor Network ontology (SOSA/SSN) defines an observation as an act of observing, typically to estimate or determine the value of a property of a feature of interest. It keeps the property, result, sensor, procedure, phenomenon time and result time as separate terms. It also states that phenomenon time is not necessarily the same as result time.
The OGC SensorThings API uses a similar model for sensor web services. An Observation belongs to a Datastream and carries a result, phenomenon time, result time, result quality, valid time and parameters.
You need not adopt either model. Store meaning, observation time, quality and provenance as separate fields, whatever schema you use.
The minimum observation record
| Field or relationship | Question it answers | Failure if omitted |
|---|---|---|
| Site, asset and point identity | Which physical thing and channel produced this? | Values from different sources are merged or compared incorrectly |
| Observed property | Is this active power, cumulative energy, temperature or status? | A number has no stable engineering meaning |
| Value and unit | What magnitude and unit were reported? | W, kW and MW, or °C and °F, are silently confused |
| Phenomenon or source time | When did the measured condition apply? | Buffered old data looks current |
| Result or ingestion time | When was the result produced or received here? | Transport delay and replay cannot be diagnosed |
| Quality or state | Is the value good, uncertain, bad, stale, substituted or unavailable? | Invalid data looks authoritative |
| Procedure and configuration identity | Which range, scaling, firmware, calibration and calculation applied? | A configuration change looks like a physical change |
| Sequence or stable event identity | Has this observation already been processed? | Retries become duplicate records or double-counted energy |
Here is one record for the energy register on a three-phase distribution board, reporting on 15-minute clock boundaries. It is the 09:45 reading from the outage above, delivered after the uplink came back:
{
"site": "plant-02",
"asset": "db-3-compressor-house",
"point": "db-3/energy-active-total",
"property": "active energy, cumulative register",
"value": 184213.4,
"unit": "kWh",
"observed_at": "2026-09-14T09:45:00Z",
"received_at": "2026-09-14T12:02:41Z",
"quality": "good",
"config_revision": "point-list r7",
"sequence": 88412
}The two times differ by 2 hours 17 minutes. That difference tells the consumer the reading was held during an outage and is not current load. The sequence number lets the consumer discard a second copy of the same reading.
Not every source can fill every field. Record which fields a source cannot fill, and do not invent precision. For example, a third-party Zigbee sensor reports a value with no timestamp, so the Gateway stamps it on arrival. That timestamp is an arrival time. Label it as one in the point list.
Designing the point
Define the measurement first
Start with the measurand: the quantity you intend to measure. "Temperature" is incomplete if the decision depends on air temperature at a named location, pipe surface temperature, immersion temperature or an average across rooms.
Write down the range, the required uncertainty over that range, the response time and the installation conditions. Write down the decision, alarm or calculation that uses the value. Then write down what happens if the value is wrong, missing or late. That last entry sets the stale limit and the alarm behaviour later.
A component's accuracy label does not give the performance of the installed chain. NIST's metrological-traceability policy treats traceability as a property of a measurement result. A documented chain of calibrations supports it, and each calibration adds to the measurement uncertainty. A calibrated instrument alone does not make every later result traceable. For electricity, the high-accuracy metering guide separates meter class, complete-chain uncertainty and legal approval. For an RTD, the Pt100 and Pt1000 calculator gives the standard curve and the tolerance-class limit. Add the transmitter, wiring and installation effects to that.
Keep scaling and conversions traceable
A raw register or analogue input can need signed decoding, word ordering, a CT or VT ratio, linear scaling or a non-linear curve before it becomes an engineering value. Store the parameters, or the configuration revision that applied them. A wrong word order can decode to a plausible number, so decode one raw value by hand. The Modbus commissioning guide has a worked decode.
For a calculated point, also keep the input point identities, the formula and its revision, and the units before and after conversion. Record the time-alignment rule and what the formula does with a bad, missing or stale input. Mark any substitution, clamping or manual override on the result.
Do not replace a bad value with zero. Zero is a valid reading for a stopped motor or an empty tank, and it leads to a different operational decision from "unknown".
Keep observation time separate from arrival time
A reading carries the time the condition applied and the time each system received it, and the two separate whenever data is buffered and replayed. Telemetry timestamps, source time and arrival time covers which time to keep, UTC and summer time, and how to test replay on site.
Transport guarantees stop at each hop
MQTT 5.0 defines three quality-of-service levels between one sender and one receiver: at most once, at least once and exactly once. It also has message expiry and session state. The MQTT over TLS guide covers broker identity, client authentication and topic access.
None of this proves that the physical reading was correct or that every sample became a message. It does not prove that the payload has the right point, unit or timestamp, that the next integration hop kept the same guarantee, or that the consumer stored the message once. A retained message is the last one published on its topic, and it can be hours or days old.
Give each observation a stable identity. For a metered point, the point identity plus the observation time is usually enough. Make writes idempotent, so that a retried batch does not count twice. Test disconnect and replay on site. Do not infer the behaviour from a protocol name.
Make quality and freshness visible
A consumer must be able to tell a current value from a late, stale, substituted or missing one, and each source's quality states need mapping to yours on purpose. Stale, missing and invalid telemetry sets out those states, how to choose a stale limit and how to show it beside the value.
Aggregate from raw evidence
Each quantity needs its own aggregation function. Mean power over an interval, maximum demand, last-known state, alarm count and the change in a cumulative register all differ. Record the interval boundaries and time zone. Record whether each interval is labelled by its start or its end, the minimum completeness rule and how late data revises a closed interval. Mark each result as measured, aggregated, estimated or substituted.
Never average a cumulative energy register. Take the difference between the readings at the interval boundaries. Never add instantaneous power readings without multiplying each one by the time it represents.
Counter rollover and counter reset both produce a negative difference. They need different treatment. Take a 32-bit Wh counter, which wraps from 4,294,967,295 to 0. At an average of 40 kW it takes about 12 years to get there, so few systems test the wrap before it happens.
At 23:45 the register reads 4,294,960,296 Wh. At 00:00 it reads 3,000 Wh. As a rollover, the interval energy is (4,294,967,296 − 4,294,960,296) + 3,000 = 10,000 Wh. That is 10 kWh in 15 minutes, an average demand of 40 kW:
Now the meter is replaced. The last reading from the old meter is 1,250,400 Wh and the first from the new meter is 3 Wh. The same wrap formula gives 4,293,716,899 Wh, about 4.3 GWh in 15 minutes.
Accept a wrap only when the energy it implies is possible in the interval. A 100 kW supply can deliver at most 25 kWh in 15 minutes. A larger result is a reset. Close that interval as incomplete, start a new baseline and record the event. On a ZEM-6x electricity monitor such as the ZEM-65, Edge shows the register's rollover point, in kWh, as a read-only setting. The per-phase kWh registers are writable, so an engineer can preset them. Log every preset as a reset.
Keep the cumulative register alongside any interval values. After a telemetry gap, the register difference recovers the total energy across the gap. It cannot recover the power profile inside the gap.
Sampling, reporting and dashboard refresh
"One-minute data" can mean five different things. A sensor sampled once a minute. A meter sampled fast and calculated a one-minute mean. A device published its latest value once a minute. A platform stored a one-minute roll-up. A dashboard refreshed once a minute over a point that changes at another rate. State each rate that matters in the point list.
EpiSensor Zigbee sensors show the distinction. They can report on wall-clock boundaries set in minutes, at a fixed rate in seconds, or on change.
Rate has a cost. A point that reports every 10 s produces 8,640 messages a day. At every 60 s it produces 1,440. A thousand points at one report a minute produce 1.44 million points a day before filtering. On a battery sensor, each report also costs radio airtime and battery life. Check which reporting interval the datasheet's battery-life figure assumes.
| Use | Typical interval | Evidence to keep |
|---|---|---|
| Energy allocation and billing | The settlement or tariff period, commonly 15 or 30 min, plus the cumulative register | Boundaries, completeness, register resets and any estimation |
| Operational trend | 1 min for power, 5 to 15 min for room temperature | Source rate, aggregation and display rate |
| Alarm | Detection delay plus stale limit must fit inside the permitted response time | Threshold, persistence, deadband, stale limit and notification delay |
| Demand management | Well inside the demand window, for example 1 min in a 15 or 30 min window, to leave time to act | Window definition, update rate, threshold logic and command evidence |
| Condition monitoring | 1 min trends of current or temperature for slow degradation. Vibration spectra need kHz sampling on a dedicated instrument | Sensor bandwidth, sampling method, feature calculation and baseline |
| Power quality, protection or safety | Instruments and methods qualified for that function, such as IEC 61000-4-30 for power quality | Applicable standard, instrument configuration, event timing and acceptance evidence |
For what current and power data can show about equipment condition, read condition monitoring from electrical data.
Commissioning tests that expose bad data
Run each test on the complete path, from the sensor to the receiving platform. Write down the expected result before you run the test.
- Trace each displayed point to the physical device, channel and asset label.
- Apply at least two known values, one of them not zero, and check every conversion. For a CT channel, compare the current in Edge and in the platform with a calibrated clamp meter reading on the same conductor.
- Check polarity, the import and export sign convention, and the phase association. A reversed CT gives negative import power on that phase. Correct it at the CT or with the meter's per-phase CT direction setting, and record which one you changed.
- Compare the source and receiving clocks. Then hold data back for a known time and check that both timestamps stay distinct.
- Disconnect or invalidate the source. The application must show bad or unavailable, not zero.
- Stop updates without clearing the last value. The point must become stale at the agreed limit.
- Break the onward connection for a fixed time, for example 30 minutes. Watch the queue grow. Reconnect, then check replay order, timestamps and completeness at the receiver.
- Replay a record or retry a delivery. An idempotent consumer must not count it twice.
- Restart each layer in turn: sensor, Gateway, broker and platform. Check that identity, configuration, clock and delivery recover.
- Compare one closed interval with the source readings, including an interval boundary, a missing value and a counter reset.
- Change an approved range or formula revision. The change must appear in the lineage. It must not change the meaning of earlier data.
- Compare important measurements with an independent reference under representative load.
The acceptance record names the equipment, configuration, reference, test time, expected result, observed result and the person who approved it.
Failure modes seen on site
| Symptom | Likely cause | How to detect it |
|---|---|---|
| Negative import power on one phase | CT reversed or on the wrong phase | Sign of each phase's power against the known load, and a power factor far from the expected value |
| Values 1,000 times too large or too small | W reported as kW, or a missing scale factor | Comparison with the meter display or a clamp meter at commissioning |
| Huge or near-zero values after a register-map change | Wrong word order or data type | A hand decode of one raw register |
| An hour missing in March and doubled in October | Intervals keyed on local time | A count of periods per local day: 92, 96 or 100 |
| A gap, then a burst of readings stamped at the same minute | Replayed data stamped on arrival | Received time minus observed time, per reading |
| Readings dated years in the past or future | A device or gateway clock not set after power loss | Observed time outside a plausible window around received time |
| A flat line at the last value | The source stopped and the display holds its last value | Stale state and the age of the newest observation |
| An energy step of several GWh in one interval | A meter reset read as a rollover | A limit on the maximum energy possible in the interval |
Monitoring data quality in operation
Commissioning proves one state on one day. Operation must detect when that state changes. For each point, track expected against received observations and the age of the newest valid observation. Count late, out-of-order, duplicate, bad, uncertain, substituted and out-of-range readings. Track clock offset, queue depth and the age of the oldest queued item. Log rejected schema, unit or identity changes, and reconcile interval totals against the cumulative register.
Alert on thresholds that match the use. For example: newest observation older than the stale limit, oldest queued item older than one hour, or clock offset above 5 s for one-minute data.
Define the denominator before you quote availability. "99.9% data availability" means nothing until the expected point set, reporting schedule, exclusions, quality rule and treatment of late data are fixed. For one point at one-minute reporting, 99.9% over 30 days still allows 43 missing readings.
ISO/IEC 25012 gives a general model for the quality of structured data. It has 15 characteristics, split between an inherent view and a system-dependent view, and ISO last reviewed and confirmed it in 2025. It supplies the vocabulary. The project still has to turn the relevant characteristics into measurable requirements.
Sensor data with EpiSensor
EpiSensor devices, such as ZEM electricity monitors and ZHT temperature and humidity sensors, report over Zigbee to a ZGW-20 Gateway. The Gateway runs EpiSensor Edge, which keeps local history (30 days by default), runs calculations and rules, and forwards data to other platforms.
Readings from EpiSensor Zigbee devices reach Edge with a timestamp on each data point, and Edge keeps that timestamp. Edge stamps a third-party Zigbee device's value with the Gateway clock when it arrives. Edge discards a data point with an invalid timestamp and logs a warning.
Edge calculates quality for each sensor over a 15-minute window. For a sensor on a fixed schedule, quality is a percentage of expected points. A sensor that reports every minute expects 15 points in the window, so 13 received gives 87%. For a sensor that reports on change, or less often than every 15 minutes, Edge reports a count of received points instead. Read the quality type before you compare two sensors.
When a supported destination fails, Edge moves the undelivered data to a queue on disk. Some integrations own their retry policy instead, so check the integration's guide. It checks the destination every minute and replays the queue when the destination returns. Replayed readings keep their original timestamps. They do not run local calculations or automation again. For MQTT, Edge counts a delivery as complete when the broker confirms the publish. A batch can arrive twice if the destination received it before Edge recorded completion, so the receiver must discard duplicates. Pending replay data has no age limit, so a long outage keeps growing the queue. Size the Gateway's free storage for the longest outage you plan to ride through.
Inside Edge, each reading carries a sensor ID, an export ID made from the device serial number and sensor ID, a timestamp and a value. The unit, scaling and quality come from the device configuration. Agree in the point list how the receiving platform gets them. Also agree which layer applies scaling, how Edge quality maps to the platform's quality model, and what alarms and control do when an input goes stale.
Once the point list is fixed, continue to IoT data-storage architecture to choose where each store sits.
Common questions
What should an IoT sensor-data record contain?
It needs the site, asset and point identity, the observed property, the value and unit, the time the observation applies to, the time it was received, and a quality state. Billing, sub-metering allocation and demand response settlement also need the scaling, calibration and calculation revision that produced the value.
What is the difference between a source timestamp and an ingestion timestamp?
A source timestamp says when the measured condition applied. An ingestion timestamp says when another system received or stored the record. The two differ during buffering, retries and clock errors. If you replace the source time with the ingestion time, old data looks current.
Does MQTT quality of service guarantee complete sensor data?
No. MQTT quality of service covers delivery between one sender and one receiver. It does not prove that the sensor measured correctly or that every sample became a message. It does not prove that the next application stored the message once, or that the payload has the correct unit and timestamp.
How should an IoT reporting interval be selected?
Start from the fastest change the application must detect and from the settlement or calculation period. Then check the permitted detection delay, device and network capacity, and storage. Sampling, calculation, publication and dashboard refresh are separate rates. Set each one on purpose.
Is a zero sensor value the same as missing or stale data?
No. Zero can be a valid measurement. Missing means no value arrived for an expected interval. Stale means the latest value is older than the freshness limit for that use. Store and show the three states separately.