How to Deal With Missing Sensor Data on Ocean Drones (October 2026)

How to deal with missing sensor data on an ocean drone comes down to a fixed order: detect the gap first, work out why it happened, mark it as missing, then fill only the short stretches where an estimate is defensible — and say so in the record. Anything longer than roughly 30 to 60 seconds at typical sampling rates should be reported as a gap, not invented.

That sounds obvious until you look at a six-week deployment log. A frozen temperature probe, a GNSS receiver that stopped reporting under a jetty, and a zero-filled pressure channel can all look identical in a spreadsheet, and each one quietly poisons whatever runs downstream.

The workflow that holds up in the field is seven steps, and the order matters more than the tools:

  1. Confirm the gap is real, and separate it from a stale or merely noisy channel.
  2. Diagnose the cause from sensor health, sample timing and marine conditions.
  3. Mark the interval as missing with a reason code, never as a valid reading.
  4. Choose a fill method sized to the gap length and to the signal.
  5. Keep estimates in separate columns with an age and an uncertainty field.
  6. Fallback or degrade deliberately when a filled channel feeds a controller.
  7. Test the recovery, then watch for silence and alert on it.

Here is each one in detail.

Table of Contents

What You Need

What You Need

None of this works without a declared expected sample rate per channel. Write it down for every sensor you log, including the ones that only sample once a minute, because gap length has to be counted in samples and converted with that number.

You also need the raw stream kept untouched, separate from anything processed. Ingestion normalisation, deduplication and gap repair all belong downstream of an immutable log, so a bad fill can be rebuilt later when you learn more.

Time synchronisation has to be good enough that a gap is a real gap. GNSS discipline with a holdover, or NTP with a measured offset, gets you there. If the clock steps backwards, you have manufactured a gap and then a sudden burst of late samples.

A per-sample quality field is the cheapest tool in the whole workflow. One column carrying a value such as measured, interpolated, extrapolated, stale or invalid means nothing downstream has to guess.

Beyond that you want a second source of information that does not share a failure mode with the first, some onboard buffering so a radio blackout does not destroy data already collected, and an alert path that reaches a human when a critical channel goes quiet.

Finally, build a test harness that can delete samples on purpose. Every recovery path you write needs to be exercised before the day you need it.

Step-by-Step

1. Confirm That the Data Is Actually Missing

Missing data is a gap, a frozen value, or an unrecorded interval in a series where a reading was expected at a known cadence. Noisy and drifting data are a different problem: those samples are present but unreliable, and smoothing them does not repair a hole.

Start by separating four states that routinely get merged into one. They look alike in a chart and demand opposite treatment.

StateSignatureHow to detect itCorrect treatment
MissingNo record at all for a timestamp where one was expectedCount of samples in a window is below expected countMark the gap, fill only if short
ZeroA genuine reading that happens to be 0.0Value is present and timestamp is on cadenceKeep it. Never overwrite with zero
Stale or frozenRecords keep arriving but the value never changesVariance near zero across many samples, or the last value repeatsDiscard and raise a sensor fault
Noisy or driftingSamples arrive on time but scatter or creep away from truthVariance above expectation, or a slow trend versus a referenceFilter and recalibrate, do not treat as a gap

A zero-filled series is the most damaging case because it looks complete. Every point is on cadence, so completeness checks pass, and downstream code has no way to know it is reading a fabrication. The same applies to a sensor that has frozen at a plausible value: the chart is continuous, the numbers look reasonable, and the fault is invisible.

Delayed packets are a fourth case. If data is arriving late but within a tolerance window, it is not missing, and a reorder buffer fixes it. Malformed records, a corrupted frame or a truncated string, are missing in effect, because the measurement exists but cannot be parsed. Store the reason, because a channel that always parses badly and a channel that sometimes loses packets need completely different fixes.

One practical check separates a hardware fault from a pipeline fault. Look at the timestamps rather than the fault flags: developers working on long robotics runs in the ethz-asl maplab repository diagnose inertial dropouts from message timing gaps alone, because the device keeps reporting and only the timing shows the loss. If the sensor’s own heartbeat continued but the samples never reached your logger, the sensor is fine and the link or buffer is not.

2. Check Sensor Health and Sampling Timing

Now diagnose rather than patch. Compute the interval between consecutive samples per channel and plot the distribution; a healthy channel clusters tightly around its nominal period, and anything wider means jitter that will confuse a later fill. Then check the status fields most drivers expose, plus power events and driver messages around the gap.

Look for the physical tells. A repeated final value across many samples is a stalled or disconnected probe, and a flatline that steps once an hour is often a logger writing cached data after a link loss. An offset that appears at exactly a power-cycle boundary points at brownout recovery rather than the sensor itself.

The frozen case is worth dwelling on. A templated integration that fires only when the expected value changes reports nothing at all when a device dies, so the stale-entity check has to be built explicitly rather than assumed to arrive with the monitoring software.

On a boat, add the marine causes. A connector with salt in it produces intermittent contact that looks like packet loss. Biofouling on an optical or acoustic transducer does not kill the channel outright, it degrades it, so the symptom is a drifting signal and rising missing rate as the deployment runs on. Under a bridge or a heavy canopy, GNSS stops without any hardware fault at all.

Grouping the cause into link and power, sensor health and environment, or design and clock, keeps the fix pointed. You cannot pick a fill method until you know whether the channel is coming back.

3. Mark the Gap Instead of Treating It as a Valid Reading

Every gap needs four things in the record: a null value in the numeric column, an explicit quality flag, start and end timestamps, and a reason code from a fixed list such as link_loss, power_event, sensor_fault, clock_step or parse_error.

Null rather than zero, always. Zero is a value the sensor could legitimately have reported, and it pulls means, minima and trend lines in directions that look real. A null propagates honestly through your aggregation and tells you the coverage was partial.

Reason codes earn their keep later. When a deployment shows 40 percent loss concentrated in the acoustic channel, you want the breakdown by code to know whether you are buying a better modem or a better connector.

4. Reconstruct Only Data That Is Safe to Estimate

This is where practitioners go wrong most often, by interpolating by default and finding the artifact in a trend weeks later. The rule is simple: the longer the gap and the faster the underlying signal changes, the worse the estimate gets, and past a point any fill is fiction.

Gap length in samplesWhat it usually meansMethodWhere it breaks down
1 to 2Single dropped packetForward fill or linear interpolationFast-changing channels such as accelerometer spikes
3 to 10Brief dropout or buffering hiccupLinear interpolation, cubic spline if the signal is smoothWave-period motion and manoeuvres
11 to 60Sensor fault, power event, extended shadowingModel-based estimate or Kalman filter driven by peer channelsAnything where the model was not fitted to this condition
Above 60The channel is goneDo not fill. Report the gap and lower the confidence for the windowEverything

The threshold in the last row is the one most pages skip. It is a judgement, not a law, and it should scale with how fast your signal moves. For a slow-moving channel such as hull temperature, a gap of a few minutes is far more defensible than the same number of samples on an IMU.

Forward fill, carrying the last good value forward, is the single most common method and it is only safe on near-constant signals. On a rising or falling channel it flattens the trace and quietly erases the slope. Linear interpolation assumes the truth moved in a straight line between the two endpoints, which is a decent assumption for a thermistor over a few seconds and a bad one through a wave impact.

Filtering is not a fill. A moving median or exponential smoothing hides noise but produces no value for a missing sample, and applying it across a gap smooths across the boundary and hides the hole you were trying to expose.

A short Python pass covering detect, fill and flag:

import numpy as np, pandas as pd

MAX_FILL_SAMPLES = 10   # above this, leave the gap open

df = pd.read_parquet("flight_014.parquet").sort_values("t")
period = 1.0                     # seconds, declared expected cadence
grid = pd.date_range(df["t"].min(), df["t"].max(),
                     freq=pd.Timedelta(seconds=period))

df = df.set_index("t").reindex(grid)          # absent samples become NaN
was_missing = df[["temp_c", "depth_m"]].isna()

df["temp_c"] = df["temp_c"].interpolate(limit=MAX_FILL_SAMPLES, limit_area="inside")
df["depth_m"] = df["depth_m"].ffill(limit=2)

df["temp_quality"] = np.where(~was_missing["temp_c"], "measured",
                       np.where(df["temp_c"].isna(), "missing", "interpolated"))
df["depth_quality"] = np.where(~was_missing["depth_m"], "measured",
                       np.where(df["depth_m"].isna(), "missing", "forward_filled"))

completeness = (df["temp_quality"] != "missing").mean() * 100
print(f"completeness {completeness:.1f}%")

Two details in that snippet carry more weight than they look. Reindexing onto an explicit time grid is what turns an absent sample into a NaN instead of silently shifting every later timestamp, and limit_area=”inside” stops interpolation from inventing data beyond the end of the flight. The quality columns come from the mask captured before filling, so nothing can be filled without being marked.

5. Preserve the Uncertainty and the Original Record

Keep the raw values and write estimates into separate columns. If you overwrite the measurement in place, you destroy the evidence, and you will not be able to tell later whether an odd turn in a track was real or an artifact of a fill.

Attach two more fields to anything you estimate: the age of the estimate, meaning how far it was carried from real data, and an uncertainty value. A forward-filled point one sample old and one two hundred samples old are both labelled interpolated, which hides a big difference in trust.

Make every correction traceable. Store which method produced a value, the version of the code that ran, and the reason code inherited from the gap. When a model gets retrained in six months, you want to know which gaps were closed with the old model.

Quantify the error rather than assuming it. The cheapest estimate is to withhold a run of real samples of the length you actually fill, run the method, and measure the mean absolute error against the truth you kept. Do that across several gap lengths and several conditions, and you get a number you can act on instead of a feeling.

6. Prevent One Sensor Failure From Corrupting Decisions

A filled channel that feeds a controller is a liability. Decide in advance what the system does when confidence drops below a threshold: fall back to a second sensor, enter a degraded mode with slower sampling and reduced speed, surface on a preplanned heading, or hold station and wait for the link to recover.

Redundancy is about complementarity, not duplication. Two identical IMUs share every failure mode, including the same firmware bug and the same power rail. The useful pairing is a fast, responsive sensor with a slow, accurate one, which is the arrangement described in the r/AskRobotics discussion of this problem: IMU fast but drifting, GNSS slow but accurate, outages expected, and a system that degrades gracefully rather than failing.

Sensor fusion handles medium gaps better than interpolation because it brings in a channel that is still alive. A Kalman filter or an equivalent state-space estimate uses the IMU for the short term and the GNSS fix to correct the drift, so when the fix drops out the state estimate keeps running with growing covariance instead of pretending the number is exact.

Tie the confidence value to behaviour, not just to a chart. If navigation depends on position and the position confidence rises above a threshold, the drone slows down, stops accepting new mission legs and logs an escalation. The alert should name the channel, the gap length, the last good timestamp and the reason code, because an alert that says sensor error gets ignored and one that says depth_m missing, 412 samples, link_loss, last good 14:02:10 gets acted on.

7. Test the Recovery Workflow and Prevent Recurrence

Simulate gaps on real logs before you trust the recovery. Delete a random 2 percent of samples, cut a 30-minute block, freeze one channel at its last value, and shift the clock forward. Run the pipeline and check that every gap is flagged, every flag matches the injected fault, and no fill crossed a region you told it not to fill.

Then review real incidents. Every field gap should produce a short note: what the signature was, what the cause turned out to be, how long detection took and what the system did in the meantime. This is the artefact that stops the same failure recurring, and it is the part most teams skip.

On the prevention side, add a watchdog that alerts when no sample arrives within two or three expected periods, not when a value goes out of range. Silence is the signal. Improve time synchronisation so gaps are real rather than artefacts, size the onboard buffer to cover a realistic outage, schedule connector inspection and transducer cleaning into the deployment cycle, and keep a reference sensor aboard when calibration drift is the dominant failure mode.

Common Mistakes

Zero-filling gaps. The series looks complete, so every completeness check passes and the fabricated points are indistinguishable from measurements. Use nulls.

Interpolating silently. Filled points land in the same column as measured points, with no marker, and weeks later someone reads them as truth. A quality flag is not polish, it is the record.

Ignoring stale values. A channel that keeps reporting a frozen number is failing quietly, and interpolation over a flatline produces a flat fill that looks like calm conditions. Detect near-zero variance first.

Using one method for every channel. A fill tuned for hull temperature will ruin an IMU trace. Match the method to how fast the signal actually moves.

Overwriting the raw record. Corrections applied in place cannot be audited or reversed. Keep raw and estimated columns apart.

Filling long gaps to keep a pipeline happy. Downstream code that breaks on nulls will push you to fabricate, so give it a completeness tolerance and a degraded mode instead.

Treating noise as missing data. These are separate problems. Filters reduce noise, gaps need provenance, and a smoothing pass across a gap hides the hole rather than closing it.

One habit ties them together: practitioners distrust anything that hides a gap. Surface it.

Frequently Asked Questions

Should missing sensor data be filled with zero?

No. Zero is a value the sensor could genuinely have reported, so a zero-filled series looks complete and every completeness check passes on it. Means, minima and trend lines all shift toward it, and nothing downstream can tell the fabrication from a measurement. Use a null value plus an explicit quality flag and start and end timestamps, then fill only the short gaps with a method chosen for the signal.

How long can a sensor gap be before interpolation becomes misleading?

It depends on how fast the signal moves, so the threshold has to be stated in samples and converted with your sample rate. A common working rule: forward fill for one or two samples, linear or spline interpolation for three to ten, a model-based or fusion estimate for tens of samples, and no fill at all beyond roughly 60 samples. For a slow channel like hull temperature you can defend far more than for an IMU.

Is linear interpolation safe for GPS, temperature, and pressure data?

It is reasonable for a slow, smooth channel over a few samples, such as temperature or barometric pressure, because the truth stays close to the straight line between endpoints. It is unsafe for GPS during a manoeuvre, for anything wave-driven, and for any channel that can spike. Interpolation also hides the fact that the sensor was absent, so always mark the filled points as imputed with an age field.

What is the difference between missing data and a stale sensor value?

Missing data means no record arrived where one was expected, so the gap shows up as absent timestamps. A stale value means records keep arriving on schedule with the same number repeated, so the gap is invisible in a completeness count. Detect stale channels by variance across consecutive samples or by the last value repeating, and treat them as a sensor fault rather than something to fill.

Should raw sensor readings be changed when a gap is repaired?

No. Keep the raw stream immutable and write estimates into separate columns carrying a quality flag, the method used, the reason code inherited from the gap and an uncertainty value. If you overwrite in place you lose the evidence, you cannot audit the repair later, and a retrained model cannot be compared against the version that produced the original trend.

How can an ocean drone remain safe when a critical sensor fails?

Decide the degraded behaviour before deployment, not during the incident. Drop the affected estimate below a confidence threshold, fall back to a complementary sensor rather than a duplicate, and enter a defined mode such as reduced speed, holding a preplanned heading, or surfacing. Pair the alert with channel name, gap length, last good timestamp and reason code so the operator can act on it.

If you take one thing from how to deal with missing sensor data on a drone, make it this: declare your expected cadence per channel and compute completeness against it before the next deployment leaves the dock. Every other choice in this workflow, from gap-length thresholds to alert thresholds, follows from that single number, and it costs an afternoon to add rather than a season of arguing about which trend was real.

Leave a Comment