A live walkthrough of a cleaning pipeline flagging five real-world problem types in solar telemetry, one stage at a time.
Note: Real raw (uncleaned) solar SCADA datasets are either registration-gated (DKASC) or
already cleaned before release (Mendeley PV SCADA). To demonstrate this, I generated a
synthetic dataset built to reproduce documented real-world failure patterns from the
literature — but the cleaning logic and model below come from real, running code.
Step 1 / 6
raw datacleaned datamissing (gap)suspected CT-ratiosensor spike
What the five flag types mean
duplicate
Same record sent twice after a reconnect — the second copy is dropped.
missing (gap)
Connectivity outage. A short gap (≤30 min) is interpolated; a long gap is flagged "missing", never fabricated.
suspected CT-ratio
This inverter's value deviates sustainably by ~2×/4×/10× from the fleet median — a likely current-transformer calibration issue. Not auto-corrected; needs a field check.
sensor spike
A single-point spike that reverts immediately — sensor noise, not real production. Kept distinct from real clipping (a sustained flat-top at rated capacity).
model anomaly
Output is notably below what irradiance and temperature predict — a candidate for underperformance (shading, soiling, degradation).