CORTEXA
← Browse
arxivcs.LG2026-07-15

Clustering algorithms for multivariate wind farm SCADA data filtering

Nicolò Italiano, Vasilis Pettas, Tuhfe Göçmen, Nicolaos A. Cutululis

During wind farm operation, Supervisory Control and Data Acquisition (SCADA) systems record numerous anomalies, transients, and specific operational modes, leading to large datasets. However, for a wide range of applications, only measurements corresponding to normal operation are required and, therefore, the SCADA data must be filtered. For this purpose, several methods have been proposed to automate and replace manual filtering conducted by experts via visual inspection of the data. In this paper, we compare the filtering accuracy of multiple clustering algorithms against manual filtering, introducing evaluation metrics that are suitable for unlabeled data and robust across potential applications. Based on the results, we provide recommendations for generalizing model calibration to different datasets and discuss potential use cases for each model. The models are applied to the SCADA data of three turbines of an existing offshore wind farm, using 10-minute statistics across multiple data channels. In addition to the anomalies and operational modes typically recorded, the dataset presents a large number of non-evident outliers due to several field tests. Overall, the results highlight the importance of extending the analysis beyond the power curve, both in feature selection and in the design of evaluation metrics. In most cases, cluster-based methods are able to detect both evident and subtle outliers, achieving higher accuracy than manual filtering. However, the accuracy and the amount of data retained vary considerably depending on the model, and expert involvement remains necessary, though to a reduced extent compared to manual filtering.

View free PDFSource page

Related papers

arxivcs.LG2026-07-23

CASC: Causal Adversarial Subspace Clustering for Multivariate Spatiotemporal Data

Francis Ndikum Nji, Vandana Janeja, Jianwu Wang

Deep subspace clustering plays a critical role in applications involving multivariate spatiotemporal data, such as sea ice monitoring, disease spread analysis, and tracking neuro-degeneration over time. Despite recent advances, existing methods primarily rely on geometric self-ex…

View free PDFSource page
arxivcs.LG2026-06-29

Toward an Energy-Optimized Operation of Data Centers Located in Wind Farms Using Reinforcement Learning

Jan Stenner, Alexander Kilian, Sebastian Peitz, Hermann de Meer

This paper studies Reinforcement Learning as an online controller for curtailment-aware workload shifting in wind-turbine-integrated high-performance computing (HPC) data centers. We introduce a reproducible fixed-day simulation framework with synthetic wind and price signals and…

View free PDFSource page
arxivcs.LG2026-07-17

Data-Native Global Optimization for Big Data K-means Clustering

Ravil Mussabayev, Rustam Mussabayev, Zukhra Yerdaliyeva, Kuldeyev Nursultan

Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids. We target arbitrarily tall data: a fixed feature space ma…

View free PDFSource page