arXiv AI

A Discrete-sequence Dataset for Evaluating Online Unsupervised Anomaly Detection Approaches for Multivariate Time Series

By Advanced AI EditorApril 10, 2025No Comments2 Mins Read

[Submitted on 21 Nov 2024 (v1), last revised 8 Apr 2025 (this version, v4)]

View a PDF of the paper titled PATH: A Discrete-sequence Dataset for Evaluating Online Unsupervised Anomaly Detection Approaches for Multivariate Time Series, by Lucas Correia and 3 other authors

View PDF
HTML (experimental)

Abstract:Benchmarking anomaly detection approaches for multivariate time series is a challenging task due to a lack of high-quality datasets. Current publicly available datasets are too small, not diverse and feature trivial anomalies, which hinders measurable progress in this research area. We propose a solution: a diverse, extensive, and non-trivial dataset generated via state-of-the-art simulation tools that reflects realistic behaviour of an automotive powertrain, including its multivariate, dynamic and variable-state properties. Additionally, our dataset represents a discrete-sequence problem, which remains unaddressed by previously-proposed solutions in literature. To cater for both unsupervised and semi-supervised anomaly detection settings, as well as time series generation and forecasting, we make different versions of the dataset available, where training and test subsets are offered in contaminated and clean versions, depending on the task. We also provide baseline results from a selection of approaches based on deterministic and variational autoencoders, as well as a non-parametric approach. As expected, the baseline experimentation shows that the approaches trained on the semi-supervised version of the dataset outperform their unsupervised counterparts, highlighting a need for approaches more robust to contaminated training data. Furthermore, results show that the threshold used can have a large influence on detection performance, hence more work needs to be invested in methods to find a suitable threshold without the need for labelled data.

Submission history

From: Lucas Correia [view email]
[v1]
Thu, 21 Nov 2024 09:03:12 UTC (3,409 KB)
[v2]
Mon, 25 Nov 2024 14:24:57 UTC (3,409 KB)
[v3]
Wed, 15 Jan 2025 17:16:22 UTC (3,409 KB)
[v4]
Tue, 8 Apr 2025 15:26:49 UTC (3,437 KB)

Previous ArticleStanford’s 2025 AI Index Reveals an Industry at a Crossroads

Next Article Google DeepMind imposes tough non-competes on UK staff as AI talent wars intensify | Technology News

Advanced AI Editor

Leave A Reply