Journal of Siberian Federal University. Engineering & Technologies / Automatic Classification of Power Outage Causes in 0.4 kV Networks Based on Text Records Using Machine Learning

Full text (.pdf)
Issue
Journal of Siberian Federal University. Engineering & Technologies. 2026 19 (5)
Authors
Bolshev, Vadim E.; Vinogradova, Alina V.; Volynets, Igor O.
Contact information
Bolshev, Vadim E.: Federal Scientific Agroengineering Center VIM (Moscow, Russian Federation); ; ORCID: 0000-0002-5787-8581; Vinogradova, Alina V.: Federal Scientific Agroengineering Center VIM (Moscow, Russian Federation); Volynets, Igor O. : Usetech Bel LLC (Minsk, Belarus)
Keywords
automatic cause identification of power outages; 0.4 kV electrical networks; emergency and planned outage log; machine learning; multi-class classification problem; vector representation of text documents; text classification
Abstract

This work presents a study dedicated to the development and evaluation of machine learning algorithms for automating the identification of the causes of power outages in 0.4 kV electrical networks. The urgency of this task is due to inaccuracies in maintaining reporting documentation, which are caused by non-standardized wording and errors in the log of emergency and planned outages, thus hindering subsequent analysis and the adoption of measures to improve reliability. The proposed approach is based on the analysis of textual information concerning the progress of fault elimination and its subsequent comprehensive preprocessing, including normalization, lemmatization, and vectorization using the TF-IDF weighting method, which allows the text descriptions to be converted into a «term-document» matrix. Since identifying the causes of power outages is a multiclass classification problem, three algorithms were tested: logistic regression, random forest, and gradient boosting (LightGBM). Model quality was evaluated using cross-validation with stratified splitting and on an independent test sample. The best results were demonstrated by the logistic regression model, which achieved an Accuracy of 0.864 and an F1-score with micro-averaging of 0.866 on the test set, confirming the high practical applicability of the proposed approach. Error analysis showed that the main barriers to improving accuracy are spelling errors in the source data and the ambiguity of the classification system itself, where a single event can be assigned to multiple categories. Based on the results obtained, recommendations for further improving classification quality were formulated. These include both data enhancement (increasing the sample size, error correction at the input stage, improving the cause classifier) and the application of more complex approaches, such as automatic text correction, the use of embeddings to account for semantics, and the application of deep neural networks

Pages
684–695
EDN
BNQMFJ
Paper at repository of SibFU
https://elib.sfu-kras.ru/handle/2311/159307

Creative Commons License This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0).