Dataset of bibliometric records and topic modelling outputs for AI-driven boiler optimisation in thermal power plants (2014–2025)
收藏资源简介:
This dataset supports the study of artificial intelligence (AI)-driven boiler optimisation in thermal power plants through bibliometric analysis and topic modelling. It comprises bibliographic records retrieved from the Scopus database for English-language journal articles and review papers published between 2014 and 2025. The dataset was developed to examine the evolution of research on AI applications for boiler performance optimisation, combustion control, emissions reduction, predictive maintenance, fault diagnosis, and intelligent monitoring in thermal power generation. It includes raw bibliometric records, processed text data, document–term matrices, Latent Dirichlet Allocation (LDA) outputs, temporal topic trends, and supporting bibliometric statistics. The LDA model identifies seven latent research topics from document titles, abstracts, and author keywords. The dataset contains document–topic probability distributions (theta.csv), topic–term probability distributions (beta.csv), dominant topic assignments, top-ranked topic terms, topic labels, and annual topic prevalence. These files enable users to investigate thematic structures, analyse the evolution of research topics over time, and compare topic distributions across publications. The dataset also includes processed text files (cleaned_corpus.csv and dtm.csv) to facilitate replication of the topic modelling workflow. Supporting files such as document_topic_full.csv and country_stats.csv provide integrated bibliometric metadata and publication statistics for further analysis. The data can be interpreted at multiple levels. Bibliometric records support analyses of publication trends, research productivity, collaboration patterns, and institutional or country contributions. The LDA outputs provide probabilistic representations of document topics, where higher topic probabilities indicate stronger thematic relevance. Topic–term probabilities identify the most representative terms within each topic, while annual topic prevalence enables assessment of changes in research emphasis over time. Researchers may use this dataset to reproduce the published analysis, evaluate alternative topic modelling approaches, benchmark text mining methods, conduct scientometric studies, or investigate emerging trends in AI applications for thermal power plants and energy systems. The dataset is compatible with R, Python, MATLAB, and other software environments that support CSV-formatted data.




