A Curated Dataset of Research Abstracts on AI and Large Language Models in Information Management and Librarianship (2020-2024)
收藏资源简介:
This dataset contains 303 curated abstracts of research articles focusing on the application of Artificial Intelligence (AI) and Large Language Models (LLMs) in information management, library science, and related services. The data was systematically collected from the Scopus database for publications between 2020 and 2024. The original search result yielded 498 documents. This was refined to 307 records upon exporting abstract data. A final set of 303 high-quality abstracts was established after a rigorous preprocessing and cleaning pipeline to ensure data integrity and suitability for natural language processing (NLP) and semantic analysis tasks. This collection is ideal for researchers working in: Text Mining and Semantic Clustering Topic Modeling and Research Trend Analysis Natural Language Processing (NLP) Applications Bibliometric and Scientometric Studies AI in Libraries and Information Services Keywords: Artificial Intelligence; Large Language Models; Natural Language Processing; Information Management; Librarianship; Research Abstracts; Text Dataset; Scopus. 2. For the Cluster Assignments File File Name: abstracts_AI_LLM_librarianship_cluster_assignments.csv Title: Semantic Cluster Assignments for a Corpus of AI/LLM in Librarianship Research Abstracts Description: This file provides the semantic cluster assignments for the corresponding dataset "A Curated Dataset of Research Abstracts on AI and Large Language Models in Information Management and Librarianship (2020-2024)". The clusters were generated by applying K-means clustering to average GloVe word embeddings of the article abstracts. The analysis identified 7 distinct, semantically coherent research themes, which are detailed in the associated research article. This data is provided to facilitate reproducibility, further analysis, and to serve as a ground truth for comparative studies in semantic clustering and topic discovery. The file includes the document identifier and its corresponding cluster label (0-6). The dominant themes for each cluster are: Cluster 0: Healthcare Information Systems Cluster 1: Information Retrieval Systems Cluster 2: Research Data Management (RDM) Cluster 3: Digital Library Adoption Cluster 4: Knowledge Management Cluster 5: Reference Services Cluster 6: Scientific Publishing Keywords: Semantic Clustering; K-means; Cluster Labels; Research Themes; Topic Discovery; GloVe Embeddings; Text Mining.
本数据集收录了303篇经精心甄选的研究论文摘要,主题围绕人工智能(Artificial Intelligence, AI)与大语言模型(Large Language Models, LLMs)在信息管理、图书馆学及相关服务领域的应用。所有数据均系统采集自Scopus数据库2020年至2024年间发表的文献。 最初的检索结果共得到498篇文献,导出摘要数据后筛选至307条记录。随后经过严格的预处理与清洗流程,最终得到303篇高质量摘要,以保障数据完整性与适配自然语言处理(Natural Language Processing, NLP)及语义分析任务的要求。 本数据集适用于以下研究方向的科研人员: - 文本挖掘与语义聚类 - 主题建模与研究趋势分析 - 自然语言处理(NLP)应用 - 文献计量与科学计量研究 - 图书馆与信息服务中的人工智能应用 关键词:人工智能;大语言模型;自然语言处理;信息管理;图书馆学;研究论文摘要;文本数据集;Scopus。 2. 聚类分配文件 文件名:abstracts_AI_LLM_librarianship_cluster_assignments.csv 标题:面向图书馆学领域人工智能/大语言模型研究摘要语料库的语义聚类分配 描述:本文件为对应数据集《2020-2024年信息管理与图书馆学领域人工智能及大语言模型研究摘要精选数据集》提供语义聚类分配结果。聚类通过对论文摘要的平均GloVe词嵌入应用K-means聚类算法生成。 本次分析共识别出7个语义连贯的独立研究主题,相关细节已在配套研究论文中详述。本数据的发布旨在助力研究可复现性、后续分析,并可作为语义聚类与主题发现领域对比研究的基准真值。 本文件包含文献标识符及其对应的聚类标签(0至6)。各聚类的主导主题如下: 聚类0:医疗信息系统 聚类1:信息检索系统 聚类2:研究数据管理(Research Data Management, RDM) 聚类3:数字图书馆应用 聚类4:知识管理 聚类5:参考咨询服务 聚类6:科学出版 关键词:语义聚类;K-means;聚类标签;研究主题;主题发现;GloVe词嵌入;文本挖掘。




