Research Data and Supplementary Material for "Functional Roles of Large Language Models in Subject Indexing: A Content Analysis of the Literature"
收藏资源简介:
his repository contains the research data and supplementary materials associated with the study “Functional Roles of Large Language Models in Subject Indexing: A Content Analysis of the Literature.” The study examines how the scientific literature characterizes the participation of Large Language Models (LLMs) across the stages and operations of the subject indexing process. The analytical corpus comprises 11 documents identified through structured searches in Scopus, Web of Science, and BRAPCI and analyzed using thematic categorical Content Analysis. The deposited materials document the search and selection procedures, corpus composition, analytical categories, and final validated coding matrix. The categorical system comprises three operational categories of LLM participation — C1: Subject Analysis and Generation of Thematic Representation; C2: Evaluation and Validation of Thematic Representation; and C3: Selection and Refinement of Thematic Representation — and two cross-cutting dimensions: DT1: Semantic Regulation through Knowledge Organization Systems (KOS) Structures; and DT2: Human–LLM Mediation. The repository includes: (1) supplementary methodological documentation; (2) the consolidated research dataset in XLSX format; (3) the final validated analytical matrix in CSV format; and (4) a README describing the deposited files. Artificial Intelligence was used as analytical support during Content Analysis to identify potentially relevant textual segments and suggest codings. All suggestions were subjected to human auditing, and only the researcher-validated analytical matrix is included as research data. These materials are provided to support transparency, traceability, and reproducibility of the analytical procedures reported in the associated article.



