Bibliometric Dataset: Defense Mechanisms in Large Language Models, 2024–2026
收藏资源简介:
Reproducibility package for the bibliometric mapping of defense mechanisms against adversarial attacks on Large Language Models (LLMs), accompanying the manuscript "The Evolution of Defense Mechanisms in Large Language Models: Emerging Trends and Strategic Mapping of the 2024–2026 Research Landscape" submitted to Scientometrics. This deposit contains: Raw corpus: a Scopus BibTeX export of 137 Q1 journal articles published 2024–2026, extracted on 13 February 2026 from 17 explicitly white-listed Q1 sources (SCImago Journal Rank 2024). 2. Search queries: the verbatim Boolean queries used in Scopus, including the main corpus query and the methodological precedent query. 3. Analysis code: a fully reproducible R script using Bibliometrix 4.1 that regenerates every quantitative claim in the manuscript — main bibliometric indicators, annual scientific production, Bradford's Law, Lotka's Law, country productivity and collaboration networks, author co-authorship networks, most cited documents, Callon's strategic centrality–density diagram, thematic evolution analysis, Multiple Correspondence Analysis (MCA), hierarchical clustering, and keyword co-occurrence networks. 4. Pre-computed results: 13 Excel tables exported directly from Bibliometrix on 13 February 2026 (the actual files that produced the reported values), plus the country-level production CSV and the preliminary knowledge-gap analysis generated with Scopus AI Deep Research. 5. Visualizations: 23 figures used in or supporting the manuscript, including VOSviewer overlay, density, and network visualizations. 6. Documentation: a comprehensive README, data dictionary, search strategy document, citation file (CITATION.cff), and CC BY 4.0 license. Key reproducible findings: 137 documents from 19 unique Q1 sources, 78.38% annual growth rate, 516 authors, China–Singapore as the most active bilateral collaboration dyad (15 links), and "adversarial machine learning" as the highest-frequency motor theme (n = 23 keyword occurrences). Multiple Correspondence Analysis explains 53.40% of conceptual variance across two dimensions. Software requirements: R ≥ 4.3, Bibliometrix 4.1.4, igraph 2.0+, VOSviewer 1.6.20. Suggested citation: Vargas Yáñez, S., & Tobón, S. (2026). Bibliometric Dataset: Defense Mechanisms in Large Language Models, 2024–2026 [Dataset]. Zenodo.



