Data and code for "Mapping the Convergence of Information Retrieval and Large Language Models"
收藏资源简介:
This deposit contains the data and code supporting the bibliometric study "Mapping the Convergence of Information Retrieval and Large Language Models." The study maps the intellectual structure and temporal evolution of the information retrieval (IR) and large language model (LLM) literatures using bibliographic coupling analysis. Bibliographic records were retrieved from OpenAlex (https://openalex.org) for the period 2015–2025 using two query arms (eight bridge-term abstract searches and sixteen IR-title × LLM-abstract combinations), yielding an analytic corpus of 15,024 records. Coupling networks were constructed in R using the bibliometrix package, thresholded at a minimum of five shared references, restricted to the giant connected component, and cleaned of off-topic clusters arising from term polysemy. The final network of 4,064 documents and 71,534 coupling links was analysed in Gephi (Louvain community detection, ForceAtlas2 layout, betweenness and eigenvector centrality), and four temporal slices were analysed to trace the field's evolution. Contents: the R scripts for data retrieval and network construction; the final coupling network and four period networks (GraphML); the node-level community and centrality table (CSV); a provenance file documenting record counts at each collection stage; and a README describing each file and the pipeline order. Bibliographic records derive from OpenAlex and are released under its CC0 terms; all derived files and code in this deposit are made available for reproducibility.



