Bibliometric Accuracy of Generative AI–Recommended Target Journals: A Comparative Study of Four Large Language Models Across Ten Manuscripts
收藏资源简介:
Background: Researchers increasingly consult generative AI chatbots to select target journals, but the accuracy of the bibliometric indicators reported (impact factor, quartile, publisher) has not been systematically quantified. Objective: To quantify and compare bibliometric error rates across four generative AI tools recommending journals for real manuscripts. Methods: Ten manuscripts spanning diverse disciplines were submitted to ChatGPT, Gemini, Claude, and Kimi, each proposing 10 candidate journals with metrics, yielding 400 recommendations and 192 unique journals verified independently by two blinded reviewers. Error rates were compared using Wilson confidence intervals, chi-square tests, ANOVA, and clustering-adjusted models (GEE, mixed-effects) accounting for manuscript-level nesting. Results: No fabricated journals were found (0/400); 1.75% named a real journal under an outdated title. At least one metric error occurred in 71.5% of recommendations (64.0%-79.0% by tool; GEE P = .001). Impact-factor error was comparable across tools (P = .80), but quartile error differed significantly (22.1%-57.3%; P < .001). Inter-rater agreement was high (82.7%-99.5%), and results replicated with a second reviewer. No manuscript received unanimous four-tool agreement on its top-ranked journal. Conclusions: Generative AI tools rarely fabricate journals but frequently misreport their bibliometric metrics, with error profiles differing by tool. AI-generated journal recommendations should be treated as an exploratory aid requiring systematic verification.



