遇见数据集

Polish literary studies citation data: source coverage and OAI-PMH-to-OpenCitations case studies

收藏
Zenodo2026-08-05 更新2026-08-13 收录
官方服务:

资源简介:

This dataset documents research on how citation infrastructures shape the observable representation of Polish literary studies and on how publisher-supplied reference lists can be transformed into interoperable citation data. It was produced within the National Science Centre, Poland project Analysis of trends in Polish literary studies using digital methods (grant 2025/09/X/HS2/00585). The dataset has three connected components. First, it documents a comparison of coverage for 82 Polish literary-studies journals in OpenCitations and Scopus. The source snapshots contained 8,421 article records in OpenCitations and 17,833 in Scopus. Record linkage produced a union of 21,017 articles: 5,237 found in both sources, 3,184 only in OpenCitations and 12,596 only in Scopus. The public release contains OpenCitations data and journal-level derived comparison results; the record-level Scopus export is excluded because third-party redistribution rights are not established. Second, the dataset contains two journal case studies comparing structured references supplied through OAI-PMH with retrospective extraction from PDF files using the OpenCitations Citation Extraction Service and GROBID. The Forum Poetyki workflow processed 339 eligible articles and 7,727 OAI-PMH reference instances. Its stratified sample comprised 50 PDFs; 32 produced at least one bibliographic candidate, yielding 903 candidates against 1,089 publisher-supplied references in the same sample. The Zagadnienia Rodzajów Literackich workflow processed 391 eligible records, of which 234 contained at least one OAI-PMH reference, with 7,707 reference instances in total. In its 50-document sample, 44 PDFs produced 1,211 candidates; the corresponding sample contained 1,082 OAI-PMH references. Third, publisher-supplied references were parsed with AnyStyle, enriched with Crossref candidates, deduplicated, normalised and exported as OpenCitations-compatible CSV files. The two workflows produced 11,857 metadata rows and 12,305 citation links in total. Both exports passed closure and structural validation with zero errors. The validator returned 12 non-blocking warnings: 11 for uppercase titles and one for a page interval. The deposited files include harvesting inventories, deterministic sample manifests and allocation tables, PDF-extraction outputs and failures, publisher-supplied reference lists for the samples, parsed and Crossref-enriched references, candidate-level matching diagnostics, deduplication mappings, normalisation logs, export exclusions, final metadata and citation tables, summary files and validator reports. Full-text PDFs, service caches, credentials and production API keys are excluded. The dataset supports research on source-dependent coverage, missing bibliographic metadata, retrospective citation extraction and reproducible transformation of humanities reference lists. It is a pair of journal case studies rather than a representative sample of all Polish literary-studies publishing platforms, document types or periods.

提供机构:
Zenodo
创建时间:
2026-08-05
二维码
社区交流群
二维码
科研交流群
商业服务