MULDER: A Multilingual Dataset for Term-Definition Extraction in Scientific Literature
收藏资源简介:
Finding relevant papers from the rapidly growing scientific literature is increasingly overwhelming. Recent advances in Information Retrieval (IR) and Large Language Models (LLMs) have substantially improved semantic scientific search by enabling retrieval systems to identify semantically related studies. However, existing systems still primarily rely on lexical overlap or embedding similarity, often retrieving papers that are semantically similar to a query document rather than studies that explicitly share the same concepts, methods, or research tasks. As a result, researchers seeking more fine-grained and concept-specific search results may still struggle to find truly relevant work. To resolve this, we investigate a two-stage semantic retrieval pipeline in which users can iteratively refine retrieved papers using explicit conceptual information (i.e., term-definition pairs) extracted from scientific publications. To support this setting, we introduce MULDER (MULtilingual Definition Extraction for Retrieval), a dataset of 344 machine learning abstracts (172 English and 172 Turkish) with human-annotated term-definition pairs in English and Turkish, together with cross-lingual context links that align equivalent concepts across languages. Additionally, we provide a retrieval-oriented augmentation benchmark containing 777 automatically retrieved ArXiv abstracts annotated for term-definition extraction. Finally, we present benchmark experiments and a reproducible evaluation framework demonstrating how these datasets can support future research on concept-aware scientific retrieval.



