anon-muses-neurips/muses
收藏资源简介:
MUSES(挖掘未探索的科学证据以激发新假设生成)是首个用于前瞻性智力根源预测的百万实例基准。给定作者在时间t的已记录发表历史,任务是根据每篇科学论文进入作者下一篇论文参考文献的可能性,对一个固定的包含233万篇科学论文的池进行排名。该基准在两个方面具有挑战性:熟悉度(CiteNext → CiteNew → CiteNew-Isolated)和功能性(任何引用 → 修辞性根源证据 → 作者认可)。数据集包含实例分割、不同层级的正样本集、候选池及其派生标志,但不包括S2ORC文本,需从上游获取。
MUSES (Mining Unexplored Scientific Evidence to Spark novel hypothesis generation) is the first million-instance benchmark for prospective intellectual-roots prediction. Given an authors documented publication history at time t, the task is to rank a fixed pool of 2.33M scientific papers by how likely each one is to enter the authors next papers bibliography. The benchmark is hard along two orthogonal axes: Familiarity (CiteNext → CiteNew → CiteNew-Isolated) and Functional (any citation → rhetorical ROOT evidence → author endorsement). The dataset includes instance splits, tier-specific positive sets, candidate pool and derived flags, but does not include S2ORC text, which must be obtained from the upstream release.



