FICSIM
收藏资源简介:
FICSIM是一个为长篇小说的多方面语义相似性评估而构建的数据集。该数据集由长篇小说、最近创作的小说组成,包括12个维度的相似度评分,这些评分由作者产生的元数据和数字人文学者验证。数据集来源于Archive of Our Own(AO3),一个拥有超过1500万作品的数字存档。为了保证数据质量,作者们获得了每位作者的同意,以确保他们的作品能够被用于研究和分析。FICSIM旨在解决数字人文领域中,特别是在计算文学研究任务中,评估语言模型在处理长篇文本方面的能力问题。
FICSIM is a dataset constructed for multi-faceted semantic similarity assessment of full-length novels. The dataset comprises newly created full-length novels, with similarity scores across 12 dimensions validated by author-generated metadata and digital humanities scholars. The dataset is sourced from Archive of Our Own (AO3), a digital archive housing over 15 million works. To ensure data quality, the dataset developers have obtained informed consent from each original author to utilize their works for research and analysis. FICSIM aims to address the gap in evaluating large language models' capabilities in processing long-form texts within the digital humanities, particularly for computational literary research tasks.

- 1FicSim: A Dataset for Multi-Faceted Semantic Similarity in Long-Form Fiction卡内基梅隆大学语言技术研究所,俄克拉荷马大学图书馆与信息研究学院 · 2025年



