adugeen/personal-facts-msc
收藏资源简介:
Personal Facts (MSC) — Multi-Dimensional Annotation是一个手动标注的数据集,包含2,779个从Multi-Session Chat (MSC)语料库中抽取的个人事实。这些事实标注了七个维度,包括主题、时间锚定、指称、生命周期、有效性以及对话延续潜力。数据集分为训练集(2,223个样本)和测试集(556个样本),并提供了每个字段的详细描述和标签清单。数据集的构建过程包括从MSC语料库中抽取、去重、嵌入和聚类,然后进行手动标注。数据集可用于训练和评估多标签分类器、质量过滤、记忆策略研究以及审计更大规模的人设语料库。数据集的局限性包括单一标注者、类别不平衡、源偏差以及标注噪声。
Personal Facts (MSC) — Multi-Dimensional Annotation is a manually annotated dataset of 2,779 personal facts sampled from the Multi-Session Chat (MSC) corpus, labeled across seven dimensions that jointly characterize a facts topic, temporal anchoring, referent, lifetime, validity, and dialogue-continuation potential. The dataset is split into a training set (2,223 examples) and a test set (556 examples), with detailed descriptions of each field and label inventory provided. The datasets construction involves sampling from the MSC corpus, deduplication, embedding, clustering, and manual annotation. It is intended for training and benchmarking multi-label classifiers, quality filtering of persona corpora, memory-policy research, and auditing larger persona corpora. Limitations include single annotator, class imbalance, source bias, and noisy annotations.





