遇见数据集

PolyglotFakeFacts-ITPS: Reproducibility Package for Multilingual Information Threat Prioritization

收藏
Mendeley Data2026-09-09 收录
官方服务:

资源简介:

This repository provides the reproducibility package associated with the experimental evaluation of the Information Threat Priority Score (ITPS), an interpretable framework for multilingual information-threat prioritization. The package is based on a controlled subset of the PolyglotFakeFacts V2 dataset and contains 2,000 news articles (1,000 Fake and 1,000 Real) with identical language distributions across classes. The subset was constructed using a deterministic language-balanced sampling procedure with random seed 20260819. The experimental ITPS evaluated in this study combines three components: leakage-controlled authenticity risk (A), embedding-derived security relevance (S), and cross-lingual narrative co-presence (P). Security relevance and cross-lingual semantic similarity were computed using the sentence-transformers/paraphrase-multilingual-mpnet-base-v2 multilingual embedding model. The package includes the experimental subset, the selection and provenance manifest, the executed experiment notebook, leakage and robustness audits, precomputed out-of-fold authenticity probabilities, final article-level A/S/P/ITPS results, ranking-reversal analyses, sensitivity analyses, and documentation of the experimental environment and methodology. The authenticity component was evaluated using source/domain-disjoint and near-duplicate-aware five-fold cross-validation with boilerplate/source-fingerprint stripping. The final leakage-controlled authenticity probabilities used by the ITPS experiment achieved ROC-AUC = 0.943520. This package is intended to support transparency, verification, and reproducibility of the reported ITPS experiment. It is a derived experimental research artifact and does not replace or constitute a new version of the parent PolyglotFakeFacts dataset. Parent dataset: Ciobanu, Alexandru (2026), “PolyglotFakeFacts: A Multilingual Dataset of Fake and Real News across Politics, Security, and Social Domains,” Mendeley Data, V2, DOI: 10.17632/gff8bmr4ff.2.

本仓库提供了与信息威胁优先级评分(Information Threat Priority Score,简称ITPS)实验评估相关的可复现性套件,该评分是一种面向多语言信息威胁优先级排序的可解释框架。 本套件基于PolyglotFakeFacts V2数据集的受控子集,包含2000篇新闻文章(其中虚假新闻、真实新闻各1000篇),且两类样本的语言分布完全一致。 该子集通过确定性语言平衡采样流程构建,随机种子设为20260819。 本研究评估的实验版ITPS包含三个核心组件:可控泄露真实性风险(leakage-controlled authenticity risk,简称A)、嵌入衍生安全相关性(embedding-derived security relevance,简称S)以及跨语言叙事共存度(cross-lingual narrative co-presence,简称P)。安全相关性与跨语言语义相似度均通过sentence-transformers/paraphrase-multilingual-mpnet-base-v2多语言嵌入模型计算得到。 本套件包含实验子集、样本选择与来源说明文档、已执行的实验笔记本、泄露与鲁棒性审计报告、预计算的折外真实性概率、最终的文章级A/S/P/ITPS评分结果、排名反转分析、敏感性分析,以及实验环境与方法论说明文档。 真实性组件的评估采用了源/域不相交且感知近重复的五折交叉验证,并移除了文本模板与源指纹。本ITPS实验所使用的最终可控泄露真实性概率达到了ROC-AUC=0.943520的性能。 本套件旨在提升已报道的ITPS实验的透明度、可验证性与可复现性。其属于衍生的实验研究成果,并非原始PolyglotFakeFacts数据集的替代版本或新版本。 原始数据集:Ciobanu, Alexandru (2026), 《PolyglotFakeFacts:涵盖政治、安全与社会领域的多语言虚假与真实新闻数据集》,Mendeley Data,V2,DOI:10.17632/gff8bmr4ff.2.

创建时间:
2026-08-23
二维码
社区交流群
二维码
科研交流群
商业服务