遇见数据集

Pairwise Duplicate News Dataset

收藏
Mendeley Data2026-04-09 收录
官方服务:

资源简介:

The Pairwise Duplicate News Dataset consists of 10,836 article pairs (5,418 duplicate and 5,418 non-duplicate) collected from seven major Pakistani English news outlets, DAWN, The Express Tribune, GEO News, The News International, The Nation, 24 News, and 92 News HD. The dataset was created from a larger corpus of 54,929 news articles covering eight categories: National, World, Business, Sports, Entertainment, Crime, Technology, and Health. The dataset was constructed using Sentence-BERT embeddings and FAISS cosine similarity search, with a similarity threshold of 0.85 for identifying duplicate articles. Each duplicate cluster represents a single real-world news event, and one pair was sampled per cluster. Non-duplicate pairs were formed by randomly pairing articles from different clusters and stratified into three similarity levels: easy (0.2-0.3), medium (0.3-0.5), and hard (0.5-0.7) to ensure balanced evaluation difficulty.

成对重复新闻数据集(Pairwise Duplicate News Dataset)共计包含10836条文章对(重复对与非重复对各5418条),数据采集自巴基斯坦7家主流英文新闻媒体:DAWN、《论坛快报(The Express Tribune)》、GEO新闻(GEO News)、《国际新闻报(The News International)》、《民族报(The Nation)》、24 News以及92 News HD。该数据集源自规模达54929条新闻文章的更大语料库,覆盖国内、国际、商业、体育、娱乐、犯罪、科技、健康共8个新闻分类。 本数据集通过句子BERT(Sentence-BERT)嵌入与FAISS余弦相似度搜索构建,用于识别重复文章的相似度阈值设定为0.85。每个重复簇对应一个真实新闻事件,且每个簇仅采样一对文章。非重复对通过随机配对不同簇的新闻文章生成,并按相似度分层为三类评估难度等级:简单(0.2-0.3)、中等(0.3-0.5)与困难(0.5-0.7),以确保评估难度分布均衡。

二维码
社区交流群
二维码
科研交流群
商业服务