遇见数据集

BhashaHMPV Dataset : Multilingual HMPV News and Fact-Check Articles Dataset for Indian Regional Languages

收藏
Zenodo2025-04-11 更新2026-05-26 收录
官方服务:

资源简介:

For the collection of Google News articles on HMPV in the Indian context, we scraped the articles using Python-based framework known as Splinter. In the script, we queried terms and phrases such as “HMPV”, “hmpv india”. The results included articles from the Google News website in a variety of languages, and the websites’ domains and languages were noted. We automated the URL by changing the language filter of the website Also in some cases, all articles were scraped and those unrelated to HMPV were filtered out in the pre-processing stage. All the samples collected were then put together into one CSV.We retrieved articles in ten Indian languages supported by Google News, namely: Bengali, English, Gujarati, Hindi, Marathi, Malayalam, Punjabi, Tamil, Telugu, Urdu, and Kannada.We also performed stemming for each language, and the stemmed outputs were added as separate columns in the respective language-specific sheets of the final CSV file. The following information was extracted along with the news articles:1) language of the Google News article2) title of the Google news article3) source of the Google news article (if available)4) link of the Google news article5) content of the Google news article6) domain of the article For the collection of Google Fact-Check articles, we used the Google Fact-Check API key to fetch the articles.In the python script, we queried terms and phrases such as “HMPV”, "hmpv india".We also performed stemming for each language, and the stemmed outputs were added as separate columns in the respective language-specific sheets of the final CSV file.The following information was extracted along with the news articles: 1) claim-text of the Google fact-check article2) claimant of the Google fact-check article3) claim-date of the Google fact-check article4) review-publisher of the Google fact-check article5) review-title of the Google fact-check article6) review-url of the Google fact-check article7) review-date of the Google fact-check article8) textual-rating of the Google fact-check article9) extracted-content of the Google fact-check article

针对印度语境下关于人偏肺病毒(HMPV)的谷歌新闻文章数据集,我们使用基于Python的Splinter框架完成文章爬取。在脚本中,我们以"HMPV""hmpv india"等词条与短语进行检索。检索结果涵盖谷歌新闻网站的多语言文章,并记录了各来源网站的域名与语言属性。我们通过调整网站的语言筛选条件实现URL自动化生成;部分场景下会爬取全部文章,随后在预处理阶段过滤掉与HMPV无关的内容。所有采集到的样本最终整合为单个CSV文件。我们共获取了谷歌新闻支持的10种印度语言的文章,分别为孟加拉语、英语、古吉拉特语、印地语、马拉地语、马拉雅拉姆语、旁遮普语、泰米尔语、泰卢固语、乌尔都语与卡纳达语。我们还针对每种语言执行词干提取操作,并将词干提取结果作为独立列,添加至最终CSV文件中对应语言的专属工作表内。 伴随新闻文章一同提取的信息包括: 1) 谷歌新闻文章的语言属性 2) 谷歌新闻文章的标题 3) 谷歌新闻文章的来源(如可获取) 4) 谷歌新闻文章的链接 5) 谷歌新闻文章的正文内容 6) 文章的所属域名 针对谷歌事实核查文章数据集,我们借助谷歌事实核查API密钥获取文章。在Python脚本中,我们同样以"HMPV""hmpv india"等词条与短语进行检索。我们也针对每种语言执行词干提取操作,并将词干提取结果作为独立列,添加至最终CSV文件中对应语言的专属工作表内。 伴随事实核查文章一同提取的信息包括: 1) 谷歌事实核查文章的主张文本(claim-text) 2) 谷歌事实核查文章的主张发布主体(claimant) 3) 谷歌事实核查文章的主张发布日期(claim-date) 4) 谷歌事实核查文章的核查发布方(review-publisher) 5) 谷歌事实核查文章的核查标题(review-title) 6) 谷歌事实核查文章的核查链接(review-url) 7) 谷歌事实核查文章的核查日期(review-date) 8) 谷歌事实核查文章的文本评级(textual-rating) 9) 谷歌事实核查文章的提取内容(extracted-content)

提供机构:
Zenodo
创建时间:
2025-04-11
二维码
社区交流群
二维码
科研交流群
商业服务