遇见数据集

Pashtu Text Sentiment Analysis Dataset

收藏
Mendeley Data2026-04-09 收录
官方服务:

资源简介:

We gathered information from respected Pashtu newspapers, such as those published by the Wahdat Newspaper, and articles written by professors from various areas. This data collection took place over a month, encompassing regular newspaper issues and contributions from professors. All text was acquired digitally from faculty members of the Pashtu department and the Wahdat Newspaper. Initially, we amassed a total of 29,000 sentences in their raw form. Later, we conducted further processing, including cleaning and preprocessing, resulting in around 21,800 sentences of refined data. Subsequently, domain experts meticulously examined the data and attributed sentiments to each sentence.

本数据集的采集来源包含两部分:一是权威普什图语(Pashto)报刊,例如瓦赫达特报(Wahdat Newspaper)发行的出版物;二是各领域学者的撰稿文章。本次数据采集周期为一个月,涵盖正规报刊刊发内容与学者投稿作品。所有文本均通过数字化方式从普什图语系(Pashtu Department)教职员工及瓦赫达特报团队处获取。初始阶段,我们共累计得到29000条原始语句。后续我们对数据开展了清洗、预处理等进一步处理工作,最终得到约21800条精炼数据语句。随后,领域专家对该数据集进行了细致审核,并为每条语句标注了情感倾向。

二维码
社区交流群
二维码
科研交流群
商业服务