遇见数据集

Language Resources for Intrinsic Plagiarism Detection in Urdu Language

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This is a dataset based on the intrinsic plagiarism . To produce a high-quality dataset to train the classification algorithm, we have gathered the Urdu essays and reports from various popular and highly trending websites such as, jang.com, urduessaypoint.blogspot.com, www.dawnnews.tv etc. All the documents gathered from the websites are then compiled in .txt format. More than 2500 plagiarized and non plagiarized documents are created systematically

创建时间:
2024-03-25
二维码
社区交流群
二维码
科研交流群
商业服务