官方服务:
资源简介:
Wikipedia Language Identification
应用场景:
创建时间:
2020-08-21
相关数据集
Wikipedia Data
This simple text data that I personally used for learning Regex expressions.
kaggle2023-12-16 更新770
Politicians by Country from the English-language Wikipedia
This project contains data on most English-language Wikipedia articles within the category "Category:Politicians by nationality" and subcategories, along with the code used to generate that data. Both
DataCite Commons2020-09-01 更新110
Data from: Robust clustering of languages across Wikipedia growth
Wikipedia is the largest existing knowledge repository that is growing on a genuine crowdsourcing support. While the English Wikipedia is the most extensive and the most researched one with over five
DataONE2017-09-19 更新100
mmnga/wikipedia-ja-20230720-50k
该数据集是从[izumi-lab/wikipedia-ja-20230720]中随机抽取的50,000条记录,包含curid、title和text三个特征。数据集分为train分割,下载大小为78354971字节,数据集大小为134082445.03326812字节。
Hugging Face2023-09-25 更新180
Labeled Dataset Of Offensive Language In Arabic
Dataset Construction for the Detection of Anti-Social Behaviour in Online Commun
kaggle2023-08-29 更新100



