遇见数据集

agentlans/en-region-classifier-dataset

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

这是一个英语区域分类数据集,用于根据文本内容识别英语文本的来源国家或地区。数据集包含来自多个英语国家/地区的文本数据,如澳大利亚、加拿大、印度、爱尔兰、英国、美国等,总计90万条数据,分为训练集(81万条)、验证集(4.5万条)和测试集(4.5万条)。数据可用于文本分类任务,支持监督学习和本地化研究,适用于NLP模型训练和评估。

This is an English region classification dataset designed to identify the country or region of origin of English text based on content. The dataset includes text data from multiple English-speaking countries/regions, such as Australia, Canada, India, Ireland, the United Kingdom, the United States, etc., totaling 900,000 entries split into training set (810,000), validation set (45,000), and test set (45,000). It is suitable for text classification tasks, supporting supervised learning and localization research, and can be used for NLP model training and evaluation.

提供机构:
agentlans
二维码
社区交流群
二维码
科研交流群
商业服务