遇见数据集

Reddit EU language dataset

收藏
Zenodo2021-08-31 更新2026-05-28 收录
数据链接:
官方服务:

资源简介:

This dataset has been created for a personal project related to the recognition of the original language of someone writing in english. <strong>Origin</strong> The dataset has been crawled from the subreddit r/europe and contains around 1.5 milions posts in it's raw form. <strong>Structure</strong> This repo contains both the raw data and the cleaned data, the latter, purged of deleted comments and of those that were not linked to the provenience of the writer, contains around 450k datapoints and has the following structure: body: the text content of the comment country_name: extended name of the country permalink: link to the comment author: username of the creator created_utc: utc creation datetime: date and time of creation alpha2: ISO country alpha2 code alpha3: ISO country alpha3 code numeric: ISO country number apolitical_name: apolitical country name

提供机构:
Zenodo
创建时间:
2021-08-31
二维码
社区交流群
二维码
科研交流群
商业服务