遇见数据集

POS Hindi English

收藏
India Data2025-07-23 更新2026-05-16 收录
官方服务:

资源简介:

Each token in the dataset is labeled with a language code like 'en' for English, 'hi' for Hindi, and 'rest' for other or unclassified tokens. Corresponding entity tags follow the BIO tagging format, denoting the beginning (B), inside (I), or outside (O) of named entities such as PERSON, PLACE, or ORGANISATION. The data is sourced from social platforms like Twitter and includes features like mentions, hashtags, emojis, and links, reflecting real-world language complexity.

创建时间:
2025-06-03
二维码
社区交流群
二维码
科研交流群
商业服务