遇见数据集

hinglishNorm

收藏
arXiv2020-10-18 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

hinglishNorm是由Vahan Inc创建的一个包含13494个Hindi-English混合语句的数据集,专门用于文本规范化任务。该数据集的特点是每个句子都配有其人工标注的规范化形式,旨在解决非标准文本到标准格式的转换问题。数据集的创建过程包括数据收集、过滤与清洗、以及人工标注等步骤。hinglishNorm主要应用于自然语言处理领域,特别是在为印度用户构建互联网应用时,处理混合语言文本的需求。

hinglishNorm is a dataset containing 13,494 Hindi-English code-mixed utterances, developed by Vahan Inc specifically for text normalization tasks. Each utterance in the dataset is paired with its manually annotated normalized form, targeting the challenge of converting unstandardized text into its standardized formal equivalent. The construction of hinglishNorm includes steps such as data collection, filtering and cleaning, and manual annotation. This dataset is primarily applied in the field of natural language processing (NLP), especially to address the demand for processing code-mixed text when building internet applications for Indian users.

提供机构:
Vahan Inc
创建时间:
2020-10-18
二维码
社区交流群
二维码
科研交流群
商业服务