遇见数据集

TAUS Language Translation Data | Parallel translation for Colloquial English into various languages for Machine Learning

收藏
Datarade2024-04-19 收录
官方服务:

资源简介:

The corpus is a great fit for training chat bots or social media content, and will give the conversation with your local audience a friendly, casual tone. From product user reviews and blog post comments to everyday business small talk, your MT engine will be able to handle even the most creative user voices. This corpus contains over 1 million words, and a total vocabulary of more than 37000 different words. Need more data? In the following months, TAUS will release more equally sized corpora for the same domain and language combinations, with a significant increase of vocabulary. English - Hindi English - Urdu English - Tamil English - Nepali English - Turkish English - Pashto English - Sorani English - Bengali English - Burmese English - Assamese English - Telugu English - Sinhalese English - Dari English - Punjabi (Pakistan) English - Punjabi (India) English - Lao English - Kurmanji (lat) English - Kurmanji (arab) Other languages are available on demand.

提供机构:
TAUS
搜集汇总
数据集介绍
TAUS Language Translation Data | Parallel translation for Colloquial English into various languages for Machine Learning 数据集图片
背景与挑战
背景概述
该数据集专为机器学习设计,提供英语与多种语言(如印地语、乌尔都语、泰米尔语等)的平行翻译语料,适用于训练聊天机器人或处理社交媒体内容,以增强对话的友好和随意感。它包含超过100万词和37000个词汇,未来还将发布更多相同领域和语言组合的扩展语料。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务