遇见数据集

Toward A Comparable Corpus Of Latvian, Russian And English Tweets

收藏
Zenodo2017-05-22 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

Twitter has become a rich source for linguistic data. Here, a possibility of building a trilingual Latvian-Russian-English corpus of tweets from Riga, Latvia is investigated. Such a corpus, once constructed, might be of great use for multiple purposes such as training machine translation models, examining cross-lingual phenomena and studying the population of Riga. This pilot study shows that it is feasible to build such a resource by building and analysing a pilot corpus, which is made publicly available and can be used to construct a large comparable corpus.

创建时间:
2017-05-22
二维码
社区交流群
二维码
科研交流群
商业服务