遇见数据集

New media corpus: Jututoad

收藏
data.europa2022-11-23 更新2025-05-31 收录
官方服务:

资源简介:

This corps contains recordings of internet chat rooms from 2003 and 2006. 300 files, 7 million words Unlike other corpuses of new media (newsgroups, forums, comments), there are no two versions of the corpus of chat rooms - with repetitions and removed with repetitions - because the corpus of chat rooms has not been massively used to quote the previous post.

本语料库(corpus)收录了2003年与2006年的互联网聊天室聊天记录,共计包含300个文件,总词量达700万。与其他新媒体语料库(涵盖新闻组、论坛、评论板块类语料库)不同,本聊天室语料库并未设置「带重复版本」与「去重版本」两类分支,究其原因在于该聊天室语料库极少被用于引用前文内容。

提供机构:
Tartu Ülikool
二维码
社区交流群
二维码
科研交流群
商业服务