遇见数据集

esCorpiusDialog: A Large-Scale Multilingual Dialogue Dataset in Spanish, Catalan, Basque, and Galician

收藏
Zenodo2026-02-04 更新2026-05-26 收录
官方服务:

资源简介:

In response to the growing need for comprehensive conversational datasets, we introduce esCorpiusDialog, a large-scale, diverse dialogue corpus designed to facilitate the development of conversational models in Spanish, Catalan, Basque, and Galician. Addressing the critical shortage of high-quality multi-turn conversational data for these languages, the current release comprises 30,518,397 dialogues with 129,705,492 conversational turns and 2,187,821,082 tokens (≈15.8 GiB). On average, each dialogue contains 4.3 turns, supporting modeling of multi-turn interactions at scale. esCorpiusDialog aggregates data from multiple sources: movie and TV subtitles (OpenSubtitles), newsgroups (Usenet), online forums (Menéame, Mediavida, Reddit), and literature (Project Gutenberg). The corpus includes approximately 30.2 million dialogues in Spanish, ~116k in Basque, >92k in Catalan, and 63k in Galician. Compared to earlier versions of this record, the forum-derived components are now released in a dehydrated (metadata-only) form to align with platform Terms of Service: we do not redistribute raw forum text or user handles. Instead, we share thread/post identifiers and ordered sequences of comment identifiers representing root-to-leaf dialogue paths, which preserve the dialogue structure and enable reproducibility. Rehydration scripts are provided to allow users to re-download the underlying content directly from the original platforms using these identifiers, under the user’s responsibility and subject to each platform’s access policies and Terms of Service. The dataset has undergone processing to clearly define dialogue boundaries and turn segmentation, including basic normalization (e.g., removal of URLs/emails and anonymization during internal processing for forum/newsgroup data). With its breadth of topics and varied dialogue styles across sources, esCorpiusDialog provides a high-coverage resource for training and analyzing open-domain conversational systems and for studying multi-turn conversational phenomena in Spanish and Spain’s co-official languages.

提供机构:
Zenodo
创建时间:
2026-02-04
二维码
社区交流群
二维码
科研交流群
商业服务