遇见数据集

Targeted domain-specific (wine and olive oil) parallel (Spanish-English) corpus

收藏
Zenodo2026-03-02 更新2026-05-26 收录
官方服务:

资源简介:

Domain-specific corpus of 19,566,171 parallel segments (Spanish-English) belonging to the domain of olive oil and wine. The dataset was compiled through targeted webcrawling of wine and olive oil producers' webpages from Spain, combined with the Spanish-English ParaCrawl corpus, hence the Restricted status of this dataset due to potential copyright conflicts. The data were automatically curated to filter domain-specific parallel segments by building small statistical n-gram language models stored in SQLite databases using kenlm (https://github.com/kpu/kenlm). The different language models were calculated based on several small, monolingual comparable corpora: one Spanish wine-related corpus, one Spanish olive oil-related corpus, one English wine-related corpus, and one English olive oil-related corpus. For research purposes only.

提供机构:
Zenodo
创建时间:
2026-03-02
二维码
社区交流群
二维码
科研交流群
商业服务