Targeted domain-specific (wine and olive oil) parallel (Spanish-English) corpus
收藏资源简介:
Domain-specific corpus of 19,566,171 parallel segments (Spanish-English) belonging to the domain of olive oil and wine. The dataset was compiled through targeted webcrawling of wine and olive oil producers' webpages from Spain, combined with the Spanish-English ParaCrawl corpus, hence the Restricted status of this dataset due to potential copyright conflicts. The data were automatically curated to filter domain-specific parallel segments by building small statistical n-gram language models stored in SQLite databases using kenlm (https://github.com/kpu/kenlm). The different language models were calculated based on several small, monolingual comparable corpora: one Spanish wine-related corpus, one Spanish olive oil-related corpus, one English wine-related corpus, and one English olive oil-related corpus. For research purposes only.



