Matina
收藏资源简介:
Matina是一个由塔比阿特莫达勒斯大学和德黑兰大学共同创建的大型波斯语文本语料库,包含72.9亿个tokens,经过精细的数据预处理和去重处理以确保高质量。该数据集整合了公开的波斯语数据集以及新收集的数据源,以确保内容的多样性和事实性信息的包含。Matina语料库旨在促进波斯语的自然语言处理,支持大型语言模型的预训练以及基于transformer架构的较小模型的发展,适用于文本分类、机器翻译、情感分析等多种NLP任务。
Matina is a large-scale Persian text corpus jointly developed by Tarbiat Modares University and University of Tehran, comprising 7.29 billion tokens. It has undergone rigorous data preprocessing and deduplication to ensure high data quality. This corpus integrates publicly available Persian datasets and newly collected data sources to guarantee content diversity and the inclusion of factual information. The Matina corpus aims to advance Persian natural language processing (NLP), support pre-training of large language models (LLMs) and the development of smaller models based on the Transformer architecture, and is applicable to various NLP tasks such as text classification, machine translation, and sentiment analysis.

- 1Matina: A Large-Scale 73B Token Persian Text Corpus塔比阿特莫达勒斯大学,德黑兰大学 · 2025年



