MHBS-IHB/fishmt5
收藏资源简介:
--- license: apache-2.0 task_categories: - translation language: - zh - la size_categories: - 1M<n<10M tags: - bilingual pretty_name: Fish Names Chinese-Latin Parallel Corpora --- # Fish Names Chinese-Latin Parallel Corpora ## Dataset Overview We curated over 60,000 authoritative Chinese-Latin bilingual parallel corpora for fish names by integrating cross-source data, including Eschmeyer's Catalog of Fishes online database. Using a dual translation approach, we applied the Multilingual Text-to-Text Transfer Transformer (mT5) model to generate missing Chinese names. *Note: The current release provides 10,000 paired data entries.* ## Dataset Details - **Total Curated Records:** > 60,000 authoritative pairs - **Current Release:** 10,000 carefully reviewed Chinese-Latin name pairs - **Languages:** Chinese, Latin - **Data Sources:** - Eschmeyer's Catalog of Fishes online database - Other cross-source authoritative fish name databases - **Methodology:** - Data integration from multiple sources - Dual translation approach using mT5 to generate missing Chinese names - Rigorous quality control and review ## Intended Use - **Research:** Fish taxonomy, biodiversity studies, and ecological research - **Translation:** Evaluation and development of bilingual translation models - **Corpus Development:** Creation of high-quality multilingual corpora for biocultural diversity studies ## Limitations - **Data Size:** Although the full dataset includes over 60,000 pairs, only a subset of 10,000 pairs is provided in the current release. - **Review Status:** The associated research article is currently under review. Future updates will expand the dataset and include additional metadata. ## Citation If you use this dataset in your research, please cite the forthcoming publication (currently under review). ## Contact For any questions or further information, please contact the dataset curators.
许可证:Apache-2.0 task_categories: - 翻译 language: - 中文(zh) - 拉丁语(la) size_categories: - 100万 < n < 1000万 tags: - 双语 pretty_name: 鱼类名称汉拉平行语料库(Fish Names Chinese-Latin Parallel Corpora) --- # 鱼类名称汉拉平行语料库 ## 数据集概览 本数据集通过整合多源数据(含埃施迈尔鱼类在线目录数据库(Eschmeyer's Catalog of Fishes)),构建了逾6万条权威的鱼类名称汉拉双语平行语料。我们采用双翻译策略,借助多语言文本到文本迁移Transformer(Multilingual Text-to-Text Transfer Transformer,简称mT5)模型生成缺失的中文名称。 *注:本次发布仅提供10000条配对数据条目。* ## 数据集详情 - **总整理记录数:** 逾6万条权威配对数据 - **本次发布量:** 10000条经过严格审核的汉拉鱼类名称配对数据 - **涉及语言:** 中文、拉丁语 - **数据来源:** - 埃施迈尔鱼类在线目录数据库(Eschmeyer's Catalog of Fishes) - 其他多源权威鱼类名称数据库 - **构建方法:** - 多源数据整合 - 采用mT5模型结合双翻译策略生成缺失的中文名称 - 严格的质量管控与审核流程 ## 预期用途 - **研究领域:** 鱼类分类学、生物多样性研究及生态学研究 - **翻译方向:** 双语翻译模型的评估与开发 - **语料构建:** 用于生物文化多样性研究的高质量多语言语料创建 ## 局限性说明 - **数据规模:** 尽管完整数据集包含逾6万条配对数据,但本次发布仅提供其中10000条的子集。 - **审核状态:** 相关研究论文目前处于审稿阶段。未来更新将扩充数据集并补充更多元数据。 ## 引用说明 若您在研究中使用本数据集,请引用即将发表的(当前处于审稿阶段的)相关论文。 ## 联系方式 如有任何疑问或需要进一步信息,请联系数据集整理团队。




