遇见数据集

saillab/alpaca-belarusian-cleaned

收藏
Hugging Face2024-09-20 更新2025-04-12 收录
官方服务:

资源简介:

--- language: - be pretty_name: Belarusian alpaca-52k size_categories: - 100K<n<1M --- This repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: [OpenReview](https://openreview.net/forum?id=02MLWBj8HP) If you have used our dataset, please cite it as follows: **Citation** ``` @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}, url={https://openreview.net/forum?id=02MLWBj8HP} } ``` The original dataset [(Alpaca-52K)](https://github.com/tatsu-lab/stanford_alpaca?tab=readme-ov-file#data-release) was translated using Google Translate. **Copyright and Intended Use** This dataset has been released under CC BY-NC, intended for academic and research purposes only. Please review the licenses and terms and conditions of Alpaca-52K, Dolly-15K, and Google Cloud Translation before using this dataset for any purpose other than research.

语言: - 白俄罗斯语 数据集展示名称:白俄罗斯语Alpaca-52K 规模分类: - 10万<n<100万 本仓库包含用于TaCo论文的数据集。 如需获取更多细节,请参阅该论文:[OpenReview](https://openreview.net/forum?id=02MLWBj8HP) 若您使用了本数据集,请按以下方式引用: **引用格式** @inproceedings{upadhayay2024taco, title={TaCo: 通过翻译辅助思维链(Chain-of-Thought)流程提升低资源语言在大语言模型(Large Language Model,LLM)中的跨语言迁移能力}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={第5届有限/低资源场景下的实用机器学习研讨会,国际学习表征大会(ICLR)}, year={2024}, url={https://openreview.net/forum?id=02MLWBj8HP} } 本数据集的原始版本[(Alpaca-52K)](https://github.com/tatsu-lab/stanford_alpaca?tab=readme-ov-file#data-release)通过谷歌翻译(Google Translate)完成译制。 **版权与使用意图** 本数据集采用CC BY-NC许可协议发布,仅用于学术与研究用途。若您将本数据集用于研究以外的任何用途,请先查阅Alpaca-52K、Dolly-15K以及谷歌云翻译(Google Cloud Translation)的许可协议与条款细则。

提供机构:
saillab
二维码
社区交流群
二维码
科研交流群
商业服务