遇见数据集

relaion2B-en-research-safe-japanese-translation

收藏
魔搭社区2026-04-28 更新2026-08-02 收录
官方服务:

资源简介:

# relaion2B-en-research-safe-japanese-translation This dataset is the Japanese translation of the English subset of ReLAION-5B ([laion/relaion2B-en-research-safe](https://huggingface.co/datasets/laion/relaion2B-en-research-safe)), translated by [gemma-2-9b-it](https://huggingface.co/datasets/laion/relaion2B-en-research-safe). We used [text2dataset](https://github.com/llm-jp/text2dataset) for translating with open-weight LLMs. By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese. ## Prompt The following is the prompt used for translation with Gemma. > You are an excellent English-Japanese translator. Please translate the following sentence into Japanese. > > You must output only the translation. > > Sentence: {passage} > > Translation: ## LICENSE This dataset is licensed under Apache 2.0, inheriting the license from the [laion/relaion2B-en-research-safe](https://huggingface.co/datasets/laion/relaion2B-en-research-safe) dataset. ## Bibtex ``` @inproceedings{sugiura-etal-2025-developing, title = "Developing {J}apanese {CLIP} Models Leveraging an Open-weight {LLM} for Large-scale Dataset Translation", author = "Sugiura, Issa and Kurita, Shuhei and Oda, Yusuke and Kawahara, Daisuke and Okazaki, Naoaki", editor = "Ebrahimi, Abteen and Haider, Samar and Liu, Emmy and Haider, Sammar and Leonor Pacheco, Maria and Wein, Shira", booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop)", month = apr, year = "2025", address = "Albuquerque, USA", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2025.naacl-srw.15/", pages = "162--170", ISBN = "979-8-89176-192-6", abstract = "CLIP is a foundational model that bridges images and text, widely adopted as a key component in numerous vision-language models.However, the lack of large-scale open Japanese image-text pairs poses a significant barrier to the development of Japanese vision-language models.In this study, we constructed a Japanese image-text pair dataset with 1.5 billion examples using machine translation with open-weight LLMs and pre-trained Japanese CLIP models on the dataset.The performance of the pre-trained models was evaluated across seven benchmark datasets, achieving competitive average scores compared to models of similar size without the need for extensive data curation. However, the results also revealed relatively low performance on tasks specific to Japanese culture, highlighting the limitations of translation-based approaches in capturing cultural nuances. Our dataset, models, and code are publicly available." } ``` # Reference - https://laion.ai/blog/relaion-5b

提供机构:
maas
创建时间:
2025-11-25
二维码
社区交流群
二维码
科研交流群
商业服务