compar:IA
收藏资源简介:
compar:IA是由法国政府开发的开源数字公共服务平台,旨在收集大规模法语人类偏好数据。该数据集包含超过60万条自由形式的提示和25万条偏好投票,其中89%的数据为法语。数据通过盲对比较界面收集,涵盖多轮对话和用户反馈。数据集发布在Hugging Face和data.gouv.fr上,采用Etalab 2.0开放许可。该数据集主要用于多语言模型训练、评估和人类-AI交互研究,旨在解决非英语语言模型性能和文化对齐不足的问题。
compar:IA is an open-source digital public service platform developed by the French government, designed to collect large-scale French human preference data. This dataset contains over 600,000 free-form prompts and 250,000 preference votes, with 89% of the data being in French. The data is collected via a blind pairwise comparison interface, covering multi-turn conversations and user feedback. The dataset is released on Hugging Face and data.gouv.fr under the Etalab 2.0 open license. It is primarily used for multilingual model training, evaluation, and human-AI interaction research, aiming to address the gaps in performance and cultural alignment of non-English language models.

- 1compar:IA: The French Government's LLM arena to collect French-language human prompts and preference data法国政府 · 2026年



