hh-rlhf-12k-ja
收藏资源简介:
# hh-rlhf-12k-ja This repository provides a human preference dataset developed by LLM-jp, a collaborative project launched in Japan. This dataset is a Japanese translation of a subset of [hh-rlhf](https://huggingface.co/datasets/Anthropic/hh-rlhf) using DeepL. This dataset consists of 12,000 entries randomly sampled from hh-rlhf. Specifically, it includes a random selection of 3,000 entries from the training splits of the four groups: harmless-base, helpful-base, helpful-online, and helpful-rejection-sampled. For more information on each group, please refer to the original dataset documentation. ## Send Questions to llm-jp(at)nii.ac.jp ## Model Card Authors The names are listed in alphabetical order. Hirokazu Kiyomaru, Hiroshi Matsuda, Jun Suzuki, Namgi Han, Saku Sugawara, Shota Sasaki, Shuhei Kurita, Taishi Nakamura, Takashi Kodama, Takumi Okamoto.



