Synthetic-JP-EN-Coding-Dataset-Magpie-69k
收藏资源简介:
# Synthetic-JP-EN-Coding-Dataset-Magpie-69k [Magpie](https://arxiv.org/abs/2406.08464)の手法を様々なモデルに対して適用し作成した、約69000件の日本語・英語のコーディング対話データセットです。 作成に利用したモデルは以下の通りです。`model`キーに該当レコードの作成に利用したモデル情報があります。 - [nvidia/Nemotron-4-340B-Instruct](https://huggingface.co/nvidia/Nemotron-4-340B-Instruct) - [microsoft/Phi-3-medium-4k-instruct](https://huggingface.co/microsoft/Phi-3-medium-4k-instruct) - [mistralai/Mixtral-8x22B-Instruct-v0.1](https://huggingface.co/mistralai/Mixtral-8x22B-Instruct-v0.1) - [cyberagent/calm3-22b-chat](https://huggingface.co/cyberagent/calm3-22b-chat) データセットの作成には[DeepInfra](https://deepinfra.com/)を利用しました。 また、[このリポジトリ](https://github.com/Aratako/magpie-nemotron)でデータセット作成に用いたコードを公開しています。これをベースに、プロンプトテンプレートやシステムプロンプト等を一部変更することで生成しています。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。



