Shaer-AI/shaer-sft-test-generations-k5
收藏资源简介:
该数据集名为Shaer SFT Test Generations K=5,包含从最终Shaer SFT适配器在测试集上生成的阿拉伯诗歌输出。具体来说,它基于源数据集Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits的测试分割(包含3,481行),每行生成5个样本,总计17,405行生成数据。生成过程使用基础模型Navid-AI/Yehia-7B-preview和适配器仓库Shaer-AI/Shaer-adapters中的适配器(位于adapters/fresh_sft/train/best子文件夹),通过带LoRA的vLLM后端实现,解码设置包括最大新令牌数768、温度0.55、top_p 0.9和重复惩罚1.08。数据集仅包含健康的生成行,排除了本地失败的重试尝试。
This dataset contains generated outputs from the final Shaer SFT adapter on the held-out test split. It is based on the source dataset Shaer-AI/ashaar-with-enhanced-descriptions-baseform-final-sft-lte20-min500-splits (test split with 3,481 rows), generating 5 samples per source row for a total of 17,405 generated rows. Generation uses the base model Navid-AI/Yehia-7B-preview and an adapter from the Shaer-AI/Shaer-adapters repository (subfolder adapters/fresh_sft/train/best), implemented with vLLM backend with LoRA, with decoding settings including max_new_tokens 768, temperature 0.55, top_p 0.9, and repetition_penalty 1.08. Only healthy generated rows are included, and local failed retry attempts are excluded.



