rejected-agriculture-instructions-dataset
收藏资源简介:
这是一个尼泊尔语源基指令数据集(被拒绝版本),由NVIDIA NeMo Data Designer从权威尼泊尔语文档(包括农业手册和法律文本)自动生成。数据集中的每条指令调优样本均严格基于源文本,答案完全来自源文档;对于无法回答的问题,模型会明确给出拒绝回答。每条记录采用聊天消息格式(messages),并附带元数据(metadata)以及每条记录的质量评分(quality_scores),评分维度包括 grounding(基础性)、correctness(正确性)和 naturalness(自然性),评分范围为 1-5,由 LLM 作为评判者给出。数据以 JSONL 格式存储,每个源文档对应一个分片文件(data/train-<shard>.jsonl),重新运行时分片会被幂等覆盖。重要提示:本数据集中的所有记录均低于预设的质量门槛,每条记录都包含拒绝原因(reject_reasons)。因此,数据集不适合直接用于指令调优;它适用于评判校准、难负样本挖掘,或者使用不同阈值进行重新过滤。
This is a Nepali language source-based instruction dataset (rejected version), automatically generated by NVIDIA NeMo Data Designer from authoritative Nepali documents (including agricultural manuals and legal texts). Each instruction tuning sample in the dataset is strictly based on the source text, with answers entirely derived from the source documents; for unanswerable questions, the model explicitly gives a refusal response. Each record is in a chat message format (messages), accompanied by metadata and quality scores (quality_scores) for each record, with dimensions including grounding, correctness, and naturalness, scored on a scale of 1-5, given by an LLM as the judge. The data is stored in JSONL format, with each source document corresponding to a shard file (data/train-<shard>.jsonl), and shards are idempotently overwritten upon re-execution. Important note: All records in this dataset are below the preset quality threshold, and each record contains rejection reasons (reject_reasons). Therefore, the dataset is not suitable for direct instruction tuning; it is suitable for judge calibration, hard negative mining, or re-filtering with different thresholds.
数据集概述
该数据集是一个尼泊尔语、基于权威来源的指令数据集(被拒绝版本),由 NVIDIA NeMo Data Designer 从尼泊尔官方文档(农业手册、法律文本)中生成。
核心特点
- 语言:尼泊尔语
- 任务类别:文本生成
- 标注/标签:尼泊尔语、指令微调、农业、法律、合成、基于来源
数据说明
- 答案严格限定于源文档内容;对于无法回答的问题,会给出明确拒绝的回复。
- 记录采用聊天
messages格式,并附带metadata和逐条记录的quality_scores(基于来源性/正确性/自然度,评分范围1-5,由LLM作为评判者)。 - 数据以
data/train-<shard>.jsonl形式存储,每个源文档对应一个分片;分片在重新运行时会被幂等地覆盖。
数据性质与用途
- 这些记录均被判定为低于质量门槛,每条记录都包含
reject_reasons。 - 适用场景:评判标准校准、困难负样本挖掘,或使用不同阈值进行重新过滤。
- 不适用场景:不应用于直接的指令微调。





