davanstrien/f1lift-augmented
收藏资源简介:
davanstrien/f1lift-augmented是一个由classify-and-augment工具生成的LLM标注数据集。该数据集使用HuggingFaceTB/SmolLM3-3B模型进行标注,包含positive和negative两个标签。原始输入数据有180行,经过处理后输出202行数据。标签分布显示,negative标签有140条真实数据,positive标签有40条真实数据和22条合成数据。合成审计部分详细记录了positive标签的合成过程,包括生成64条候选数据,验证并保留了22条,接受率为93.8%。
davanstrien/f1lift-augmented is an LLM-annotated dataset produced by classify-and-augment. The dataset uses the HuggingFaceTB/SmolLM3-3B model for annotation and includes two labels: positive and negative. The original input consists of 180 rows, and the processed output contains 202 rows. The label distribution shows 140 real data points for the negative label and 40 real data points plus 22 synthetic data points for the positive label. The synthesis audit details the process for the positive label, including generating 64 candidate data points, validating and keeping 22, with an acceptance rate of 93.8%.



