遇见数据集

davanstrien/imdb-classify-augment-v5-distfilter

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

这是一个由LLM(大型语言模型)标注的IMDB电影评论数据集,使用classify-and-augment工具生成。数据集包含正面(positive)和负面(negative)两类情感标签,使用HuggingFaceTB/SmolLM3-3B模型处理。原始输入20行数据,最终输出25行数据(包含真实数据和合成数据)。标签分布显示:负面标签13个(9真实+4合成),正面标签12个(11真实+1合成)。合成数据经过模型自一致性检查验证,接受率为50%。

This is an LLM-annotated IMDB movie review dataset produced using the classify-and-augment tool. The dataset contains two sentiment labels: positive and negative, processed by the HuggingFaceTB/SmolLM3-3B model. It started with 20 input rows and produced 25 output rows (including real and synthetic data). Label distribution shows: 13 negative labels (9 real + 4 synthetic) and 12 positive labels (11 real + 1 synthetic). Synthetic data was validated through model self-consistency checks with a 50% acceptance rate.

提供机构:
davanstrien
二维码
社区交流群
二维码
科研交流群
商业服务