Nels2/brainrot-custom-dataset
收藏资源简介:
Brainrot Fine-Tuning Dataset是一个用于监督微调实验的清洗后数据集,旨在训练模型将标准英语文本翻译成随意的“brainrot”网络用语风格。该数据集采用聊天格式,每个条目包含系统角色(说明翻译任务)、用户角色(提供原始标准英语文本)和助手角色(输出brainrot风格的重写文本)。风格特点包括小写措辞、随意俚语、表情符号、夸张表达和网络文化短语。数据集包含训练集(5663行)、验证集(503行)和测试集(505行),源自原始数据集shvn22k/brainrot-dataset。
The Brainrot Fine-Tuning Dataset is a cleaned dataset intended for supervised fine-tuning experiments, designed to train models to translate standard English text into casual brainrot internet-speak style. It uses a chat format, with each entry consisting of a system role (explaining the translation task), a user role (providing the original standard English text), and an assistant role (outputting the brainrot-style rewritten text). The style includes lowercase phrasing, casual slang, emojis, exaggerated expressions, and internet-culture phrasing. The dataset includes training (5663 rows), validation (503 rows), and test (505 rows) splits, derived from the original dataset shvn22k/brainrot-dataset.





