yeniguno/turkish-gibberish-detection
收藏官方服务:
资源简介:
这是一个包含文本和标签特征的数据集,总大小约为1.1GB,划分为训练集、验证集和测试集三个部分,分别包含945785、202668和202669个样本。数据集提供了默认配置,指定了各个数据集的文件路径。
This dataset includes text and label features, with a total size of approximately 1.1GB, divided into three parts: training set, validation set, and test set, containing 945785, 202668, and 202669 samples respectively. The dataset provides a default configuration specifying the file paths for each dataset split.
提供机构:
yeniguno


