相关数据集
atlasia/FineWeb2-Moroccan-Arabic-Predictions
该数据集包含文本和语言两个特征,数据类型均为字符串。数据集仅包含训练集,共有10,000个样本,总大小为2,742,010字节。数据集的下载大小为1,363,591字节。默认配置下的数据文件路径为data/train-*。
Hugging Face2024-12-13 更新150
Number of articles that mention different TCPR methods.
Number of articles that mention different TCPR methods.
NIAID Data Ecosystem130
sumyeongahn/sst2
--- dataset_info: features: - name: idx dtype: int64 - name: sentence dtype: string - name: label dtype: int64 - name: paraphrase dtype: string splits: - name: train
Hugging Face2024-05-28 更新80
llangnickel/long-covid-classification-data
该数据集包含与长COVID相关的PubMed文章摘要,这些文章由信息专家手动收集。数据集分为训练集、开发集和测试集,其中正例和负例的数量分别为:训练集215正例和199负例,开发集76正例和62负例,测试集70正例和68负例,总计690篇文章。
Hugging Face2022-11-24 更新150
tyzhu/find_sent_after_sent_train_100_eval_40
--- configs: - config_name: default data_files: - split: train path: data/train-* - split: validation path: data/validation-* dataset_info: features: - name: inputs dtype: string
Hugging Face2023-11-20 更新100



