相关数据集
Self-GRIT/wikitext-2-raw-v1-preprocessed-1k
该数据集包含一个名为text的字段,数据类型为字符串。数据集只有一个train分割,包含1000个样本,占用的字节数为301261.42491421103。数据集的下载大小为438681字节,数据集大小为301261.42491421103字节。
Hugging Face2024-07-24 更新250
g3_sum
Urdu Talk Show Script Summarization Dataset with Training and Testing Files
kaggle2024-05-20 更新120
dis-project-1-preprocessed-documents
Project corpus documents with stopwords removal, stemming and tokenization.
kaggle2024-11-03 更新100
Pre-processing steps.
Example of pre-processing steps performed in a extract for the purpose of creating a adjacency network. The pre-processing step eliminates punctuation marks and words conveying low semantic content. T
NIAID Data Ecosystem130
[ Pig ] Normalized Requirement Files
The documents of requirements are normalized by standard pre-processing techniques including splitting identifiers, special token elimination, stemming, and stop word removal.
NIAID Data Ecosystem110



