相关数据集
Self-GRIT/wikitext-2-raw-v1-preprocessed-1k
该数据集包含一个名为text的字段,数据类型为字符串。数据集只有一个train分割,包含1000个样本,占用的字节数为301261.42491421103。数据集的下载大小为438681字节,数据集大小为301261.42491421103字节。
Hugging Face2024-07-24 更新250
BR4-004 - Bislama stops
Pilot Bislama stop wordlist with LE. Language as given:
Research Data Australia120
omarelsayeed/good_chats_dataset_pre_tokenization
--- dataset_info: features: - name: Chat_ID dtype: string - name: text dtype: string splits: - name: train num_bytes: 97515949 num_examples: 33735 download_size: 30316154
Hugging Face2024-02-23 更新160
[ Pig ] Normalized Requirement Files
The documents of requirements are normalized by standard pre-processing techniques including splitting identifiers, special token elimination, stemming, and stop word removal.
NIAID Data Ecosystem110



