官方服务:
资源简介:
StopWords_DatesandNumbers
应用场景:
创建时间:
2021-02-17
相关数据集
nkandpa2/wiki-dolma
--- configs: - config_name: default data_files: - split: train path: - "wiki/archive/v3/documents/*.jsonl.gz" - config_name: wikiteam data_files: - split: train path: - "wiki/a
Hugging Face2024-10-31 更新150
kth8/text-cleanup-20000x
--- license: apache-2.0 lang: en --- Dataset for training small models to clean up noisy text. Each instance follows this format: ```json { "messages": [ { "role": "system", "conten
Hugging Face2026-03-18 更新50
Webis Gmane Email Corpus 2019
The Webis Gmane Email Corpus 2019 is a dataset of more than 153 million parsed and segmented emails crawled between February and May 2019 from gmane.io covering more than 20 years of public mailing li
NIAID Data Ecosystem40
[ Maven ] Normalized Requirement Files
The documents of requirements are normalized by standard pre-processing techniques including splitting identifiers, special token elimination, stemming, and stop word removal.
NIAID Data Ecosystem70
Dictionary for the CST Lemmatizer
Binary wordlists for the CST lemmatizer as suplement to the rules of the lemmatizer. Works with both tagged and untagged input. Use: cstlemma -d NAME-OF-WORDLIST
B2FIND40



