anggawarjaya18/simple-wiki
收藏官方服务:
资源简介:
该数据集是一个包含英文维基百科条目及其简化版本配对的集合。具体来说,每个数据点由原始文本(text列)和对应的简化文本(simplified列)组成,可用于训练句子嵌入模型,例如通过Sentence Transformers进行特征提取或句子相似性任务。数据集规模在10万到100万之间,属于单语(英语)数据集,主要用于特征提取和句子相似性任务。
This dataset is a collection of pairs of English Wikipedia entries and their simplified variants. It can be used directly with Sentence Transformers to train embedding models, and includes columns for text (original) and simplified (simplified version). The dataset is monolingual (English), with a size between 100K and 1M examples, and is categorized for feature-extraction and sentence-similarity tasks.
提供机构:
anggawarjaya18


