遇见数据集

anggawarjaya18/simple-wiki

收藏
Hugging Face2026-05-06 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个包含英文维基百科条目及其简化版本配对的集合。具体来说,每个数据点由原始文本(text列)和对应的简化文本(simplified列)组成,可用于训练句子嵌入模型,例如通过Sentence Transformers进行特征提取或句子相似性任务。数据集规模在10万到100万之间,属于单语(英语)数据集,主要用于特征提取和句子相似性任务。

This dataset is a collection of pairs of English Wikipedia entries and their simplified variants. It can be used directly with Sentence Transformers to train embedding models, and includes columns for text (original) and simplified (simplified version). The dataset is monolingual (English), with a size between 100K and 1M examples, and is categorized for feature-extraction and sentence-similarity tasks.

提供机构:
anggawarjaya18
二维码
社区交流群
二维码
科研交流群
商业服务