相关数据集
bianet
# Dataset Card for Bianet ## Table of Contents - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leade
魔搭社区2025-12-05 更新180
Appendix 1 CoMeT and MLE sampling frames
Appendix 1 represents the sampling frame for the Corpus of Medical Textbooks (CoMeT) for the article On the Replicability of Corpus-Derived Medical Word Lists published in Applied Corpus Lingui
DataCite Commons2025-06-01 更新90
aimaralab/aimaralab
这是一个由AiMara Lab创建的Aymara-Spanish平行语料库,基于manythings.org的Spanish-English平行语料库(来源于Tatoeba)。该语料库经过母语为Aymara的人手动校对,使用Google Translate模型生成,并由人类验证者进行后期编辑。此资源的目的是促进NLP工具和针对Aymara语言的自动翻译模型的开发。
Hugging Face2025-11-06 更新100
commoncrawl/gneissweb-annotation-host-testing-v1
GneissWeb注释数据集是一个应用于Common Crawl语料库的质量和类别注释数据集,由IBM Research的GneissWeb方法提供支持。该数据集支持对医疗、教育、技术和科学领域网络内容的精确过滤,便于为研究项目、语言模型和专业应用构建高质量语料库。数据集包含两个层次的注释粒度:主机级别(整个域的聚合统计)和URL级别(单个URL分类)。数据集利用了IBM公开提供的GneissWe
Hugging Face2025-12-11 更新50



