官方服务:
资源简介:
This corpus contains blog posts. The corpus is available through a dedicated concordancer.
应用场景:
创建时间:
2023-10-13
相关数据集
GEIZIG1^ (MEINE DGS – annotiert. Öffentliches Korpus der Deutschen Gebärdensprache, 1. Release)
Ein Type aus der Annotation von MEINEDGS mit allen seinen Tokens aus MEINEDGS: The Public DGS Corpus consists of more than 50 hours of video data from the DGS-Korpus project made available together wi
DataCite Commons2020-09-14 更新100
siimh/estonian_corpus_2021
该数据集包含两个版本的Estonian National Corpus 2021,分别是有形态学标注的文本(corpus_et.jsonl)和清理后的纯文本(corpus_et_clean.jsonl)。数据集总大小约为43GB,包含约1.96亿个句子、24亿个单词、1170万篇文档和6450万个段落。这些数据可以用于形态学分析、自然语言理解、语言模型微调等多种自然语言处理任务。
Hugging Face2024-10-25 更新70
BRECHEN2^ (MEINE DGS – annotiert. Öffentliches Korpus der Deutschen Gebärdensprache, 1. Release)
Ein Type aus der Annotation von MEINEDGS mit allen seinen Tokens aus MEINEDGS: The Public DGS Corpus consists of more than 50 hours of video data from the DGS-Korpus project made available together wi
DataCite Commons2020-09-14 更新60
Prague Dependency Treebank 3.5
This corpus is manually annotated at several levels – aside from syntactic parsing and morphological information, it is annotation for sentence information structure, multiword expression, coreference
SSH Open MarketPlace2023-10-13 更新80
AbNC: Abkhaz National Corpus
This corpus includes Abkhaz texts published between 1920 and 2016. The corpus is encoded in TEI. The corpus is available for online browsing through the Corpuscle concordancer (CLARINO distribution).
SSH Open MarketPlace2025-03-25 更新60



