遇见数据集

ybracke/lexicon-dtak-transnormer-v1

收藏
Hugging Face2025-03-11 更新2025-04-12 收录
官方服务:

资源简介:

Lexicon-DTAK-transnormer数据集是从dtak-transnormer-full-v1平行语料库派生出的词汇表,包含1600年至1899年期间德国文本的原始拼写ngram和正规化拼写ngram的对齐及其频率信息。该数据集分为三个子数据集,每个子数据集对应一个世纪的数据。每对ngram包括原始拼写ngram、正规化拼写ngram、总频率和出现在多少文档中。

The Lexicon-DTAK-transnormer dataset is a lexicon derived from the dtak-transnormer-full-v1 parallel corpus, containing ngram alignments between original spelling ngrams and normalized spelling ngrams observed in German texts from 1600 to 1899 along with their frequency information. The dataset is divided into three sub-datasets, each corresponding to a centurys data. Each ngram pair includes the original spelling ngram, normalized spelling ngram, total frequency, and the number of documents in which it occurs.

提供机构:
ybracke
二维码
社区交流群
二维码
科研交流群
商业服务