遇见数据集

Training corpus SUK 1.1

收藏
B2FIND2026-04-29 收录
官方服务:

资源简介:

The SUK training corpus contains about 1 million tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, and lemmatisation, with...

SUK训练语料库包含约100万个经人工标注的词元(token),其标注覆盖标记化(tokenisation)、分句(sentence segmentation)、形态句法标注(morphosyntactic tagging)以及词形还原(lemmatisation)等多个层面,且……

二维码
社区交流群
二维码
科研交流群
商业服务