遇见数据集

Cross-Register Authorship Attribution Corpus

收藏
Zenodo2021-09-16 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This corpus contains writings of eight authors known to have written in both vernacular and classical Chinese. The corpus has 4.2 million Chinese characters and can be useful in authorship identification research. The file README.md contains a full description of the data. All materials in this archive are in the public domain.

本语料库收录了8位同时擅长白话与文言写作的知名作家的作品。该语料库总计包含420万汉字,可应用于作者归属识别相关研究。README.md文件中包含该数据集的完整说明。本归档内的全部素材均属于公有领域。

提供机构:
Zenodo
创建时间:
2021-09-16
二维码
社区交流群
二维码
科研交流群
商业服务