遇见数据集

From Character to Poem: Nested Contexts and Scalar Limits of Parallelism in Classical Chinese Poetry

收藏
Zenodo2026-01-11 更新2026-05-26 收录
官方服务:

资源简介:

This dataset accompanies the manuscript "From Character to Poem: Nested Contexts and Scalar Limits of Parallelism in Classical Chinese Poetry." It provides resources for benchmarking parallelism detection in Chinese classical regulated verse (律诗). The core data consists of silver_standard_train.json (80 MB, ~60K poems) and silver_standard_test.json (842 KB) containing pentasyllabic regulated poems from the Tang (唐), Song (宋), Yuan (元), Ming (明), and Qing (清) dynasties. Each poem entry includes: dynasty—the historical period; couplets—an array of four couplets, each containing two five-character lines; char_match—character-level binary parallelism labels indicating whether each of the five corresponding character positions in a couplet pair exhibits semantic/grammatical correspondence (as determined by the shared community membership); line_match—couplet-level binary labels indicating whether each couplet as a whole is parallel, as determined by the SikuBERT-based binary classifier; confidence_scores—model confidence values for each couplet's parallelism classification. The char_communities.json file provides a character-to-cluster mapping that groups ~6,700 Chinese characters into semantic communities used for generating silver-standard annotations. The evaluation_results.json file contains benchmark results from 100 experimental trials across multiple model architectures (character-level, couplet-level, and poem-level classifiers) with accuracy, precision, recall, and F1 metrics. The source corpus penta_4c_poems.csv includes the raw poems with metadata (dynasty, title, author, form type).

提供机构:
Zenodo
创建时间:
2026-01-11
二维码
社区交流群
二维码
科研交流群
商业服务