遇见数据集

CEFR-level expert and large language model annotations of English lexical units across ten texts

收藏
Zenodo2026-06-09 更新2026-06-12 收录
官方服务:

资源简介:

This dataset accompanies the study "Comparison of language models in the CEFR-level assessment of English lexical units" and contains: (1) A corpus of ten English text excerpts spanning CEFR levels A2–C2 across four functional registers (graded news, popular science, literary essay/nonfiction, legal/academic); (2) Expert annotations of every content word and multi-word expression with CEFR levels following the English Vocabulary Profile (Cambridge University Press), totalling 1085 lexical units; (3) Parallel annotations of the same corpus produced by seven contemporary large language models: DeepSeek V4 Pro, Kimi K2.6, Llama 3.3 70B Instruct, Qwen 3.7 Max, ChatGPT 5.5, Claude Sonnet 4.6, and Gemma 4 31B; (4) The Python analysis pipeline (compare_cefr_models.py) computing token-selection precision/recall/F1, lemma and POS accuracy, CEFR classification accuracy, macro-F1, Cohen's quadratic-weighted kappa, MAE with 95% bootstrap confidence intervals, and pairwise McNemar tests between models; (5) The resulting tables (CSV) and figures (300 dpi, monochrome PNG). All annotations follow a unified JSON schema. Source attribution for every text excerpt - original URL, publication date, access date, and Wayback Machine snapshot where available - is provided in corpus/corpus_manifest.json. Text excerpts are short (135–296 words each) and used solely for research purposes under fair use; original copyright in each source article remains with its respective publisher. The annotations are released under CC BY 4.0.

提供机构:
Zenodo
创建时间:
2026-06-09
二维码
社区交流群
二维码
科研交流群
商业服务