遇见数据集

Hammer-Purgstall Correspondence TEI Evaluation Dataset

收藏
Zenodo2025-12-22 更新2026-05-26 收录
官方服务:

资源简介:

This dataset supports LLM-supported TEI encoding tasks, output evaluation and comparison. It comprises a representative sample of 100 letters from the correspondence of Joseph von Hammer-Purgstall (1774-1856), an Austrian orientalist, historian, and diplomat. The correspondence spans six decades (1790s-1850s) and exhibits significant linguistic diversity, containing letters primarily in German alongside English, French, and Italian, with instances of code-switching (Latin, Greek, and Arabic text segments). The dataset includes four main components for LLM processing:1. Input component: 100 plain text letter transcriptions (UTF-8)2. Reference component: 100 manually encoded TEI XML (P5) reference annotations following TEI Guidelines for correspondence3. Output component: LLM-generated TEI XML encodings from four models (GPT-5-mini, Claude Sonnet 4.5, Qwen3-14B-Q6, OLMo2-32B-instruct-Q4), totaling 400 encodings4. Evaluation component: Comprehensive assessment results in JSON format with aggregate Excel reports and a visualization of the cross-model comparison The sample was selected through systematic stratified sampling based on language distribution, writer diversity, and letter length variation. Files included:- hpe-correspondence-metadata.json: Metadata for 100 letters (language, sender, recipient, date, letter length, etc.)- hpe-correspondence-transcriptions.zip: Plain text letter transcriptions- hpe-correspondence-tei-reference.zip: Manually encoded TEI XML reference files- hpe-correspondence-llm-encodings.zip: LLM-generated TEI XML encodings from four models- hpe-correspondence-evaluation.zip: Evaluation results in JSON format with Excel reports- hpe-prompt-templates.zip: Five prompt scenarios for LLM processing with encoding instructions and few-shot samples Code Availability The evaluation results in this dataset were generated using the TEI LLM Evaluation Framework. The source code is available at: https://github.com/strubrina/tei-evaluation.git

本数据集可支撑基于大语言模型(Large Language Model,LLM)的文本编码倡议(Text Encoding Initiative,TEI)编码任务,并用于相关输出的评估与对比。数据集包含奥地利东方学家、历史学家与外交官约瑟夫·冯·哈默-普尔格斯塔尔(Joseph von Hammer-Purgstall,1774-1856)往来书信中的代表性样本,共计100封书信。这批书信的时间跨度达六十年(1790年代至1850年代),语言多样性显著:主体为德语,同时包含英语、法语与意大利语书信,还存在语码转换现象(如插入拉丁语、希腊语及阿拉伯语文本片段)。 数据集包含四大模块以适配大语言模型处理需求: 1. 输入模块:100封书信的纯文本转录稿(UTF-8编码) 2. 参考模块:100份按照《TEI书信编码指南》手动标注的TEI XML(P5标准)参考注释 3. 输出模块:由四款大语言模型生成的TEI XML编码结果,分别为GPT-5-mini、Claude Sonnet 4.5、Qwen3-14B-Q6及OLMo2-32B-instruct-Q4,总计400份编码文件 4. 评估模块:包含综合评估结果的JSON格式文件、汇总型Excel报表,以及跨模型对比可视化内容 该样本通过系统分层抽样方法选取,抽样依据涵盖语言分布、作者多样性及书信长度差异。 数据集包含以下文件: - "hpe-correspondence-metadata.json":100封书信的元数据(语言、发件人、收件人、日期、书信长度等) - "hpe-correspondence-transcriptions.zip":书信纯文本转录稿压缩包 - "hpe-correspondence-tei-reference.zip":手动标注的TEI XML参考文件压缩包 - "hpe-correspondence-llm-encodings.zip":四款大语言模型生成的TEI XML编码结果压缩包 - "hpe-correspondence-evaluation.zip":包含JSON格式评估结果与Excel报表的压缩包 - "hpe-prompt-templates.zip":适用于大语言模型处理的五类提示模板压缩包,内含编码指令与少样本示例 代码可用性 本数据集的评估结果基于TEI大语言模型评估框架(TEI LLM Evaluation Framework)生成,其源代码可在以下地址获取:https://github.com/strubrina/tei-evaluation.git

提供机构:
Zenodo
创建时间:
2025-11-18
二维码
社区交流群
二维码
科研交流群
商业服务