seshat-perspective
收藏资源简介:
Seshat-perspective-eval是一个首个采用透视主义方法的历史数据库,利用多个大语言模型(Deepseek、Llama和Mistral)对历史文本进行合成标注,以提供多视角的历史分析。数据集适用于文本生成任务,语言为英文,样本数量在1千到1万之间,标签包括历史和文化分析。该数据集旨在支持历史与文化分析中的多视角文本生成研究。
Seshat-perspective-eval is a pioneering historical database using perspectivism, leveraging multiple large language models (Deepseek, Llama, and Mistral) to synthesize annotations on historical texts, providing multi-perspective historical analysis. The dataset is suitable for text generation tasks, is in English, has sample sizes between 1,000 and 10,000, and includes labels for historical and cultural analysis. It aims to support multi-perspective text generation research in historical and cultural analysis.
数据集概述:Seshat-perspective-eval
基本信息
- 数据集名称:Seshat-perspective-eval(又名 Seshat-perspective)
- 许可证:CC BY-NC-SA 4.0(知识共享-非商业性使用-相同方式共享4.0)
- 语言:英语(en)
- 数据类型:文本生成(text-generation)
- 数据规模:1,000 < n < 10,000 条样本
- 标签:历史(history)、文化分析(cultural_analytics)
数据集内容
Seshat-perspective 是首个采用视角主义方法进行合成标注的历史数据银行。该数据集通过多个大型语言模型(LLMs)生成标注,具体使用的模型包括:
- Deepseek(标注版本标记:dr1)
- Llama(标注版本标记:l31l)
- Mistral(标注版本标记:m3m)
相关资源
- 论文(含字段描述与验证流程):https://github.com/facells/fabio-celli-publications/blob/main/docs/2026_perspective_seshat_clicit26.pdf
- 复现代码:https://colab.research.google.com/drive/1_4aUNGjl7_uhLZZKE7mAYHPhWYUZ9jvr?usp=sharing





