遇见数据集

SIT-Model

收藏
Zenodo2025-10-16 更新2026-05-26 收录
官方服务:

资源简介:

Syncretic Interpretive Transformer (SIT) Overview Syncretic Interpretive Transformer (SIT) is an advanced deep semantic modeling framework designed to analyze and interpret Chinese and Western cultural classics.It integrates multimodal encoding, tradition-specific interpretive layers, and a Cross-Civilizational Attention Module (CCAM), enhanced with a Dialectical Alignment Protocol (DAP) to align and contrast philosophical texts across civilizations while preserving their unique cultural logic. SIT is particularly effective for studying Confucian, Taoist, and Buddhist classics as well as ancient Greek and Roman philosophical texts and Renaissance literature. It provides a foundation for cross-cultural semantic computing and digital humanities research. ✨ Features Multimodal Symbolic Modeling Integrates textual and image features (e.g., classical book scans, annotated versions). Employs a Spiking Layer + Frequency-Aware Mixer (see Fig. 1, p.7) to improve fine-grained semantic representation. Tradition-Sensitive Interpretation Two separate interpretive pathways: Chinese tradition: metaphor–analogy–resonance reasoning. Western tradition: deductive–axiomatic–structural reasoning. Cross-civilizational semantic alignment is performed through CCAM. Dialectical Alignment Protocol (DAP) Captures cultural–semantic tension dynamically. Treats divergence as interpretive richness rather than noise. See Fig. 3 (p.10) for the structural diagram. Cross-Cultural Contrastive Learning Uses contrastive objectives to align classical texts across traditions. Integrates knowledge graphs with transformer representations for accurate cross-cultural semantic mapping. 📊 Datasets Dataset Content Key Features Chinese Cultural Texts Semantic Dataset Confucian, Taoist, Buddhist classics, poetry, histories Rich in metaphors and cultural imagery Western Literary Classics Dataset Plato, Aristotle, Shakespeare, and others Includes original and modern translations Cross Cultural Semantic Analysis Dataset Aligned Chinese-Western pairs Annotated for themes, metaphors, value frames Multilingual Classics Classification Dataset 10+ languages Supports zero-shot and transfer learning Full dataset description is available on pages 12–13, including annotation schema and multilingual labeling strategy. ⚙️ Installation Clone the repository git clone https://zenodo.org/records/17365382 🚀 Usage Latent semantic vectors Dialectical Divergence Score (philosophical tension) Concept cluster labels and alignment mapping 🧪 Applications Cross-civilization philosophical thought analysis Symbolism and metaphor detection across traditions Digital humanities and cultural computing research Interpretability tools (attention heatmaps, divergence graphs) 🧩 Model Components Syncretic Interpretive Transformer (SIT) — dual-path interpretive architecture Cross-Civilizational Attention Module (CCAM) — semantic space alignment Epistemic Modulator — multi-layer interpretive weighting Dialectical Alignment Protocol (DAP) — cultural–semantic tension modeling Fusion Module — multimodal feature alignment and aggregation 📈 Performance Dataset Accuracy F1 Score AUC Chinese Cultural 92.76 90.91 94.02 Western Literary 91.65 89.93 93.10 Cross-Cultural 91.48 89.41 93.11 Multilingual Classics 90.62 88.67 92.38 Experimental results are reported on pages 14–16, Tables 1–2. 🧭 Future Work Extend support to Arabic, Sanskrit, and Hebrew classics Enhance adaptive symbolic encoding to reduce reliance on predefined vocabularies Develop interactive visualization interface for philosophical tension mapping Expand cross-domain transfer learning to contemporary thought texts 📜 License This project is licensed under the MIT License. 🙏 Acknowledgments This research is supported by the Education Department of Shaanxi Province (Project No. 21JK0108).Special thanks to the cross-cultural semantic computing and digital humanities research community.Author: Xinxin Wang, Shangluo University.

融合阐释Transformer(Syncretic Interpretive Transformer,简称SIT) ## 概述 融合阐释Transformer(SIT)是一款先进的深度语义建模框架,旨在分析与阐释中西经典文化典籍。该框架集成了多模态编码、专属传统阐释层以及跨文明注意力模块(Cross-Civilizational Attention Module,简称CCAM),并通过辩证对齐协议(Dialectical Alignment Protocol,简称DAP)进行增强,可在保留不同文明独特文化逻辑的前提下,对齐并对比跨文明哲学文本。 SIT尤其适用于研究儒、释、道经典,古希腊罗马哲学文本以及文艺复兴时期文学作品,可为跨文化语义计算与数字人文研究提供支撑。 ✨ 核心特性 ### 多模态符号建模 - 集成文本与图像特征(如经典典籍扫描件、带注释版本)。 - 采用脉冲层+频率感知混合器(详见第7页图1),以提升细粒度语义表征能力。 ### 传统敏感型阐释 包含两条独立阐释路径: - 中国传统路径:隐喻-类比-共振推理。 - 西方传统路径:演绎-公理-结构推理。 跨文明语义对齐通过CCAM实现。 ### 辩证对齐协议(DAP) - 可动态捕捉文化-语义张力。 - 将文本分歧视为阐释丰富性而非噪声。 结构示意图详见第10页图3。 ### 跨文化对比学习 - 采用对比目标函数实现不同传统下经典文本的对齐。 - 将知识图谱与Transformer表征相结合,实现精准的跨文化语义映射。 📊 数据集 | 数据集名称 | 数据集内容 | 核心特性 | | --- | --- | --- | | 中国文化典籍语义数据集 | 儒、释、道经典,诗词与史籍 | 富含隐喻与文化意象 | | 西方文学经典数据集 | 柏拉图、亚里士多德、莎士比亚等相关文本 | 包含原文与现代译本 | | 跨文化语义分析数据集 | 对齐后的中西文本对 | 针对主题、隐喻与价值框架进行标注 | | 多语种经典分类数据集 | 覆盖10余种语言 | 支持零样本学习与迁移学习 | 完整数据集说明详见第12-13页,包含标注框架与多语种标注策略。 ⚙️ 安装 克隆代码仓库: git clone https://zenodo.org/records/17365382 🚀 使用方法 - 潜在语义向量 - 辩证分歧得分(哲学张力) - 概念簇标签与对齐映射表 🧪 应用场景 - 跨文明哲学思想分析 - 跨传统符号与隐喻检测 - 数字人文与文化计算研究 - 可解释性工具(注意力热力图、分歧图谱) 🧩 模型组件 - 融合阐释Transformer(SIT):双路径阐释架构 - 跨文明注意力模块(CCAM):语义空间对齐 - 认知调制器:多层阐释权重分配 - 辩证对齐协议(DAP):文化-语义张力建模 - 融合模块:多模态特征对齐与聚合 📈 模型性能 | 数据集名称 | 准确率 | F1值 | AUC值 | | --- | --- | --- | --- | | 中国文化典籍数据集 | 92.76 | 90.91 | 94.02 | | 西方文学经典数据集 | 91.65 | 89.93 | 93.10 | | 跨文化语义分析数据集 | 91.48 | 89.41 | 93.11 | | 多语种经典分类数据集 | 90.62 | 88.67 | 92.38 | 实验结果详见第14-16页的表1至表2。 🧭 未来工作 - 扩展对阿拉伯语、梵语与希伯来语经典的支持 - 优化自适应符号编码,降低对预定义词汇表的依赖 - 开发面向哲学张力映射的交互式可视化界面 - 将跨域迁移学习扩展至当代思想文本 📜 许可证 本项目采用MIT许可证进行授权。 🙏 致谢 本研究受陕西省教育厅资助(项目编号:21JK0108)。特别感谢跨文化语义计算与数字人文研究社群。作者:王欣欣,商洛学院。

提供机构:
Zenodo
创建时间:
2025-10-16
二维码
社区交流群
二维码
科研交流群
商业服务