PhillyMac/Conflict_Resolution_Content_1
收藏资源简介:
--- license: cc0-1.0 task_categories: - text-generation - feature-extraction language: - en tags: - corpus - leadership - historical - deku-corpus-builder size_categories: - 1K<n<10K --- # Conflict Resolution Content 1 This corpus was automatically generated by the **Deku Corpus Builder** for use in RAG-based AI applications. ## Dataset Description - **Subject**: Conflict Resolution Leadership - **Subject Type**: topic - **Total Items**: 1,750 - **Items Requiring Attribution**: 0 - **Has Embeddings**: Yes (all-MiniLM-L6-v2) - **Created**: 2026-03-21 ## Dataset Structure Each record contains: - `text`: The content text - `source_url`: Original source URL - `source_title`: Title of the source document - `source_domain`: Domain of the source - `license_type`: License classification (e.g. `public_domain`, `cc_by`, `cc_by_sa`) - `attribution_required`: Boolean — True for CC BY / CC BY-SA and other attribution-required licenses - `attribution_text`: Formatted Creative Commons attribution string (empty if not required) - `license_url`: URL to the CC license deed (empty if not required) - `relevance_score`: Relevance to the subject (0-1) - `quality_score`: Content quality score (0-1) - `topics`: JSON array of detected topics - `character_count`: Length of the text - `subject_name`: The subject this content relates to - `subject_type`: "personality" or "topic" - `extraction_date`: When the content was extracted - `embedding`: Pre-computed 384-dimensional embedding vector ## Attribution 0 of 1,750 chunks in this corpus require attribution under their source license. When building lessons from these chunks, the `attribution_text` field must be surfaced in the lesson output per the Legend Leadership Attribution Tracking Spec. ## Usage ```python from datasets import load_dataset dataset = load_dataset("PhillyMac/Conflict_Resolution_Content_1") # Access attribution-required chunks for item in dataset["train"]: if item["attribution_required"]: print(item["attribution_text"]) ``` ## Integration with RAG This dataset is designed to be integrated with existing embedded corpuses. The embeddings use the `sentence-transformers/all-MiniLM-L6-v2` model, compatible with FAISS indexing. ## License Content is sourced from public domain and Creative Commons licensed materials. See individual `license_type` fields for per-chunk licensing details. ## Generated By [Deku Corpus Builder](https://github.com/PhillyMac/deku-corpus-builder) - An automated corpus building system for AI applications.
--- 许可证:CC0 1.0 任务类别: - 文本生成 - 特征提取 语言: - 英语 标签: - 语料库 - 领导力 - 历史 - deku-corpus-builder 规模类别: - 1000 < 条目数 < 10000 --- # 冲突解决内容集1 本语料库由**Deku语料库构建工具(Deku Corpus Builder)**自动生成,专为基于检索增强生成(Retrieval-Augmented Generation, RAG)的人工智能应用打造。 ## 数据集描述 - **主题**:冲突解决领导力 - **主题类型**:话题 - **总条目数**:1750 - **需标注来源的条目数**:0 - **已生成词嵌入**:是(使用all-MiniLM-L6-v2模型) - **创建时间**:2026年3月21日 ## 数据集结构 每条数据记录包含以下字段: - `"text"`:内容文本 - `"source_url"`:原始来源链接 - `"source_title"`:来源文档标题 - `"source_domain"`:来源域名 - `"license_type"`:许可证分类(例如`"public_domain"`、`"cc_by"`、`"cc_by_sa"`) - `"attribution_required"`:布尔值,若为`"CC BY"`/`"CC BY-SA"`等需标注来源的许可证则为真 - `"attribution_text"`:格式化的知识共享署名字符串(无需标注时为空) - `"license_url"`:知识共享许可证协议链接(无需标注时为空) - `"relevance_score"`:与主题的相关度评分(范围0-1) - `"quality_score"`:内容质量评分(范围0-1) - `"topics"`:检测到的话题的JSON数组 - `"character_count"`:文本字符数 - `"subject_name"`:该内容关联的主题名称 - `"subject_type"`:可选"人物"或"话题" - `"extraction_date"`:内容提取日期 - `"embedding"`:预计算的384维词嵌入向量 ## 署名要求 本语料库的1750个文本块中,无任何条目需要按照其来源许可证要求标注来源。在基于这些文本块构建课程内容时,需按照《Legend领导力署名追踪规范》(Legend Leadership Attribution Tracking Spec)在课程输出中展示`"attribution_text"`字段。 ## 使用方法 python from datasets import load_dataset dataset = load_dataset("PhillyMac/Conflict_Resolution_Content_1") # 筛选需标注来源的文本块 for item in dataset["train"]: if item["attribution_required"]: print(item["attribution_text"]) ## 与检索增强生成系统集成 本数据集专为与现有嵌入语料库集成而设计,其词嵌入采用`"sentence-transformers/all-MiniLM-L6-v2"`模型,兼容Facebook AI相似性搜索(Facebook AI Similarity Search, FAISS)索引构建。 ## 许可证 本数据集内容来源于公有领域及知识共享许可协议授权的素材,各文本块的具体许可信息请查阅每条数据的`"license_type"`字段。 ## 生成方 [Deku语料库构建工具(Deku Corpus Builder)](https://github.com/PhillyMac/deku-corpus-builder)——一款面向人工智能应用的自动化语料库构建系统。



