遇见数据集

Automotive-Diagnostic Case Study for "Automating Corporate Knowledge Mapping: A Dual-Domain Semantic Pipeline for Warm-Start Interviews"

收藏
Zenodo2026-08-19 更新2026-08-20 收录
官方服务:

资源简介:

This upload contains only the raw experimental data underlying the automotive-diagnostic case study reported in Section 4 ("Evaluation") of "Automating Corporate Knowledge Mapping: A Dual-Domain Semantic Pipeline for Warm-Start Interviews." It does not include any analysis code — this is a data-only deposit. Access This record is restricted. See ACCESS_POLICY section below. Files File Description Structure data/assessment_graph_topology.json The expert-authored question graph $G=(V,E)$ referenced in Section 4 ("Initial expert structure") 45 nodes, 44 edges, 4 hierarchical depth levels. Each node has an id, a label, a depth, and a passport string. data/automotive_work_orders.csv The "Active trace" corpus (Table 4 of the paper): a single technician's own documented work 148 rows, columns order_id, technician_id, job_category, work_description. All rows belong to one technician (Tech_Diagnostic_01), matching the case study's "digital footprint of a particular technician." Every work_description is a unique, independently worded narrative. data/automotive_knowledge_corpus.csv Combined text corpus referenced in Table 4 648 rows, columns source, topic_id, text. source="External Context" (500 rows) stands in for the OEM manuals / repair-manual corpus; source="Active Trace" (148 rows) is a short excerpt of the corresponding automotive_work_orders.csvrow, tagged with a topic_id recording which of the four real thematic clusters it belongs to. Ground-truth facts you can check directly from these files, with no modeling required: the graph has exactly 45 nodes / 44 edges (a valid tree); node v_3 ("Mechanical Drive & Valvetrain Architecture") has exactly 30 descendant nodes; there are exactly 148 work orders, of which exactly 2 concern an EGR-burnout fault (job_category = "Electronic Systems & EGR"); everywork_description and every Active Trace / External Context text value in automotive_knowledge_corpus.csv is unique. Relationship between raw text and the paper's topic model Each Active Trace row's topic_id records which of the four real thematic clusters that work order belongs to (fuel_delivery, electronic_obd, mechanical_timing, egr_burnout), matching the topic proportions theta_m used in Eq. 7 of the paper. The individual row text is a natural, independently worded technician narrative rather than a fixed template — analogous to how BERTopic itself separates a cluster's representative keyword profile (Eq. 1, c-TF-IDF) from the individual documents assigned to that cluster. License / citation Please cite both the paper and this dataset if you use this data in published work. Redistribution of the raw files outside the scope of an approved access request is not permitted; seeACCESS_POLICY. Access Policy Who may request access Access may be granted to: Editors and peer reviewers handling the manuscript "Automating Corporate Knowledge Mapping: A Dual-Domain Semantic Pipeline for Warm-Start Interviews" during its review process at the journal/venue it is submitted to. Researchers affiliated with an accredited university, research institute, or comparable research organizationwho want to independently verify, replicate, or build on the results reported in that specific paper. Students or independent researchers without an institutional affiliation, only if their request is accompanied by a clear research purpose and, where applicable, the endorsement of a supervising academic (handled as an exception, not the default case). Access is not granted for: general public curiosity without a stated research purpose, commercial use or product development, journalistic requests, bulk/automated harvesting, or any request that does not identify a real, verifiable individual. What a request must include Every access request must contain: Full name and institutional affiliation (with an institutional email address wherever possible — personal email addresses require additional justification, see above). A short (2–4 sentence) description of the intended use — e.g. "verifying the reproducibility of Table 5" or "extending the pruning method to a different domain." Confirmation that the requester will not redistribute the raw files, and will cite both the paper and this dataset in any resulting publication. Review process Requests are reviewed personally by the corresponding author. Expect a response within 10 business days. Approval is at the sole discretion of the corresponding author; a denial does not require extensive justification, and decisions may be revisited if circumstances change (e.g. a reviewer role ends, or a requester later provides missing affiliation details). Approved access is granted for the specific stated purpose; a materially different future use requires a new request.

提供机构:
Zenodo
创建时间:
2026-08-19
二维码
社区交流群
二维码
科研交流群
商业服务