OpenFCS — A Knowledge Graph of the 2025 UNESCO Framework for Cultural Statistics: Extraction Artefacts and a Controlled Comparison of Anchoring Strategies
收藏资源简介:
Artefacts of a knowledge-graph extraction experiment on the 2025 UNESCO Framework for Cultural Statistics (Parts I and II). Three knowledge graphs (GraphML, Turtle, JSON-LD) produced under free, anchored and anchors-only extraction from the same 198-page corpus with a local open-weights model; the seed vocabularies; the blind human-validation spreadsheets and their pre-registered samples; extraction diagnostics; and the scripts that reproduce the pipeline from OCR to graph. Main finding: anchoring LLM extraction in a controlled vocabulary is a trade-off rather than a free improvement. It raises the largest connected component from 25% to 70% of concepts and reduces the graph diameter from 13 to 9, while reducing precision from 54% to 35% (n=120 per arm, p=0.003). The cost is independent of vocabulary size: 20 and 52 seed concepts yielded identical precision. Inter-annotator agreement in the blind validation: Cohen's kappa 0.932 and 0.966. Hallucination rate (triples whose cited evidence is absent from the page): 0.00% (free) and 0.39% (anchored). Source document: 2025 UNESCO Framework for Cultural Statistics, UNESCO Institute for Statistics (https://uis.unesco.org).



