遇见数据集

Formalizing Narratives for Spatial Context Extraction in GeoA

收藏
Zenodo2026-07-28 更新2026-08-13 收录
官方服务:

资源简介:

Minimal Reproducibility Package for Formalizing Narratives for Spatial Context Extraction in GeoAI This Zenodo record contains the minimal reproducibility package associated with the following manuscript: Menšík, M., Číhalová, M., Rapant, P., Duží, M., Žáček, M., and Albert, A.Formalizing Narratives for Spatial Context Extraction in GeoAI2026 The manuscript has been submitted for publication and is currently under review. The package implements the text-processing pipeline described in the manuscript. It processes English sentences describing movement using spaCy and a set of custom rules for identifying the values of the directional valency functors DIR1, DIR2, DIR3, and MANN. The identified information is transformed into a structured, Prolog-oriented representation intended for subsequent computational processing. The package also includes a simple web interface implemented using Flask. Package Contents The archive contains: the complete Python source code required to run the processing pipeline; a Flask web application providing a graphical interface to the pipeline; the rule-based implementation for identifying the values of DIR1, DIR2, DIR3, and MANN; plain-text files containing independent example sentences; generated visually formatted outputs; generated structured outputs; project documentation in README.md; dependency specifications in requirements.txt. Each example input file contains one English sentence per line. Every sentence describes a single main movement event and may explicitly specify its origin, path, destination, or relative direction. The sentences are processed independently and do not form a continuous narrative. Reproducibility Scope The supplied example sentences are designed so that the expected values of DIR1, DIR2, DIR3, and MANN can be identified directly from their meaning. The package can be used to: process the supplied example sentence files using the pipeline described in the associated manuscript; generate a visually formatted representation of the identified functor values; generate structured output intended for subsequent computational processing; process additional English sentence files that follow the supported input format; inspect which extraction rule produced each identified value. The supplied text files serve as demonstration inputs for the implemented processing workflow. They do not constitute an annotated evaluation dataset, a quantitative benchmark, or a continuous narrative dataset. Input and Output Input files must be UTF-8-encoded plain-text files containing exactly one complete English sentence per line. The files contain no header or additional metadata, and empty lines between sentences should be avoided. For each sentence, the application generates four records corresponding to DIR1, DIR2, DIR3, and MANN. Each record contains: the numerical identifier of the narrative; the numerical identifier of the sentence; the identifier of the rule that produced the result; the functor type; the identified functor value; the associated motion verb. Functors for which no value was identified are omitted. The output is a structured, TIL Script representation. Reproducing the Processing Workflow The accompanying README.md provides detailed instructions for: creating and activating a Python virtual environment; installing the required dependencies; installing the required spaCy language model; starting the Flask web application; processing the supplied example files; processing additional files in the supported format; inspecting and downloading the generated outputs. Following these instructions enables independent researchers to run and inspect the complete sentence-processing workflow implemented in the package. Technical Environment The package was developed and tested using: Python: 3.12.6 Flask: 3.1.3 spaCy: 3.8.14 spaCy language model: en_core_web_trf 3.8.0 Operating system: Windows 11 All additional Python dependencies are listed in requirements.txt. Known Limitations The implementation is intended for English sentences describing a single main movement event with a limited amount of explicitly expressed spatial information. Each input line is processed independently. The application does not use information from preceding or following sentences and does not perform cross-sentence coreference resolution. The quality of the generated output depends on the tokenization, lemmatization, part-of-speech tagging, and dependency analysis produced by spaCy. Sentences containing unusual word order, implicit spatial information, multiple movement events, or syntactic ambiguity may not be processed correctly. The application does not resolve identified place names against an external geographic database. Data Availability The archive includes the example input files and the corresponding generated outputs used to demonstrate the processing workflow. The example sentences do not represent a complete experimental dataset and are not intended for quantitative evaluation of the method. License The source code included in this reproducibility package is distributed under the MIT License. The example input sentences and the accompanying documentation are distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Third-party software and language models used by the application, including spaCy, Flask, and en_core_web_trf, remain subject to their respective licenses and are not relicensed as part of this package. Purpose This reproducibility package enables independent researchers to inspect, run, verify, and reuse the implementation of the processing pipeline described in the associated manuscript. The archived materials support methodological transparency, verification of the implemented extraction rules, long-term preservation of the source code, and further research on the extraction of spatial information from natural-language descriptions. Contact For questions concerning the source code or the processing workflow, contact: Adam AlbertDepartment of Computer Science, FEECSVSB – Technical University of OstravaEmail: adam.albert@vsb.cz Funding Work is partially supported by SP2026/056 Application of Formal Methods in Knowledge Modelling and Software Engineering IX, VŠB - Technical University of Ostrava, Czech Republic.

提供机构:
Zenodo
创建时间:
2026-07-28
二维码
社区交流群
二维码
科研交流群
商业服务