遇见数据集

INPERIA-Interview: A Synthetic Spanish Dataset of Penitentiary Pseudo-Interviews for Risk Factor Assessment

收藏
Zenodo2026-07-21 更新2026-08-01 收录
官方服务:

资源简介:

INPERIA-Interview is a synthetic Spanish text dataset composed of penitentiary pseudo-interviews designed to support exploratory natural language processing research on interview answers related to risk factors for temporary release permits. The dataset was developed in the context of the INPERIA academic project, focused on the use of artificial intelligence and natural language processing techniques for the analysis of penitentiary interview responses. The resource contains 55 fictional pseudo-interviews and 550 individual answers generated from fictional inmate profiles. The interviews are structured around ten risk-related factors inspired by the Spanish Tabla de Valoración del Riesgo (TVR): foreign status, drug dependence, criminal professionalisation, recidivism, breach of sentence, Article 10 of the Spanish General Penitentiary Organic Law, absence of previous prison permits, cohabitation deficiencies, distance from the release destination and internal pressures. The dataset is provided in long format, where each row corresponds to an individual synthetic answer linked to an interview identifier, split, question identifier and TVR-inspired factor. The package also includes predefined training and test partitions, a question-factor mapping file and documentation files describing the structure, intended uses, limitations and ethical considerations of the resource. A deliberately limited subset of 33 response-level records includes preliminary reference levels for three representative factors: foreign status, criminal professionalisation and distance from the release destination. These annotations are included in the test partition and are intended only for exploratory evaluation, prompt testing and preliminary comparison of language models. They should not be interpreted as expert-validated ground truth. No real penitentiary interviews, prison records, judicial files, clinical data, administrative documents, personal data or identifiable information from inmates are included. All profiles and answers are fictional and synthetically generated. The dataset is intended only for research, educational purposes, prompt engineering experiments, exploratory analysis, annotation protocol design and preliminary evaluation of language models. It must not be used to make, support or automate real penitentiary, judicial, administrative or clinical decisions. A data article describing the dataset, its construction process, limitations and potential uses is currently under preparation.

提供机构:
Zenodo
创建时间:
2026-07-08
二维码
社区交流群
二维码
科研交流群
商业服务