遇见数据集

INPERIA-Interview: A Synthetic Spanish Dataset of Penitentiary Pseudo-Interviews for Risk Factor Assessment

收藏
Zenodo2026-06-09 更新2026-06-12 收录
官方服务:

资源简介:

INPERIA-Interview is a synthetic Spanish dataset composed of penitentiary pseudo-interviews designed for experimental research on automatic analysis of responses associated with risk assessment factors. The dataset was developed in the context of the INPERIA academic project, focused on the use of artificial intelligence and natural language processing techniques to support penitentiary interview analysis. The resource contains synthetic responses generated from fictional profiles and structured around ten risk-related factors inspired by the Tabla de Valoración del Riesgo (TVR): extranjería, drogodependencia, profesionalidad delictiva, reincidencia, quebrantamiento, artículo 10 LOGP, ausencia de permisos, deficiencia convivencial, lejanía and presiones internas. The dataset is provided in long format, where each row corresponds to an individual answer linked to an interview identifier, split, question identifier and TVR factor. The package also includes the train/test partitions, the question-factor mapping file and documentation files describing the structure, intended uses, limitations and ethical considerations of the dataset. No real penitentiary interviews, personal data or identifiable information from inmates are included. All profiles and answers are fictional and synthetically generated. The dataset is intended only for research, educational purposes, prompt testing, exploratory analysis and preliminary evaluation of language models. It must not be used to make or support real penitentiary, judicial, administrative or clinical decisions. A paper describing the dataset and its preliminary evaluation may be added in a future version of this record.

提供机构:
Zenodo
创建时间:
2026-06-09
二维码
社区交流群
二维码
科研交流群
商业服务