From Voice To Shell: Speech To Docker Dataset
收藏资源简介:
Dataset with 3,192 audio files, 3.92 hours, comprising voice recordings of 12 subjects (4 women and 8 men) while enunciating the prompts of 'text-to-docker' dataset test samples in English (EN). Followed protocol has been approved by the Social Research Ethics Committee (SREC) of the University of Castilla-La Mancha (UCLM) under reference number CEIS-2025-119450. Human participation is justified and the benefits and risks have been adequately assessed to participants, developing a participation mechanism which warrants equal opportunities for research collaboration. Every participant has been provided an informed consent document including the necessary information of research purposes and requirements. It also complies with the regulations in force regarding the personal data protection. Audio files are encrypted, which could be extracted using the decrypt.py script: user@host:#~/data/$ lsen decrypt.py metadatauser@host:#~/data/$ pip install cryptography[... install dependencies ...]user@host:#~/data/$ python3 decrypt.py -p <password> .[... decrypt files ...]
本数据集包含3192条音频文件,总时长3.92小时,由12名受试对象(4名女性、8名男性)朗读英文(EN)的`text-to-docker`数据集测试样本提示文本录制而成。本研究采用的实验流程已通过卡斯蒂利亚-拉曼恰大学(University of Castilla-La Mancha, UCLM)社会研究伦理委员会(Social Research Ethics Committee, SREC)审批,审批编号为CEIS-2025-119450。本次人类受试者研究具备合理性,已针对每位受试者充分评估研究的收益与风险,并建立了保障研究合作机会均等的参与机制。所有受试者均已获取并签署知情同意书,文件中包含研究目的与相关要求的全部必要信息。本数据集同时符合现行个人数据保护相关法规要求。 音频文件均已加密,可通过`decrypt.py`脚本完成解密操作: user@host:#~/data/$ ls decrypt.py metadata user@host:#~/data/$ pip install cryptography[... 安装依赖项 ...] user@host:#~/data/$ python3 decrypt.py -p <密码> .[... 解密文件 ...]



