European Nucleotide Archive: qualitative interviews with data practitioners
收藏资源简介:
Abstract This study consists of qualitative interviews about nucleotide sequence data curation, submission, reuse, and infrastructural labour within the European Nucleotide Archive (ENA). Interview partners include data submitters, curators, bioinformaticians, and infrastructure staff who generate, process, or steward ENA records. Some of the main topics covered include the organisation of sequencing and data production workflows; the history and institutional embedding of ENA within the International Nucleotide Sequence Database Collaboration; the perceived value of open nucleotide data and public archives; challenges involved in preparing, annotating, and submitting sequence data and metadata; opportunities and limitations of reusing ENA records across research contexts; practices of data validation, standardisation, and quality control within the archive; and the role of ENA data in downstream analysis, comparative genomics, surveillance, and broader life science research infrastructures. This study is part of the project A Philosophy of Open Science for Diverse Research Environments (PHIL_OS). Methods The data in this study was collected using semi-structured qualitative interviews. Interview partners were recruited by snowball sampling through their engagement with ENA. 23 interviews were conducted in total. Interviews were conducted between September 2023 and November 2023. The interviews took place at EMBL's European Bioinformatics Institute (EMBL-EBI), Hinxton, UK or online using Zoom videoconferencing software. Interviews lasted 18-90 minutes. When participants provided their written consent, interviews were audio-recorded and transcribed smart verbatim using otter.ai and manual proofreading. Sensitive information was removed before publishing transcripts. Transcripts were analysed using thematic coding. Codes were organised into parent codes using an inductive approach based on emergent categories. Description of the data and file structure Documentation files include interview guides, the information sheet and consent form, ethics approval, and the data narrative. Documentation files are named according to the structure: authorname_filename_DOCUMENTATION. These identifiers recorded three elements: the case study (withe.g.“ENA_” ), the professional role of the interviewee (e.g., Data Technician, DT; Data Educator, DE; Data Manager, DM; Data Director, DD; or Data Curator, DC; see supplementary material 3 for details on full roles), and finally a sequential number corresponding to the order of interviews. A full list of files is provided in the README file. Notes This study was conducted as part of the project A Philosophy of Open Science for Diverse Research Environments (PHIL_OS). More information can be found at https://opensciencestudies.eu This project has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (grant agreement No. 101001145). The author N.S. was funded via a doctoral training grant awarded as part of the UKRI AI Centre for Doctoral Training in Environmental Intelligence (UKRI grant number EP/S022074/1).



