遇见数据集

<b>CODE - </b><i>Comprehensive Oral mucosa Database with Explanations</i>

收藏
DataCite Commons2025-12-10 更新2026-02-09 收录
官方服务:

资源简介:

Comprehensive Oral Mucosa Database with Explanations (CODE) is an annotated image-based dataset developed to advance research, education, and computational modelling in the domain of oral mucosal health and disease. The dataset comprises high-quality intraoral images collected from participants representing a spectrum of oral mucosal conditions ranging from normal mucosa to Oral Potentially Malignant Disorders (OPMDs) and Oral Cancers. Each image is accompanied by extensive metadata and expert annotations, making CODE a unique, interpretable resource that can be utilised in the field of clinical and computational research.<b>Data Collection and Methodology</b>A total of 110 participants were prospectively recruited and intraoral images were captured using standardized mobile phone based intraoral photography protocols. Participants for this dataset were recruited from one of the spoke centers under a Hub-and-Spoke model established as part of an Indian Council of Medical Research (ICMR) funded study. In this model, Ragas Dental College and Hospital, Chennai, serves as the Hub, while several private Dental institutions and Non-Governmental Organizations (NGOs) across different regions of India function as Spokes, contributing to data collection and community-level oral screening. The present dataset represents the images and annotations collected from one such Spoke center following uniform methodological guidelines. All intraoral images were taken in accordance with a copyrighted Standard Operating Procedure (SOP) for intraoral photography, developed by the Hub institution to ensure standardization, reproducibility, and consistency of imaging practices across all participating Spokes. The data collection was approved by the Institutional Ethics Committee (Approval No: RIEC/20231021/PHD), and informed consent was obtained from all participants prior to participation, confirming full adherence to ethical standards, confidentiality protocols, and data protection policies in accordance with the Declaration of Helsinki.Before each imaging session, lenses were cleaned, and images were taken at an approximate distance of 4–5 cm from the oral cavity, primarily under natural light, with auxiliary lighting used when necessary. Retraction aids (mouth mirrors or wooden retractors) were employed for optimal visualization. Each participant’s intraoral region was systematically photographed, capturing eight standard sites namely dorsal tongue, ventral tongue, right buccal mucosa, left buccal mucosa, upper labial mucosa, lower labial mucosa, upper (maxillary) arch and lower (mandibular) arch. Image quality was assessed based on parameters such as centering, illumination, sharpness, and absence of motion blur or artifacts. Images not meeting these criteria were reacquired. The dataset is dynamic and will be updated periodically with additional participants and newly annotated images to enhance its diversity and representativeness over time.<br><b>Annotation and Regional attributes</b>All images underwent expert review and annotation by Oral Pathologist and Public Health Dentist specialists using the VGG Image Annotator (VIA) tool, version 3.0.13. The annotation process involved delineating Regions of Interest (ROI) using polygonal outlines to mark lesion boundaries, mucosal sub-sites, and other diagnostic features (e.g., surface texture, color variation, and border irregularity).Annotations were saved in JSON format, ensuring compatibility with deep learning frameworks. Each image file follows the structured naming convention:<br>A_B_C.jpeg, where:<br><br>“A” represents a unique anonymized patient ID,“B” indicates the site of data collection, and“C” corresponds to the intraoral region (e.g., DT for dorsal tongue, LB for left buccal mucosa).<br><b>Data Structure</b>CODE is organized into four main diagnostic categories:Normal Mucosa – healthy oral tissues without lesionsVariations from Normal – minor deviations from typical mucosaOral Potentially Malignant Disorders (OPMDs) – including Erythro-leukoplakia, Erythroplakia, Homogenous Leukoplakia, Lichen planus, Lichenoid lesion, Non-homogenous Leukoplakia, Oral submucous fibrosis, Proliferative verrucous Leukoplakia and Speckled leukoplakiaOral Cancer – histopathologically confirmed squamous cell carcinomaEach image is linked with its corresponding JSON annotation file and patient's metadata sheet in Excel format, containing demographic details (age, sex), habit history (tobacco, alcohol, areca nut use), and clinical diagnosis. All files are indexed using a unique anonymized identifier (SMITA-ID) to maintain traceability while ensuring full de-identification.<br><b>Ethical and Legal Compliance</b>The study and data collection were approved by the Institutional Ethics Committee of Ragas Dental College and Hospital, Chennai (Approval No: RIEC/20231021/PHD). Written informed consent was obtained from all participants before inclusion.<br>All data were collected, processed, and shared in accordance with the Declaration of Helsinki, institutional data protection policies, and national ethical guidelines. Personal identifiers were removed, and all images were anonymized before dataset compilation.<b>Potential Applications and Reuse</b>The CODE dataset serves as a comprehensive and reproducible benchmark resource for advancing research in:AI-based oral lesion detection, segmentation, and classificationExplainable machine learning and visual reasoning modelsStandardization of mobile-based intraoral image qualityClinical education and digital diagnostic trainingBy integrating expert annotations, explanatory insights, and structured metadata, CODE represents a first-of-its-kind interpretative oral mucosa dataset with detailed annotation aimed at bridging clinical expertise with computational innovation. The dataset will be continuously expanded to include additional cases, improving class balance and supporting robust AI model development for early detection of OPMD and oral cancers.

提供机构:
figshare
创建时间:
2025-11-06
二维码
社区交流群
二维码
科研交流群
商业服务