遇见数据集

SMART (Systematic Oral Mucosal Annotated Images for Research and Training) Intraoral images Dataset

收藏
Zenodo2025-08-02 更新2026-05-26 收录
官方服务:

资源简介:

The SMART (Systematic Oral Mucosal Annotated Images for Research and Training) dataset was developed under the Indian Council of Medical Research (ICMR)-funded project titled "Development of a Simple Mobile Interface Technology Application (SMITA) for the Screening of Oral Disorders Among High-Risk Rural Populations in Tamil Nadu". This initiative aims to facilitate advancements in mobile-assisted oral health screening and the development of artificial intelligence (AI) tools within the field of Dentistry. The dataset comprises 1,878 high-resolution intraoral images obtained from 235 patients enrolled through multiple spoke centers participating in the SMITA project. Data collection is ongoing, and the dataset will be updated on a monthly basis to expand its coverage and utility. For each subject, a standardized image acquisition protocol was employed to capture eight anatomically distinct intraoral sites: dorsal tongue, ventral tongue, right and left buccal mucosa, upper and lower lips, and upper and lower dental arches. These photographs were acquired using mobile phone cameras operated by Community Health Workers (CHW) under real-world, community-based conditions, thereby enhancing the dataset’s applicability in low-resource settings. The protocol was based on a training manual specifically developed and copyrighted under the project. Images were categorized into four clinically defined diagnostic groups: Normal, Variations from normal, Oral Potentially Malignant Disorders (OPMDs) and Oral Cancer (OC). All images underwent detailed annotation using the VGG Image Annotator (VIA), employing region-specific labelling guided by a customized set of clinically validated descriptors to ensure annotation consistency and diagnostic accuracy. Further, patient metadata which includes demographic details, personal history, habit history such as usage of tobacco, betel nut and alcohol consumption and clinical findings has been recorded and included in the dataset in CSV format. The annotated intraoral images along with patient metadata, can be leveraged to develop an AI-driven algorithm for risk stratification and predictive modeling. The SMART dataset is designed to support a wide range of applications including the training and validation of AI models for automated diagnostic decision-making and real-time image based triaging. All data collection procedures were conducted in accordance with institutional ethics committee approvals, with strict adherence to data privacy, participant confidentiality, and responsible data sharing protocols.

提供机构:
Zenodo
创建时间:
2025-08-02
二维码
社区交流群
二维码
科研交流群
商业服务