遇见数据集

Edelweiss Hospital Google Review Dataset

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Dataset Overview This collection consists of 1,445 patient reviews extracted from Edelweiss Hospital’s Google Maps page. It serves as the core evidence for the sentiment analysis and operational strategy research. 1. RAW Data (- RAW.csv) This file represents the original, unprocessed extraction from the web. Content: Includes the hospital title, URL, star ratings (1-5), reviewer names, and the authentic text of the reviews. Characteristics: It captures the "noise" of online electronic Word-of-Mouth (eWOM), featuring informal Indonesian language, local slang, emojis, and special characters. It is the primary source used for initial data exploration. 2. CLEAN Data (- CLEAN.csv) This file is the refined and annotated version optimized for Natural Language Processing (NLP). Content: Features a structured layout with unique index numbers (No) and, most importantly, a sentiment_label column. Characteristics: The text has been cleaned to remove noise and labeled as "positive" or "negative." This dataset is the ready-to-use input for training the IndoBERTweet model and performing topic modeling with Latent Dirichlet Allocation (LDA).

数据集概览 本数据集包含从雪绒花医院(Edelweiss Hospital)谷歌地图页面提取的1445条患者评论,可作为情感分析与运营策略研究的核心佐证材料。 1. 原始数据(- RAW.csv) 该文件为未经处理的原始网页提取数据。 内容:包含医院名称、网页链接、1至5星评分、评论者姓名及评论原文。 特征:保留了在线电子口碑(electronic Word-of-Mouth, eWOM)的“噪音”,文本采用非正式印尼语、本地俚语,包含表情符号与特殊字符,是初始数据探索的主要数据源。 2. 清洗后数据(- CLEAN.csv) 该文件为经优化处理并标注的版本,适配自然语言处理(Natural Language Processing, NLP)任务。 内容:采用结构化布局,包含唯一索引编号(No),尤为重要的是设有sentiment_label列。 特征:文本已完成降噪清洗,并被标注为“积极”或“消极”两类。本数据集可直接作为训练IndoBERTweet模型及使用潜在狄利克雷分配(Latent Dirichlet Allocation, LDA)进行主题建模的输入数据。

创建时间:
2026-01-29
二维码
社区交流群
二维码
科研交流群
商业服务