遇见数据集

The INCOME dataset (INdonesian COMmerce and sEntiment)

收藏
Zenodo2025-08-14 更新2026-06-05 收录
官方服务:

资源简介:

Abstract The INCOME dataset (INdonesian COMmerce and sEntiment) contains 1,000 unique entries comprising e-commerce transaction records and social media posts related to digital consumption, sustainability, and support for Indonesian Micro, Small, and Medium Enterprises (MSMEs). The dataset is compiled from two primary sources: transaction logs voluntarily provided by selected Indonesian online marketplaces and publicly available user-generated content from Twitter and Instagram, collected through targeted manual searches over a defined observation period. The transaction data (500 entries) captures product category, purchase frequency, and seller origin (local or imported), enabling analysis of consumer behavior and MSME visibility trends. Product names and marketplace identifiers are deliberately omitted to avoid any form of endorsement or marketing. The social media data (500 entries) consists of posts manually identified and documented using hashtags such as #BeliLokal, #UMKM, and #BanggaBuatanIndonesia, enriched with authentic Indonesian slang expressions and informal writing styles common on these platforms. All data was recorded, cleaned, and validated through a manual documentation process, with no automated scraping tools used. Both original and preprocessed formats are provided, supporting research in sentiment-informed topic modeling, consumer behavior analysis, and system dynamics modeling for policy simulation. Dataset Overview The INCOME dataset integrates behavioral (transaction) and perceptual (social media) data to examine how public sentiment influences local MSME participation in Indonesia’s digital economy. It is designed to support research that combines Natural Language Processing (NLP) with system dynamics modeling for policy and platform strategy evaluation. Collection was performed through a combination of direct provision from marketplace partners and manual review of public social media timelines. No automated scraping tools were used; entries were documented and cross-checked by hand to preserve accuracy, privacy, and cultural context. In total, the dataset contains 1,000 unique entries: 500 e-commerce transaction records (250 original, 250 preprocessed) 500 social media posts (250 original, 250 preprocessed) Contents /transaction_original/ Contains the original e-commerce transaction logs, including product metadata, timestamps, purchase counts, and seller origin. Maintains a ~90% Imported vs ~10% Local distribution consistent with observed market dominance of imported products. Excludes actual product names and any references to marketplace-specific product origins. /transaction_preprocessed/ Contains cleaned and standardized transaction data with normalized category names, date formats, and seller origin labels. /socialmedia_original/ Contains raw Twitter and Instagram posts (text only) with hashtags, metadata, and Indonesian slang phrases reflective of real online discourse, including expressive punctuation, emojis, and local vernacular. /socialmedia_preprocessed/ Contains normalized, anonymized, and tokenized versions of the social media text. Preprocessing removes emojis, strips hashtags, and preserves slang in lowercase format for analysis. File Details /transaction_original/ • Total size: ~16 KB • Rows: 250 • Format: CSV /transaction_preprocessed/ • Total size: ~16 KB • Rows: 250 • Format: CSV /socialmedia_original/ • Total size: ~19 KB • Rows: 250 • Format: CSV /socialmedia_preprocessed/ • Total size: ~18 KB • Rows: 250 • Format: CSV Dataset Usage Research Applications This dataset is ideal for: Sentiment-informed topic modeling in Indonesian, including slang and informal speech. Consumer behavior analysis in e-commerce ecosystems. MSME visibility studies in platform algorithms. Policy simulation using system dynamics. Loading and Accessing the Data Transaction data can be loaded using pandas or other tabular data libraries. Social media text can be processed using NLP libraries such as nltk, spaCy, or transformers. Preprocessing Notes Transaction Data: Standardized date formats, normalized category names, and unified seller origin labels. Deduplication of repeated transaction entries. Exclusion of marketplace-specific product names and product origin references. Social Media Data: Public posts were located by searching for specific hashtags and keywords during the observation period. Text content was copied manually from each relevant post, ensuring accurate capture of slang, repetition, and punctuation. Usernames and identifiable metadata were removed by hand to maintain privacy. Hashtags and emojis were removed only in the processed version, while the original version preserves the text in its authentic form.

提供机构:
Zenodo
创建时间:
2025-08-14
二维码
社区交流群
二维码
科研交流群
商业服务