Sri Lankan Classified Ads Dataset for Ad Matching and Retrieval
收藏资源简介:
This dataset was developed to support research in ad matching, semantic retrieval, and intent alignment across offering and wanted ads in Sri Lankan classified marketplaces. It consists of 54,489 ad pairs, sourced from five major platforms—ikman.lk, patpat.lk, Riyasewana, adz.lk, and Hitad.lk—and includes both human-verified real and LLM-generated samples. The dataset is particularly valuable for training and evaluating machine learning models that require generalization across low-resource subcategories, especially where wanted ads are underrepresented. Each row contains: Full offering and wanted ad texts Titles and descriptions of both ad types Main and subcategory labels The dataset covers three main categories: electronics (37.17%), vehicles (33.25%), and property (29.58%), with further breakdown into 20 subcategories, including cars, vans, houses, land, mobile phones, and more. Language distribution: English (60.34%) Mixed Sinhala/English or Tamil/English (34.46%) Sinhala (5.20%) Tamil (0.01%) Construction details: Human-matched ad pairs, double-annotated and adjudicated. Synthetic matching wanted ads for existing real offering ads generated using Gemini 2.0 Flash (LLM) for rare subcategories, with structured prompt engineering. All ads were collected from publicly accessible online marketplaces, and personally identifiable information (such as phone numbers) was replaced by placeholders through preprocessing. The dataset is anonymized and suitable for academic and non-commercial use.



