遇见数据集

CultureMarkers

收藏
魔搭社区2026-06-17 更新2026-08-02 收录
官方服务:

资源简介:

# The Culture Funnel: You Can't Align What isn't in the Data This dataset contains 5.6M culturally tagged samples as presented in the paper [The Culture Funnel: You Can't Align What isn't in the Data](https://huggingface.co/papers/2606.13808). ## Dataset Summary This dataset is designed to help researchers study and mitigate the "cultural data funnel" in Large Language Model (LLM) pipelines. We use a multidimensional tagging framework to identify cultural signals, domains, geographic locations, and task specialization across pretraining, fine-tuning, alignment, and reasoning datasets. Tags for cultural dimensions, domain, task intent, and geographic location are obtained via prompting open-weights [Command-A model](https://huggingface.co/papers/2504.00698), and language identification is obtained via [FastText langid](https://huggingface.co/papers/1612.03651). Although these annotations are automatically generated and may be imperfect, human evaluation shows they provide reliable signals for large-scale analysis. Our analysis shows that explicit cultural signals often diminish throughout post-training stages, resulting in a long-tail distribution of cultural representation across modern LLM datasets. We find adding these tags can be used to improve downstream cultural benchmark performance following the “treasure marking” approach proposed by [(D’souza et al., 2025)](https://huggingface.co/papers/2506.14702). ## Dataset Structure The dataset consists of over 5.6 million samples with the following features: - `inputs`: The raw text content. - `language`: The language of the text. - `culture_tag`: Tags identifying the cultural subdimension of the text (among knowledge, preference, dynamics, bias, general culture, and no culture). - `geolocation_tag`: Information regarding the geographic region associated with the data. - `domain_tag`: The domain the data belongs to (e.g., math, humanities, conversation, social sciences). - `task_intent_tag`: The intended purpose or task type of the sample. - `dataset_source`: The original dataset source from which the sample was extracted (among CulturaX, Dolci Instruct SFT, UltraFeedback, OpenThoughts, Aya Dataset, PRISM, ShareLM, CultureBank, MultiNRC, GeoFact-X) - `custom_id`: A unique identifier for each sample. See the table below for the full set of tags for each dimension <p align="left"> <img src="https://cdn-uploads.huggingface.co/production/uploads/697b96e1dba8c2adebb95ae7/rw8jo0IDsBoQm0_M3pccT.png" width="500"> </p> ## Citation If you use this dataset in your research, please cite: ```bibtex @article{culturefunnel2026, title={The Culture Funnel: You Can't Align What isn't in the Data}, author={Sahu, Ananya and Mofakhami, Mehrnaz and D'Souza, Daniel and Euyang, Thomas and Kreutzer, Julia and Fadaee, Marzieh}, journal={arXiv preprint arXiv:2606.13808}, year={2026} } ```

提供机构:
maas
创建时间:
2026-06-12
二维码
社区交流群
二维码
科研交流群
商业服务