anonymous-654/LUCid
收藏资源简介:
LUCid(潜在用户上下文基准)是一个用于在更现实的相关性概念下评估终身个性化系统的数据集。与传统基准将相关性与语义相似性等同不同,LUCid引入了潜在用户上下文——这些信息与查询语义距离较远,但对生成正确的个性化响应至关重要。该基准旨在测试系统是否能从长交互历史中检索、推断和利用用户特定信号。数据集包含1,936个查询,交互历史长达500个会话(约620K词元),涵盖年龄组、位置/国家、宗教/文化、健康状况、领域归属和沟通风格等个性化维度。每个示例要求:1. 从历史中识别潜在用户上下文;2. 推断用户属性;3. 生成个性化响应。数据集提供多个基准变体,如LUCid-5(超短历史)、LUCid-10(短历史)、LUCid-C(受控重排)、LUCid-S(小规模评估)、LUCid-B(标准基准)和LUCid-L(长上下文压力测试),适用于不同实验设置。
LUCid (Latent User Context benchmark) is a dataset for evaluating lifelong personalization systems under a more realistic notion of relevance. Unlike traditional benchmarks that equate relevance with semantic similarity, LUCid introduces latent user context—information that is semantically distant from the query but crucial for generating the correct personalized response. The benchmark is designed to test whether systems can retrieve, infer, and utilize user-specific signals from long interaction histories. It includes 1,936 queries with interaction histories up to 500 sessions (~620K tokens), covering personalization dimensions such as age group, location/country, religion/culture, health conditions, domain affiliation, and communication style. Each example requires: 1. identifying latent user context from history, 2. inferring user attributes, and 3. generating a personalized response. The dataset offers multiple benchmark variants, including LUCid-5 (ultra-short history), LUCid-10 (short history), LUCid-C (controlled reranking), LUCid-S (small-scale evaluation), LUCid-B (standard benchmark), and LUCid-L (long-context stress test), suitable for different experimental settings.




