An LLM-driven Chinese Corpus of Human Olfactory Descriptions and Entity Annotations
收藏资源简介:
Language plays a pivotal role in artificial olfactory perception, serving as a crucial bridge that translates chemical stimuli into human experience. The mapping from olfactory stimuli to linguistic descriptions constitutes the foundational basis for modeling artificial olfaction with human-like descriptive capabilities. However, this field is currently constrained by a lack of publicly available, culturally diverse olfactory corpora, with existing datasets predominantly reflecting Western odor profiles. To address this gap, we constructed a large-scale Chinese olfactory corpus by employing a dual iterative strategy with the assistance of a large language model (LLM). Beginning with an initial set of standardized descriptors and a defined annotation template, our strategy iteratively alternated between lexicon expansion and corpus refinement throughout the process. The resulting large-scale corpus, comprising highly domain-relevant sentences, enables the training of computational language model. The model achieves a more nuanced and perceptually grounded semantic mapping of the Chinese olfactory experience. Consequently, our work provides an essential foundational resource for advancing culturally inclusive models of olfactory perception.
语言在人工嗅觉感知中扮演着核心枢纽角色,是连接化学刺激与人类嗅觉体验的关键桥梁。将嗅觉刺激映射为语言描述,是打造具备类人描述能力的人工嗅觉模型的核心基础。然而当前该领域仍存在明显局限:缺乏公开可用、具备文化多样性的嗅觉语料库,现有数据集大多仅反映西方气味特征。为填补这一研究空白,我们采用双迭代策略,并借助大语言模型(Large Language Model,LLM),构建了大规模中文嗅觉语料库。本研究以标准化初始描述词集与预设标注模板为起点,在整个流程中交替开展词汇表扩展与语料库精炼。最终构建的大规模语料库包含大量与领域高度相关的语句,可用于训练计算语言模型,该模型能够实现更为精细且贴合感知实际的中文嗅觉体验语义映射。因此,本研究为打造具备文化包容性的嗅觉感知模型提供了至关重要的基础资源。



