MBTIBENCH
收藏资源简介:
MBTIBENCH是由哈尔滨工业大学等机构创建的第一个高质量MBTI性格检测数据集,旨在解决现有数据集中自我报告标签不准确和缺乏软标签的问题。数据集通过心理学家的指导进行手动重新标注,包含286条样本,涵盖了四种MBTI维度的软标签,能够更好地反映人口性格特质的自然分布。数据集的创建过程包括数据过滤、重新标注和软标签估计,旨在解决现有数据集中的标签泄露和无关噪声问题。该数据集主要应用于心理学任务,特别是通过文本内容自动推断个体的MBTI类型,以提高性格检测的准确性和实用性。
MBTIBENCH is the first high-quality MBTI personality detection dataset created by Harbin Institute of Technology and other institutions, aiming to solve the problems of inaccurate self-reported labels and lack of soft labels in existing datasets. The dataset was manually re-annotated under the guidance of psychologists, containing 286 samples that cover soft labels across all four MBTI dimensions, which can better reflect the natural distribution of population personality traits. The creation process of the dataset includes data filtering, re-annotation and soft label estimation, which is designed to address the issues of label leakage and irrelevant noise in existing datasets. This dataset is mainly applied to psychological tasks, especially the automatic inference of an individual's MBTI type from text content, to improve the accuracy and practicality of personality detection.




