计算机科学
知识图
个性化
人格心理学
自然语言处理
人格
图形
词典
情报检索
人工智能
数据科学
理论计算机科学
万维网
心理学
社会心理学
作者
Qingren Wang,Ao Liu,Yan Kang,Juan Hou,Wei Li
标识
DOI:10.1109/ickg59574.2023.00015
摘要
Obtaining the personalities of users conveyed by their published short texts has a wide and important range of applications, from detecting abnormal behavior of online users to accurately customization recommendation. Advancement in this area can be improved using large-scale datasets with coarse- and fine-grained typologies, adaptable to multiple downstream tasks, i.e., psychology knowledge graph construction and intelligence collection of psychological questionaries. Although the LIWC lexicon helps personality prediction, the larger volume of textual dataset related to Big Five personality types still cannot meet the requirements of researches and applications, especially for tasks in Chinese. Therefore, this paper introduces BigFive, a large, high quality Chinese textual dataset manually annotated by the psychological experts. BigFive contains 13,478 Chinese phrases that belong to five categories (coarse-grained) and 30 categories (fine-grained). The reliability of five categories grouped by personality level and 30 categories grouped by dimension level is demonstrated via a detailed data analysis. In addition, a strong baseline is build based on a fine-tuning BERT model. Our BERT-based model achieves an average F1-score of. 33 (std=.24) in terms of 30 categories and an average F1-score of. 66 (std=.05) in terms of five categories. The experimental results suggest that there is much room for improvement.
科研通智能强力驱动
Strongly Powered by AbleSci AI