计算机科学
人工智能
自然语言处理
领域(数学)
易读性
语言模型
特征(语言学)
鉴定(生物学)
钥匙(锁)
自然语言
人机交互
计算机视觉
透视图(图形)
交叉口(航空)
语义学(计算机科学)
作者
Keyu Lu,Xin Zhao,Manchun Li,Zhenkang Wang,Ying Zhou,Jie Wang
出处
期刊:International journal of geographical information systems
[Taylor & Francis]
日期:2025-10-08
卷期号:40 (5): 1547-1575
被引量:2
标识
DOI:10.1080/13658816.2025.2566795
摘要
Street view imagery (SVI) can capture urban physical space features, residents’ sentiments, and soundscape characteristics, providing insights into complex relationships between human activities and the built environment. However, existing SVI-based urban perception models have limitations, including insufficient model generality, neglect of element interactions, and lack of logical reasoning capabilities. The study proposes a multimodal large language model, StreetSenser, with powerful SVI analysis and general capabilities learned from human perception. First, an optimized dataset was designed and constructed, encompassing diverse street scenarios and annotations from human perception. Subsequently, StreetSenser was developed through a two-stage Chain-of-Perception (CoP) fine-tuning method, with the ‘Object - Attribute - Relation’ framework applied to the Qwen2.5-VL model to enable perception of complex information across visual, emotional and acoustic dimensions. Results revealed that modules such as SVI description and CoP enhance urban perception capacity of StreetSenser significantly. Furthermore, with only seven billion parameters, StreetSenser demonstrates a performance comparable to or exceeding that of large models such as GPT-4o across various tasks, while also approaching or surpassing traditional models. In addition, the compact model size enables efficient and cost-effective deployment for large-scale urban applications. The results validate the potential of StreetSenser as a key decision-support tool for urban planning.
科研通智能强力驱动
Strongly Powered by AbleSci AI