管道(软件)
未来研究
鉴定(生物学)
计算机科学
钥匙(锁)
领域(数学分析)
精确性和召回率
命名实体识别
滤波器(信号处理)
人工智能
情报检索
数据科学
工程类
植物
任务(项目管理)
数学分析
生物
程序设计语言
系统工程
计算机安全
数学
计算机视觉
作者
Giovanni Puccetti,Vito Giordano,Irene Spada,Filippo Chiarello,Gualtiero Fantoni
标识
DOI:10.1016/j.techfore.2022.122160
摘要
Identifying technologies is a key element for mapping a domain and its evolution. It allows managers and decision makers to anticipate trends for an accurate forecast and effective foresight. Researchers and practitioners are taking advantage of the rapid growth of the publicly accessible sources to map technological domains. Among these sources, patents are the widest technical open access database used in the literature and in practice. Nowadays, Natural Language Processing (NLP) techniques enable new methods for the analysis of patent texts. Among these techniques, in this paper we explore the use of Named Entity Recognition (NER) with the purpose to identify the technologies mentioned in patents' text. We compare three different NER methods, gazetteer-based, rule-based and deep learning-based (e.g. BERT), measuring their performances in terms of precision, recall and computational time. We test the approaches on 1600 patents from four assorted IPC classes as case studies. Our NER systems collected over 4500 fine-grained technologies, achieving the best results thanks to the combination of the three methodologies. The proposed method overcomes the literature thanks to the ability to filter generic technological terms. Our study delineates a valid technology identification tool that can be integrated in any text analysis pipeline to support academics and companies in investigating a technological domain.
科研通智能强力驱动
Strongly Powered by AbleSci AI