预处理器
计算机科学
信息抽取
概念证明
非结构化数据
自然语言处理
数据科学
情报检索
人工智能
数据挖掘
大数据
操作系统
标识
DOI:10.1177/20597991251313876
摘要
Recent months have witnessed an increase in suggested applications for large language models (LLMs) in the social sciences. This proof-of-concept paper explores the use of LLMs to improve text quality and to extract predefined information from unstructured text. The study showcases promising results with an example focussed on historical newspapers and highlights the effectiveness of LLMs in correcting errors in the parsed text and in accurately extracting specified information. By leveraging the capabilities of LLMs in these straightforward, instruction-based tasks, this research note demonstrates their potential to improve on the efficiency and accuracy of text analysis workflows. The ongoing development of LLMs and the emergence of robust open-source options underscores their increasing accessibility for both, the quantitative and qualitative, social sciences and other disciplines working with text data.
科研通智能强力驱动
Strongly Powered by AbleSci AI