时间轴
自动汇总
计算机科学
事件(粒子物理)
情报检索
历史
量子力学
物理
考古
作者
Qiang Mao,Adam Dąbrowski,Fusheng Wei,Eric N. Olson,Robert Neary,Jingchao Yang,Han Qin,Nathaniel Huber-Fliflet
标识
DOI:10.1109/bigdata62323.2024.10826063
摘要
This paper presents a comparative study evaluating the performance of Large Language Models (LLMs) in generating timeline summaries from construction delay documents. We assessed seven open-source LLMs and two commercial chatbots (ChatGPT and Claude) on their ability to extract, organize, and summarize delay events from twenty-one carefully curated synthetic snippets of text. The evaluation framework combined automatic metrics (BERTScore and ROUGE scores) with expert human assessment across four dimensions: event description accuracy, date accuracy, event capture completeness, and language quality.Results demonstrate that while commercial solutions, particularly Claude, achieved superior performance, several open-source alternatives showed comparable capabilities. Notably, Llama-3.1-70B-Instruct showed robust performance in event capture and source tracking, while Llama-3.1-8B-Instruct offered efficient processing with balanced performance among smaller models. A critical finding was the widespread challenge in temporal information processing, with only Claude achieving complete accuracy in date extraction and event association. The study's findings suggest that open-source LLMs can serve as practical tools for construction document analysis, although model selection is a critical consideration based on specific accuracy and efficiency requirements, and resource constraints.
科研通智能强力驱动
Strongly Powered by AbleSci AI