管道(软件)
数据提取
计算机科学
数据集
数据科学
情报检索
萃取(化学)
集合(抽象数据类型)
环境数据
数据挖掘
人工智能
梅德林
化学
政治学
色谱法
程序设计语言
法学
生物化学
作者
Yushu Cheng,Huichun Zhang
出处
期刊:ACS ES&T water
[American Chemical Society]
日期:2025-07-04
卷期号:5 (8): 4897-4907
被引量:4
标识
DOI:10.1021/acsestwater.5c00551
摘要
Large language models (LLMs) have introduced a new paradigm for automated information extraction, significantly reducing manual effort. While many LLM-based pipelines perform well in fields like materials and medical research, their application in environmental domains remains limited. Document-level LLM frameworks offer new opportunities to address the unique challenges of environmental literature, where relevant information is often scattered and requires selective, multifeature extraction informed by domain context. This study introduces L2D, a structured multistep prompting pipeline that dynamically integrates domain constraints and an LLM-based self-validation process. L2D employs contextual prompting strategies to extract only relevant experimental details (e.g., chemicals undergoing biodegradation, not medium components), while validation mitigates hallucinations and misinterpretations. We benchmarked L2D against two LLM pipelines (ChatExtract and PropertyExtractor) and a deep learning model (ChemREL). While ChemREL achieved only 0.53 F1 on our chemical recognition task─reflecting the limitations of task-specific deep learning models in nuanced contexts─L2D achieved 93% accuracy in extracting eight biodegradation-related features from 79 articles, screened from an initial pool of 500 in under 30 min. Without requiring additional model training, L2D demonstrates superior precision, adaptability, and scalability, making it a practical, low-code tool for structured data set generation in environmental research.
科研通智能强力驱动
Strongly Powered by AbleSci AI