可解释性
计算机科学
背景(考古学)
语言学
自然语言处理
人工智能
历史
哲学
考古
作者
Andrés Piñeiro-Martín,Francisco-Javier Santos-Criado,Carmén García Mateo,Laura Docío-Fernández,María del Carmen López-Pérez
出处
期刊:Applied sciences
[Multidisciplinary Digital Publishing Institute]
日期:2025-01-24
卷期号:15 (3): 1192-1192
被引量:4
摘要
Large language models (LLMs) have revolutionized the field of artificial intelligence in both academia and industry, transforming how we communicate, search for information, and create content. However, these models face knowledge cutoffs and costly updates, driving a new ecosystem for LLM-based applications that leverage interaction techniques to extend capabilities and facilitate knowledge updates. As these models grow more complex, understanding their internal workings becomes increasingly challenging, posing significant issues for transparency, interpretability, and explainability. This paper proposes a novel approach to interpretability by shifting the focus to understanding the model’s functionality within specific contexts through interaction techniques. Rather than dissecting the LLM itself, we explore how contextual information and interaction techniques can elucidate the model’s thought processes. To this end, we introduce the Context-Driven Divergent Knowledge Evaluation (CDK-E) methodology, along with the Divergent Knowledge Dataset (DKD), for evaluating the interpretability of LLMs in context-specific scenarios that diverge from the model’s inherent knowledge. The empirical results demonstrate that advanced LLMs achieve high alignment with divergent contexts, validating our hypothesis that contextual information significantly enhances interpretability. Moreover, the strong correlation between LLM-based metrics and semantic metrics confirms the reliability of our evaluation framework.
科研通智能强力驱动
Strongly Powered by AbleSci AI