计算机科学
过程(计算)
编码(内存)
词汇分析
方案(数学)
人工神经网络
序列(生物学)
语言模型
人工智能
刮擦
在制品
生产(经济)
过程控制
数据建模
异常处理
机器学习
计算机集成制造
过程建模
作者
Emmanuel Stathatos,Panorios Benardos,George-Christopher Vosniakos,Dennis Gross,Helge Spieker,Arnaud Gotlieb
标识
DOI:10.1016/j.rcim.2026.103233
摘要
• Novel application of a Large Language Model (LLM) for high-level process planning. • Custom part encoding scheme combines geometric, GD&T, and order information. • LLM is trained from scratch using custom tokenization treating part features and processes uniformly. • LLM autoregressively generates all feasible process chains for a given part encoding. • LLM achieves over 99% accuracy using only 5% of the training dataset. This study applies Large Language Models (LLMs) to high-level Computer-Aided Process Planning (CAPP) in a distributed manufacturing context. It aims to generate alternative, feasible process chains for production of a wide range of parts. Parts are encoded in a custom encoding scheme supporting diverse part overall shapes, geometrical features within them, and corresponding manufacturing processes. The CAPP problem is formulated as a sequence prediction task, where a GPT-2-based LLM generates process chains autoregressively. To train and test the LLM a synthetic dataset of 7,840 unique parts and their alternative process chains was generated using expert-driven rule-based logic. The LLM is trained from scratch using a tokenization scheme treating part features and processes uniformly as discrete tokens, special tokens being employed to control sequence flow. Performance evaluation was performed for systematically reducing the size of the dataset. Finally, even with as little as 5% of the training data, the LLM achieves over 99% accuracy at the process chain-level. The extremely few spotted errors mainly involve minor secondary process mispredictions without critical failures. For comparison, a Recurrent Neural Network (RNN) was also trained with the same dataset. Since manufacturing data stemming from experts and not from sensors is notoriously difficult to collect, training a machine learning model with a dataset that is as small as possible is of utmost importance. In this light, the LLM proved superior to RNN, in fact emphatically so, the more the training dataset was limited.
科研通智能强力驱动
Strongly Powered by AbleSci AI