JSON文件
工作流程
计算机科学
忠诚
软件工程
可执行文件
程序设计语言
图形
编码(社会科学)
编配
工程设计过程
迭代设计
迭代和增量开发
建模语言
可扩展性
需求工程
适配器(计算)
系统工程
编码(集合论)
源代码
代码生成
系统设计
设计过程
理论计算机科学
功能要求
正确性
软件设计模式
过程(计算)
复杂系统
元建模
分布式计算
模块化设计
相容性(地球化学)
过程建模
场景图
封装(网络)
自动化
人机交互
作者
Soheyl Massoudi,Mark Fuge
摘要
Abstract Early-stage engineering design involves complex, iterative reasoning, yet existing large language model (LLM) workflows struggle to maintain task continuity and generate executable models. We evaluate whether a structured multi-agent system (MAS) can more effectively manage requirements extraction, functional decomposition, and simulator code generation than a simpler two-agent system (2AS). The target application is a solar-powered water filtration system as described in a cahier des charges. We introduce the design-state graph (DSG), a JSON-serializable representation that bundles requirements, physical embodiments, and python-based physics models into graph nodes. A nine-role MAS iteratively builds and refines the DSG, while the 2AS collapses the process to a generator–reflector loop. Both systems run a total of 60 experiments (2 LLMs—Llama 3.3 70B versus reasoning-distilled deepseek R1 70B × 2 agent configurations × 3 temperatures × 5 seeds). We report a JSON validity, requirement coverage, embodiment presence, code compatibility, workflow completion, runtime, and graph size. Across all runs, both MAS and 2AS maintained perfect JSON integrity and embodiment tagging. Requirement coverage remained minimal (less than 20%). Code compatibility peaked at 100% under specific 2AS settings but averaged below 50% for MAS. Only the reasoning-distilled model reliably flagged workflow completion. Powered by deepseek R1 70B, the MAS generated more granular DSGs (average 5–6 nodes) whereas 2AS mode collapsed. Structured multi-agent orchestration enhanced design detail. Reasoning-distilled LLM improved completion rates, yet low requirements and fidelity gaps in coding persisted.
科研通智能强力驱动
Strongly Powered by AbleSci AI