<p>Automated synthesis of computer-aided design (CAD) models from engineering drawings is a fundamental challenge in intelligent manufacturing, demanding not only geometric fidelity but also executable and fully editable output representations that can be directly consumed by downstream computer-aided engineering and manufacturing tools. Existing reconstruction approaches predominantly yield static, non-editable meshes or boundary representations that are incompatible with the precision and parametric editability requirements of industrial workflows. Furthermore, contemporary methods that leverage text or image descriptions as inputs impose substantial annotation overhead, severely constraining their applicability at production scale. To address these limitations, we propose the Heterogeneous fMulti-Expert Collaborative Reinforcement Learning (HMEC-CAD) framework, a principled training paradigm for direct CAD program generation from orthographic projections. HMEC-CAD systematically exploits the complementary reasoning capabilities of heterogeneous expert models through two coordinated training stages: Multi-Expert Fine-Tuning (MEFT), which distills diverse chain-of-thought reasoning styles from multiple pre-trained experts, and Multi-Expert Reinforcement Learning (MERL), which promotes cross-expert knowledge transfer and addresses reward sparsity through a hard negative sample buffering mechanism. To support rigorous evaluation under industrial conditions, we further introduce CADExpert, a large-scale open benchmark comprising 17,299 instances, each consisting of orthographic projections with precise dimensional annotations, expert-generated reasoning traces, executable CADQuery programs, and rendered three-dimensional CAD models. Extensive experiments demonstrate that HMEC-CAD achieves superior geometric accuracy, code executability, and robustness relative to both zero-shot vision-language baselines and single-expert reinforcement learning counterparts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Heterogeneous multi-expert collaborative reinforcement learning for automated CAD program synthesis from engineering drawings

  • Zehe Yin,
  • Weixin Lin,
  • Xiangbiao Kong

摘要

Automated synthesis of computer-aided design (CAD) models from engineering drawings is a fundamental challenge in intelligent manufacturing, demanding not only geometric fidelity but also executable and fully editable output representations that can be directly consumed by downstream computer-aided engineering and manufacturing tools. Existing reconstruction approaches predominantly yield static, non-editable meshes or boundary representations that are incompatible with the precision and parametric editability requirements of industrial workflows. Furthermore, contemporary methods that leverage text or image descriptions as inputs impose substantial annotation overhead, severely constraining their applicability at production scale. To address these limitations, we propose the Heterogeneous fMulti-Expert Collaborative Reinforcement Learning (HMEC-CAD) framework, a principled training paradigm for direct CAD program generation from orthographic projections. HMEC-CAD systematically exploits the complementary reasoning capabilities of heterogeneous expert models through two coordinated training stages: Multi-Expert Fine-Tuning (MEFT), which distills diverse chain-of-thought reasoning styles from multiple pre-trained experts, and Multi-Expert Reinforcement Learning (MERL), which promotes cross-expert knowledge transfer and addresses reward sparsity through a hard negative sample buffering mechanism. To support rigorous evaluation under industrial conditions, we further introduce CADExpert, a large-scale open benchmark comprising 17,299 instances, each consisting of orthographic projections with precise dimensional annotations, expert-generated reasoning traces, executable CADQuery programs, and rendered three-dimensional CAD models. Extensive experiments demonstrate that HMEC-CAD achieves superior geometric accuracy, code executability, and robustness relative to both zero-shot vision-language baselines and single-expert reinforcement learning counterparts.