Awesome AI4CAD Hub Papers · Datasets · Projects
← Back to papers

MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation

arXiv 2026 69 method

TLDR

MUSE is a benchmark for Text-to-CAD generation that evaluates manufacturability, functionality, and assemblability of B-Rep assemblies using a VLM judge.

Reasoning

The paper addresses a critical gap in Text-to-CAD evaluation by moving beyond geometric similarity to engineering-relevant criteria. Its strengths include a structured three-stage protocol and human-validated VLM judge, but reliance on VLM may introduce biases and experiments show limited success even for strong models.

Read-first score

Read-first score 69, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 89.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,evaluation,experiment,metric,validation

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
55.6

Uses existing LLM keyword relevance scores normalized to 0-100. AI for CAD,computer-aided design,neural CAD,generative CAD,parametric CAD,B-Rep,boundary representation,constructive solid geometry,CSG,sketch extrusion,CAD generation,CAD reconstruction,text-to-CAD,image-to-CAD,point cloud to CAD,CAD program

Reproducibility 25%
50

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset,github

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: sensitive; rank range: 3.

Keyword Scores

text-to-CAD
10
B-Rep
9
boundary representation
9
CAD generation
9
AI for CAD
7
generative CAD
7
CAD program
7
computer-aided design
6
CAD reconstruction
6
neural CAD
5
parametric CAD
5
sketch extrusion
3
constructive solid geometry
2
CSG
2
image-to-CAD
1
point cloud to CAD
1

Deep Analysis

Innovations

  • Introduction of MUSE, a Text-to-CAD benchmark focusing on complex, editable B-Rep assemblies with structured Design Specifications.
  • Three-stage evaluation protocol (code check, geometric check, design-intent alignment) using design-specific rubrics to assess functionality, manufacturability, and assemblability.
  • Use of a rubric-based VLM judge for scalable evaluation, validated by human annotation.

Methodology

MUSE pairs practical design instances with structured Design Specifications and evaluates generated CAD models through a three-stage protocol: code check, geometric check, and design-intent alignment. The final stage uses design-specific rubrics and a VLM judge to assess functionality, manufacturability, and assemblability, with human validation of the judge's reliability.

Key Results

Experiments reveal a failure cascade from executable code to valid geometry to engineering-ready design; even the strongest LLMs achieve limited success on fine-grained engineering criteria.

Tags

AI