MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation
TLDR
MUSE is a benchmark for Text-to-CAD generation that evaluates manufacturability, functionality, and assemblability of B-Rep assemblies using a VLM judge.
Reasoning
The paper addresses a critical gap in Text-to-CAD evaluation by moving beyond geometric similarity to engineering-relevant criteria. Its strengths include a structured three-stage protocol and human-validated VLM judge, but reliance on VLM may introduce biases and experiments show limited success even for strong models.
Read-first score
Read-first score 69, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 89.
Field roles
Rank sensitivity
Stability: sensitive; rank range: 3.
Keyword Scores
Deep Analysis
Innovations
- Introduction of MUSE, a Text-to-CAD benchmark focusing on complex, editable B-Rep assemblies with structured Design Specifications.
- Three-stage evaluation protocol (code check, geometric check, design-intent alignment) using design-specific rubrics to assess functionality, manufacturability, and assemblability.
- Use of a rubric-based VLM judge for scalable evaluation, validated by human annotation.
Methodology
MUSE pairs practical design instances with structured Design Specifications and evaluates generated CAD models through a three-stage protocol: code check, geometric check, and design-intent alignment. The final stage uses design-specific rubrics and a VLM judge to assess functionality, manufacturability, and assemblability, with human validation of the judge's reliability.
Key Results
Experiments reveal a failure cascade from executable code to valid geometry to engineering-ready design; even the strongest LLMs achieve limited success on fine-grained engineering criteria.