MM‐CAD: A Multi‐Modal CAD Dataset and Benchmark for Cross‐Modal Geometric Learning
Anush Bharathi, Ananthakrishnan A, Ramanathan Muthuganapathy
Indian Institute of Technology Madras
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Computer‐Aided Design (CAD) boosts modern manufacturing, yet design reuse remains constrained by the absence of large, openly available CAD repositories with rich multi‐modal annotations suitable for search/retrieval. Recent large‐scale efforts to annotate public datasets rely on hash‐based redundancy removal that leaves no semantic structure, and on captioning by Vision‐Language Models (VLMs) using rendered images alone, which struggles to capture geometric and procedural information. We introduce MM‐CAD, a multi‐modal CAD dataset designed to level‐up retrieval and retrieval‐augmented generation models for engineering geometry, comprising two complementary parts. MM‐CAD:A brings 33,816 unique CAD models from eleven widely used benchmark datasets under a common identifier scheme, with isometric renderings, point clouds, and humanly‐curated multi‐level text captions, and 4,376 real hand‐drawn user sketches among others. MM‐CAD:B curates 192,626 models from the 1M‐model ABC corpus through a seven‐stage pipeline centered on Manifold‐Aware Adaptive Sampling (MAAS), which organizes models into semantically coherent neighborhoods rather than merely removing duplicates, directly supplying the hard negatives that contrastive retrieval training requires. Every retained model is annotated through a metadata‐grounded pipeline that conditions caption generation on parsed construction sequences rather than rendered views alone, producing three‐level text descriptions, multi‐level contour sketches, a hierarchical application taxonomy, and photorealistic in‐context images that largely preserve source CAD geometry, a modality not previously available at this scale on CAD data. We further introduce a joint retrieval architecture that aligns sketch, text, image, B‐Rep, and point cloud encoders in a single latent space through Matryoshka‐nested contrastive objectives, establishing the first unified cross‐modal retrieval benchmark for large‐scale CAD.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
工程3D Shape Modeling and Analysis
Manufacturing Process and Optimization · Interactive and Immersive Displays
参考文献 40
此处列出前 3 条