papersTODAY 04:00 UTC
DepthBenchCAD Examines Whether More Auditing Checks Improve Generative CAD Evaluations
A new arXiv preprint introduces DepthBenchCAD, a benchmark studying how the number of edit checks affects the reliability of evaluations for generative CAD models. The work focuses on behavioral correctness after parameter edits and asks whether auditing more programs under a fixed budget actually leads to firmer conclusions. It questions the common assumption that adding edit checks is a straightforward path to more trustworthy evaluation.