IWC-Bench: Testing-Based Benchmark for LLM-Generated Web Apps
A new arXiv paper introduces IWC-Bench, a benchmark that assesses LLM-generated web applications from a software testing angle. The authors argue that existing static benchmarks can reward functionality that appears in source code but does not actually work when the app runs. Their approach aims to make automated evaluation of generated web apps track human judgments more closely.