papersSEP 12 04:00 UTC
Logit Refiner Targets Intra-Scale Dependencies in Visual Autoregressive Models
A new arXiv paper analyzes visual autoregressive models, which generate images by predicting one scale of tokens at a time and emitting all tokens in a scale in parallel. The authors argue this parallel decoding acts as a mean-field-style approximation that drops spatial dependencies within each scale. They propose a Logit Refiner method that models these intra-scale relationships to improve generation quality.