Study analyzes SGD-based learning with synthetic data in high-dimensional linear regression
A newly cross-listed arXiv paper investigates how stochastic gradient descent behaves when training combines human-generated and synthetic data in a high-dimensional linear regression setting. It engages with prior work on model collapse, a phenomenon where keeping even a fixed share of synthetic samples stops model performance from improving as training scales. The findings aim to clarify the conditions under which synthetic data can genuinely extend training beyond limited human datasets.