arXiv paper analyzes how AI-generated data affects dataset decomposition
A new preprint examines what happens to batch decomposition and downstream model performance when training sets mix human data with text or images produced by existing large language models. Using random datasets containing anomalies, the authors study criticality in dissimilar decomposition and undersampling techniques. The work aims to clarify the statistical behavior of datasets that are increasingly populated with synthetic samples.