Sparse Autoencoders Can Preserve Different Readouts at Equal Reconstruction Error
A new arXiv paper argues that matching reconstruction error and sparsity levels does not guarantee two sparse autoencoders capture the same linearly decodable information from model activations. The authors formalize this gap as a matrix-valued distortion between optimal ridge readouts and propose decoder-preserving training objectives. The work offers a way to evaluate which downstream signals survive sparse compression.