papersSEP 10 04:00 UTC
Time-Series Foundation Model Benchmarks Still Reflect Pretraining Familiarity on Later Hold-Outs
A new study questions whether time-series foundation models can be fairly evaluated using test data collected after their pretraining cutoff. It finds that even a temporally later, contamination-free hold-out does not fully isolate genuine generalization, as familiarity with the underlying data distribution absorbed during pretraining persists. The result suggests the field needs evaluation practices that go beyond simply withholding recent data.
benchmarking practicesdata-contaminationmodel-generalizationpretraining familiaritytemporal hold-out evaluationtime-series-foundation-models
COVERAGE · 1 REPORT · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.LGA Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-Out ↗SEP 10 04:00 UTC