papersSEP 10 04:00 UTC
PACE framework targets perceived latency in retrieval-augmented dialogue serving
Researchers introduce PACE, a serving framework for retrieval-augmented dialogue systems that defines Perceived Time-to-First-Response as a quality-of-experience metric and optimizes it subject to quality and cost limits. The approach combines cascaded service routing with filler response control to reduce how long users wait before receiving an initial answer. The work extends prior research on cascading and semantic caching by treating perceived latency as an explicit optimization objective.