Method Adds Streaming User Transcription to Full-Duplex Speech-to-Speech Models
A new arXiv preprint proposes a way to give full-duplex speech-to-speech models the ability to transcribe the user's speech while the conversation is ongoing. Such models can listen and speak at the same time but generally do not produce their own text transcript of what the user says, which limits uses like live captions and conversational records. The work targets streaming transcription so text is generated as the user speaks rather than after the turn ends.