
Modern spoken language systems convert raw audio into compact discrete representations that downstream models generate from and reason over. The central constraint is compression: as few tokens per second as possible while still supporting high fidelity synthesis, pushing designs toward aggressive, low frame rate encodings. Flow matching has recently emerged as a strong generative backbone for this kind of speech reconstruction and generation, known for stable training and efficient, high quality synthesis. Earlier systems relied on separate stages: explicit supervision from auxiliary models trained to capture linguistic and contextual structure, plus independent quantization and reconstruction. Flow matching based systems have instead unified these stages into a single end to end objective, reaching far lower frame rates. What has not been established is whether semantic and contextual content survives intact inside these simplified, flow matching driven representations, degrades unevenly across different kinds of speech, or requires separate treatment once bitrate becomes the binding constraint. This raises concrete open problems: how to detect semantic or contextual loss that standard reconstruction metrics do not expose, how such loss varies across ambiguous, context dependent, or low resource speech, and how any gap connects to the quality of language models built on top of these representations. These are unresolved questions with direct consequences for how the next generation of flow based speech systems should be built.
Team: Md Mubtasim Ahasan, Aman Chadha, Md Mofijul Islam, A K M Mahbubur Rahman


