Project

Speech-to-Motion: From Next-Token to Next-Gesture Prediction

Speech-to-Motion: From Next-Token to Next-Gesture Prediction

Large language models learn to predict sequences from nothing but text, yet human communication was never just words. Every sentence is accompanied by gesture: the hands, the shoulders, the tilt of a head, timed to the rhythm and emphasis of speech in ways no transcript captures. This project asks whether the sequence-modeling competence a language model develops purely from text extends to that embodied signal: predicting a speaker's hand and upper-body gesture directly from their voice, with no language involved at either end.

Co-speech gesture is a narrow but demanding test case. Current systems that generate it are purpose-built motion models trained from scratch, disconnected from the sequence competence large language models already have. Closing that gap raises specific open problems. One is data efficiency: how little paired speech-motion data is enough for real voice-to-gesture correspondence to emerge, rather than memorized fragments of a few speakers. Another is representation: how much of continuous physical movement survives being compressed into a form a sequence model can learn from, and what that loss costs in naturalness. A third is scale: whether quality keeps improving as training data grows toward realistic conversational volume, or plateaus for reasons specific to this transfer.

The broader question is how far the sequence-modeling capabilities learned by LLMs can extend beyond language into embodied modalities.

Team: Md Mubtasim Ahasan, Pranjal Nandi, Aman Chadha, Md Mofijul Islam, A K M Mahbubur Rahman