Research
Speech is more than sound
It is produced by a body: a vocal tract that moves, a face that deforms, hands that gesture - all parts of a body orchestrated as a coherent whole, precisely coordinated in time. My work tries to model that whole - to measure it, to understand it, and to generate it.
Expressive audiovisual speech synthesis and analysis
Generating speech together with the face that produces it, so that the two are not merely aligned but genuinely co-produced. This covers coarticulation modelling, the transfer of expressivity and emotion onto a talking head, and the objective and perceptual evaluation of visual intelligibility. It is also where the work becomes usable outside the lab: automatic lipsync for 3D characters, and visual dubbing of filmed speakers into another language.
- talking head
- coarticulation
- expressive synthesis
- visual intelligibility
- lipsync
- visual dubbing
Doctoral workSara Dahmani (2017–2020) · Théo Biasutto (2016–2021) · Utpala Musti (2009–2013)
Speech and gesture in interaction
Co-speech gesture is not decoration added to an utterance; it is part of it. We study how the two modalities are synchronised — at which temporal grain, around which prosodic and semantic anchors — and we build models that generate full-body gesture from the speech signal alone. The interest of the problem is that the mapping is neither deterministic nor arbitrary: the same sentence admits many gestures, but not any gesture.
- co-speech gesture
- multimodal synchronisation
- gesture generation
- prosody
- motion capture
Doctoral workLouis Abel (2021–2025) · Mickaëlla Grondin-Verdon (2021–2025)
Sign language production and processing
A more recent direction, and a demanding one: sign language is fully multimodal — manual articulation, facial expression, torso and gaze — and its resources are scarce. We work on end-to-end generation from speech or text, and on the latent representations of pose that make diffusion-based production learnable. Much of the effort goes into the representation itself, which is where the quality of the generated signing is decided.
- sign language production
- pose representation
- diffusion models
- low-resource languages
Doctoral workGuilhem Faure (2024–2027)
Clinical and educational applications
When the models are good enough to describe typical speech, their errors on atypical speech become informative. We use multimodal speech analysis for the automatic detection and characterisation of disfluencies in stuttering, for the early detection of Corhn's disease from speech and facial signals, and — on the educational side — to give second-language learners visual feedback on their own articulation.
- disfluency detection
- atypical speech
- speech biomarkers
- pronunciation training
- visual feedback
Doctoral workShakeel Ahmad Sheikh (2019–2022) · Hamza Dalhoumi (2026-2029)
Speech production, articulatory modelling and multimodal data
The measurement side of the work, and its oldest thread. Acquiring synchronised articulatory, acoustic, facial and body data; modelling the vocal tract; recovering articulation from the acoustic signal. This is also what makes the rest possible — the MultiMod platform and the tools built around it exist so that the other directions have data of the right quality, and so that other groups can use it too.
- electromagnetic articulography
- acoustic-to-articulatory inversion
- vocal tract modelling
- Arabic pharyngealisation
- corpus design
Doctoral workImen Jemaa (2009–2013)
A note on method
Across the five directions the same methodological question keeps returning: what is the right shared representation for signals that are produced together but observed in very different spaces — sound, articulator positions, meshes, skeletons. Most of our recent work sits there, at the meeting point of generative modelling and multimodal representation learning.