Slim Ouni · LORIA · Inria

Research

Speech is more than sound

It is produced by a body: a vocal tract that moves, a face that deforms, hands that gesture - all parts of a body orchestrated as a coherent whole, precisely coordinated in time. My work tries to model that whole - to measure it, to understand it, and to generate it.

Direction 01

Expressive audiovisual speech synthesis and analysis

Generating speech together with the face that produces it, so that the two are not merely aligned but genuinely co-produced. This covers coarticulation modelling, the transfer of expressivity and emotion onto a talking head, and the objective and perceptual evaluation of visual intelligibility. It is also where the work becomes usable outside the lab: automatic lipsync for 3D characters, and visual dubbing of filmed speakers into another language.

  • talking head
  • coarticulation
  • expressive synthesis
  • visual intelligibility
  • lipsync
  • visual dubbing
ProjectsANR Full3DTalkinghead (2021–2024) · EPIC Games MegaGrants (2022–2023) · Inria PuppetFace (2018–2020) · ANR ViSAC (2009–2013)
Doctoral workSara Dahmani (2017–2020) · Théo Biasutto (2016–2021) · Utpala Musti (2009–2013)

Direction 02

Speech and gesture in interaction

Co-speech gesture is not decoration added to an utterance; it is part of it. We study how the two modalities are synchronised — at which temporal grain, around which prosodic and semantic anchors — and we build models that generate full-body gesture from the speech signal alone. The interest of the problem is that the mapping is neither deterministic nor arbitrary: the same sentence admits many gestures, but not any gesture.

  • co-speech gesture
  • multimodal synchronisation
  • gesture generation
  • prosody
  • motion capture
ProjectsANR Syncogest (2025–2029, coordinator) · CNRS PRIME 80 GEPACI (2021–2024) · HumanE-AI-Net (2020–2023)
Doctoral workLouis Abel (2021–2025) · Mickaëlla Grondin-Verdon (2021–2025)

Direction 03

Sign language production and processing

A more recent direction, and a demanding one: sign language is fully multimodal — manual articulation, facial expression, torso and gaze — and its resources are scarce. We work on end-to-end generation from speech or text, and on the latent representations of pose that make diffusion-based production learnable. Much of the effort goes into the representation itself, which is where the quality of the generated signing is decided.

  • sign language production
  • pose representation
  • diffusion models
  • low-resource languages
ProjectsROGSILT, Inria-DFKI project (2026-2029) · Défi Inria COLaF (2023-2027) · ANR LOR-AI (2021–2024)
Doctoral workGuilhem Faure (2024–2027)

Direction 04

Clinical and educational applications

When the models are good enough to describe typical speech, their errors on atypical speech become informative. We use multimodal speech analysis for the automatic detection and characterisation of disfluencies in stuttering, for the early detection of Corhn's disease from speech and facial signals, and — on the educational side — to give second-language learners visual feedback on their own articulation.

  • disfluency detection
  • atypical speech
  • speech biomarkers
  • pronunciation training
  • visual feedback
ProjectsANR RHU I-DEAL (2024–2028) · ANR BENEPHIDIRE (2019–2022) · PIA-2 e-FRAN METAL (2017–2021)
Doctoral workShakeel Ahmad Sheikh (2019–2022) · Hamza Dalhoumi (2026-2029)

Direction 05

Speech production, articulatory modelling and multimodal data

The measurement side of the work, and its oldest thread. Acquiring synchronised articulatory, acoustic, facial and body data; modelling the vocal tract; recovering articulation from the acoustic signal. This is also what makes the rest possible — the MultiMod platform and the tools built around it exist so that the other directions have data of the right quality, and so that other groups can use it too.

  • electromagnetic articulography
  • acoustic-to-articulatory inversion
  • vocal tract modelling
  • Arabic pharyngealisation
  • corpus design
ProjectsINS2I MULTIMOD (2022–2023) · Inria ADT VisArtico (2013–2015) · PIA-2 e-FRAN METAL (2017–2021) · CMCU France–Tunisia (2007–2013)
Doctoral workImen Jemaa (2009–2013)

A note on method

Across the five directions the same methodological question keeps returning: what is the right shared representation for signals that are produced together but observed in very different spaces — sound, articulator positions, meshes, skeletons. Most of our recent work sits there, at the meeting point of generative modelling and multimodal representation learning.