Platform
MultiMod
A multimodal acquisition platform for speech. It records what the tongue does, what the face does, what the body does and what is heard — in one session, on one time base. Most of the corpora behind our synthesis and gesture work were captured here.
Why one platform
Speech is produced by several systems at once, and the interesting questions are almost always about how they line up: how the jaw opening relates to the acoustic signal, how a gesture stroke anchors on a prosodic peak, how the lips and the tongue coordinate through a consonant cluster. Answering any of these requires the modalities to be recorded together, and to remain comparable millisecond by millisecond afterwards.
MultiMod exists for that. Rather than a single instrument, it is a set of heterogeneous devices — each excellent at one modality — brought onto a common time base.
Equipment
| Device | What it captures | Funding |
|---|---|---|
| Carstens AG501 | Electromagnetic articulograph: 3D trajectories of sensors attached to the tongue, lips and jaw — the articulation itself, inside the mouth, where no camera can see. | EQUIPEX ORTOLANG |
| 4 Vicon cameras | Marker-based motion capture of the face and upper body. | — |
| 8 OptiTrack cameras | Second-generation motion capture system, with a denser camera ring and a wider capture volume. | CPER LCHN |
| Intel RealSense | Depth camera, for markerless data — useful where markers would disturb the speaker or the task. | Inria–Région CORExp |
| Video camera | Reference video of the session. | — |
| Microphone | The acoustic signal. | — |
Synchronisation
With hardware this heterogeneous, synchronisation is not a detail of the setup — it is the setup. Each device has its own clock, its own sampling rate and its own start-up latency, and a drift of a few tens of milliseconds is enough to destroy exactly the relationships we are trying to measure. All the devices are therefore driven by a common trigger device, which gives every stream a shared time origin and keeps the recordings alignable throughout the session.
The platform in use
From capture to data
Working with the platform
MultiMod is open to collaborations. Groups in phonetics, speech science, gesture studies and clinical research have used it for acquisitions they could not run elsewhere, and we are glad to discuss new ones.
Get in touch · Software and data · Projects using the platform