Slim Ouni · LORIA · Inria

Platform

MultiMod

A multimodal acquisition platform for speech. It records what the tongue does, what the face does, what the body does and what is heard — in one session, on one time base. Most of the corpora behind our synthesis and gesture work were captured here.

Why one platform

Speech is produced by several systems at once, and the interesting questions are almost always about how they line up: how the jaw opening relates to the acoustic signal, how a gesture stroke anchors on a prosodic peak, how the lips and the tongue coordinate through a consonant cluster. Answering any of these requires the modalities to be recorded together, and to remain comparable millisecond by millisecond afterwards.

MultiMod exists for that. Rather than a single instrument, it is a set of heterogeneous devices — each excellent at one modality — brought onto a common time base.

Equipment

DeviceWhat it capturesFunding
Carstens AG501 Electromagnetic articulograph: 3D trajectories of sensors attached to the tongue, lips and jaw — the articulation itself, inside the mouth, where no camera can see. EQUIPEX ORTOLANG
4 Vicon cameras Marker-based motion capture of the face and upper body.
8 OptiTrack cameras Second-generation motion capture system, with a denser camera ring and a wider capture volume. CPER LCHN
Intel RealSense Depth camera, for markerless data — useful where markers would disturb the speaker or the task. Inria–Région CORExp
Video camera Reference video of the session.
Microphone The acoustic signal.

The hard part

Synchronisation

With hardware this heterogeneous, synchronisation is not a detail of the setup — it is the setup. Each device has its own clock, its own sampling rate and its own start-up latency, and a drift of a few tens of milliseconds is enough to destroy exactly the relationships we are trying to measure. All the devices are therefore driven by a common trigger device, which gives every stream a shared time origin and keeps the recordings alignable throughout the session.

The platform in use

The acquisition room: the AG501 articulograph on the left, four Vicon cameras on tripods, lighting panels, a laptop, and a large screen showing a close-up of a speaker's mouth.
Layout of the acquisition hardwareThe articulograph on the left, the Vicon cameras and their lighting facing the speaker's position, the control station and the monitoring screen. The room is acoustically treated.
A speaker seated in the booth wearing the articulograph cap, sensor wires running from the face, facing the camera rig and a microphone during a recording session.
A session in progressThe speaker wears the articulograph cap; the sensor wires run from the tongue, lips and jaw to the control unit, while the cameras and microphone record the face, the body and the voice.
Close-up of the eight numbered OptiTrack cameras arranged in a ring around a monitor, with a large-diaphragm microphone in front.
OptiTrack motion captureEight cameras arranged around the speaker's position. The ring geometry is what keeps facial and hand markers visible even when the head turns.

From capture to data

Tracking the facial markersOn the left, the speaker equipped with facial sensors; on the right, the reconstructed marker mesh, tracked frame by frame. This is the step between a raw recording and data a model can be trained on.

Working with the platform

MultiMod is open to collaborations. Groups in phonetics, speech science, gesture studies and clinical research have used it for acquisitions they could not run elsewhere, and we are glad to discuss new ones.

Get in touch · Software and data · Projects using the platform