{"id":304,"date":"2016-10-27T09:35:01","date_gmt":"2016-10-27T07:35:01","guid":{"rendered":"http:\/\/members.loria.fr\/SOuni\/?page_id=304"},"modified":"2018-12-13T22:22:41","modified_gmt":"2018-12-13T20:22:41","slug":"audiovisual-speech-2","status":"publish","type":"page","link":"https:\/\/members.loria.fr\/SOuni\/accueil\/projets\/audiovisual-speech-2\/","title":{"rendered":"Audiovisual Speech"},"content":{"rendered":"<h3>Audiovisual Speech Synthesis<\/h3>\n<div style=\"width: 640px;\" class=\"wp-video\"><video class=\"wp-video-shortcode\" id=\"video-304-1\" width=\"640\" height=\"360\" poster=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2016\/09\/av-exp-ca.png\" preload=\"metadata\" controls=\"controls\"><source type=\"video\/mp4\" src=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2016\/10\/demo1-25sec-lores.mp4?_=1\" \/><a href=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2016\/10\/demo1-25sec-lores.mp4\">http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2016\/10\/demo1-25sec-lores.mp4<\/a><\/video><\/div>\n<p>Our focus in the field of audiovisual speech is\u00a0audiovisual synthesis\u00a0of highly intelligible speech. We are investigating methodologies of performing synthesis with its acoustic and visual components simultaneously. Therefore, we consider audiovisual speech as a bimodal signal with two channels: acoustic and visual. One of our purposes is to develop a highly realistic face animation, mainly the animation of the lips.<br \/>\nWhen dealing with audiovisual synthesis, we consider that it is important to study\u00a0audiovisual intelligibility\u00a0and the ability of the synthesis to send an intelligible message to the human receiver. In fact, the intelligibility of the audiovisual synthesis can be critical when considering applications addressed to hard-of-hearing humans or to learners of new languages. Our research in audiovisual speech intelligibility concerns the experimental evaluation, developing metrics to measure intelligibility.<\/p>\n<div style=\"width: 648px;\" class=\"wp-video\"><video class=\"wp-video-shortcode\" id=\"video-304-2\" width=\"648\" height=\"365\" preload=\"auto\" controls=\"controls\"><source type=\"video\/mp4\" src=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2018\/09\/heidi_bilab_withAudio.mp4?_=2\" \/><a href=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2018\/09\/heidi_bilab_withAudio.mp4\">http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2018\/09\/heidi_bilab_withAudio.mp4<\/a><\/video><\/div>\n<p style=\"text-align: center\"><em>Example of automatic lipsync using Dynalips technology<\/em><\/p>\n<h3><\/h3>\n<h3>Multimodal data acquisition<\/h3>\n<p><a href=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2017\/01\/plate-forme2s.jpg\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-343 alignleft\" src=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2017\/01\/plate-forme2s-300x200.jpg\" alt=\"\" width=\"324\" height=\"216\" srcset=\"https:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2017\/01\/plate-forme2s-300x200.jpg 300w, https:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2017\/01\/plate-forme2s-768x513.jpg 768w, https:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2017\/01\/plate-forme2s.jpg 854w\" sizes=\"auto, (max-width: 324px) 100vw, 324px\" \/><\/a><a href=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2018\/10\/mocap-optitrack.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-501 alignleft\" src=\"http:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2018\/10\/mocap-optitrack-300x232.png\" alt=\"\" width=\"274\" height=\"212\" srcset=\"https:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2018\/10\/mocap-optitrack-300x232.png 300w, https:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2018\/10\/mocap-optitrack-768x593.png 768w, https:\/\/members.loria.fr\/SOuni\/wp-content\/blogs.dir\/133\/files\/sites\/133\/2018\/10\/mocap-optitrack.png 800w\" sizes=\"auto, (max-width: 274px) 100vw, 274px\" \/><\/a><\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n<p>Our work in audiovisual speech relies on acquiring data and processing it. This can be time-consuming and costly in terms of the effort required to carry out the acquisition and processing the data. This effort is unavoidable to make more progress in modeling the processes related to human communication.<br \/>\nWe are continuously working on improving the acquisition techniques and investigating methods to make the process easier. In audiovisual speech, we used in the past sterovision technique to aquire 3D facial data. More recently, we are using more advanced motion capture techniques using VICON and also we are testing the use of cheaper hardware based on the kinect.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Audiovisual Speech Synthesis<\/p>\n<p>Our focus in the field of audiovisual speech is\u00a0audiovisual synthesis\u00a0of highly intelligible speech. We are investigating methodologies of performing synthesis with its acoustic and visual components simultaneously. Therefore, we consider audiovisual speech as a bimodal signal with two channels: acoustic and visual. One of our purposes is to develop a highly realistic face animation, mainly the animation of the lips.<br \/>\nWhen dealing with audiovisual synthesis, we consider that it is important to study\u00a0audiovisual intelligibility\u00a0and the ability of the synthesis to send an intelligible message to the human receiver. In fact, the intelligibility of the audiovisual synthesis can be critical when considering applications addressed to hard-of-hearing humans or to learners of new languages.<\/p>\n","protected":false},"author":116,"featured_media":0,"parent":22,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-304","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/pages\/304","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/users\/116"}],"replies":[{"embeddable":true,"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/comments?post=304"}],"version-history":[{"count":9,"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/pages\/304\/revisions"}],"predecessor-version":[{"id":519,"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/pages\/304\/revisions\/519"}],"up":[{"embeddable":true,"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/pages\/22"}],"wp:attachment":[{"href":"https:\/\/members.loria.fr\/SOuni\/wp-json\/wp\/v2\/media?parent=304"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}