Speech and Multimodal Interfaces Laboratory

Publications

2025

Zaburdaev Alexander, Ivanko Denis, Ryumin Dmitry. CrossMP-SENet: Transformer-Based Cross-Attention for Joint Magnitude-Phase Speech Enhancement // Lecture Notes in Computer Science / Proc. SPECOM 2025, Szeged, Hungary. 2026. vol. 16188. pp. 174–188.
More
Ryumin Dmitry, Egorova Angelina. MoDeG-Prompt: Depth-Enhanced Multimodal Gesture Recognition with Dynamic Cross-Modal Prompting for Few-Shot Learning // Communications in Computer and Information Science / Proc. 28th International Conference “Internet and Modern Society” IMS 2025. St. Petersburg, 2026. vol. 2672. pp. 347–360.
More
Axyonov Alexandr, Dolgushin Mikhail, Ryumin Dmitry. NeRF-LipSync: A Diffusion Model for Speech-Driven and View-Consistent Lip Synchronization in Digital Avatars // The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences / Proc. 6th International Workshop PSBB 2025. Moscow, 2025. vol. XLVIII-2/W9-2025. pp. 25–31.
More
Ivanko Denis, Ryumin Dmitry. Intelligent System for Automatic Bidirectional Sign Language Translation Based on Recognition and Synthesis of Audiovisual and Sign Speech // The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences / Proc. PSBB 2025. Moscow, 2025. vol. XLVIII-2/W9-2025. pp. 131–136.
More
Ryumina Elena, Ryumin Dmitry, Ivanko Denis. G-MAE: Gesture-aware Masked Autoencoder for Human-Machine Interaction // The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences / Proc. PSBB 2025. Moscow, 2025. vol. XLVIII-2/W9-2025. pp. 241–248.
More
Ryumina Elena, Markitantov Maxim, Axyonov Alexandr, Ryumin Dmitry, Dolgushin Mikhail, Karpov Alexey. Zero-Shot Multimodal Compound Expression Recognition Approach using Off-the-Shelf Large Visual-Language Models // Proc. IEEE/CVF International Conference on Computer Vision (ICCV 2025) Workshops, 9th ABAW Workshop, USA, 2025. pp. 71–79.
More
Dvoynikova A.A., Velichko A.N., Karpov A.A. ENERGI: a multimodal data corpus of interaction of participants in virtual communication // Journal of Instrument Engineering. 2025. Vol. 68, No. 12. pp. 1011–1019.
More
Axyonov A.A. Multimodal Generation of Speech, Facial Expressions, and Gestures in Digital Avatars: Current Methods and Future Prospects // Proceedings of Voronezh State University. Series: Systems Analysis and Information Technologies. 2025. No. 4. pp. 155–182.
More
Smolyaninova A.V., Pavlova T.A., Dorovskikh I.V., Karpov A.A., Dolgushin M.D., Krasnoslobodtseva L.A., Seiku Yu.V., Boldakov D.Yu. Artificial intelligence speech recognition in patients suffering from mild cognitive impairment and dementia // Psychiatry and Psychopharmacotherapy. 2025. No. 4. pp. 50–54.
More
Kiseleva K.O., Kipyatkova I.S. Code-Switching from Karelian to Russian: Morphophonological Aspects of Intra-Word Copying of Russian Units in the Speech of Karelian Speakers // Proceedings of the 11th Interdisciplinary Seminar “Analysis of Conversational Russian Speech” AR3-2025, St. Petersburg. 2025. pp. 43–47.
More
Dolgushin M.D., Guseva D.D., Karpov A.A. Investigation of Methods for Automated Diagnosis of Cognitive Impairments Based on Video Interview Data // Proceedings of the 11th Interdisciplinary Seminar “Analysis of Conversational Russian Speech” AR3-2025, St. Petersburg. 2025. pp. 25–30.
More

2024

Dresvyanskiy D., Markitantov M., Yu J., Kaya H., Karpov A. Multi-modal Arousal and Valence Estimation under Noisy Conditions // IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 2024. pp. 4773-4783.
Ryumina E., Markitantov M., Ryumin D., Kaya H., Karpov A. Zero-Shot Audio-Visual Compound Expression Recognition Method based on Emotion Probability Fusion // IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 2024. pp. 4752-4760.
Dresvyanskiy D., Karpov A., Minker W. A Cross-Multi-modal Fusion Approach for Enhanced Engagement Recognition // Lecture Notes in Computer Science, SPECOM-2024. 2024. vol. 15300. pp. 3-17.
Mamontov D., Zepf S., Karpov A., Minker W. Cross-Cultural Automatic Depression Detection Based on Audio Signals // Lecture Notes in Computer Science, SPECOM-2024. 2024. vol. 15299. pp. 309-323.