Multimodal AI Lab

School of Electrical Engineering, KAIST

We are looking for motivated students in machine learning, speech processing and computer vision. Please read this page for more information.

Recent highlights

Text-to-Speech

Tan Dat Nguyen et al., “SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS

ICASSP 2026

Audio generation from video

Kang Zhang et al., “Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

NeurIPS 2025

Tactile-driven localization

Seongyu Kim et al., “Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

CVPR 2026

Sign language recognition

Youngjoon Jang et al., “Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues

CVPR 2025

Audio-visual LLM decoding

Chaeyoung Jung et al., “AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

NeurIPS 2025

Lip to speech

Ji-Hoon Kim et al., “From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

CVPR 2025