Text-to-Speech
Tan Dat Nguyen et al., “SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS”
ICASSP 2026
School of Electrical Engineering, KAIST
Tan Dat Nguyen et al., “SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS”
Kang Zhang et al., “Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation”
Seongyu Kim et al., “Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions”
Youngjoon Jang et al., “Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues”
Chaeyoung Jung et al., “AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding”
Ji-Hoon Kim et al., “From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech”