Publications

Current and former lab members in bold. Underlined names link to a personal homepage.

2026

  1. Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

    H. Cho, S. Yoo, J. Jang, C. Kim, J. S. Chung

    Conference on Empirical Methods in Natural Language Processing

  2. Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models

    C. Jung, Y. Jang, J. Choi, J. S. Chung

    Findings of Empirical Methods in Natural Language Processing

  3. See & Sniff: Learning Visuo-Olfactory Representations

    S. Kim, S. Lee, H. Ryu, J. S. Chung, A. Senocak

    European Conference on Computer Vision

  4. SCORE: Scaling audio generation using Standardized COmposite REwards

    J. Jung, J. Kim, I. Shin, J. S. Chung

    IEEE Signal Processing Letters

  5. Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction

    D. Kwak, S. Lee, J. S. Chung

    Interspeech

  6. ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

    J. Choi, J. Kim, S. Hu, J. S. Chung

    Interspeech

  7. Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

    H. Cho, J. Jang, C. Kim, J. S. Chung

    Interspeech

  8. MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis

    T. D. Nguyen, S. Bae, J. S. Chung, J. Kim

    Interspeech

  9. A Hierarchical Feature Engineering Framework for Automated Classification of Phonotraumatic and Non-Phonotraumatic Vocal Hyperfunction

    J. Kim, K. Jang, M. Kim, H. Lee

    Interspeech

  10. Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis

    Z. Niu, S. Hu, J. Choi, Y. Chen, P. Chen, P. Zhu, Y. Yang, B. Zhang, J. Zhao, C. Wang, X. Chen

    Interspeech

  11. WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

    H. Lee, T. D. Nguyen, J. Kang, K. Shim

    Interspeech

  12. Inference-Time Scaling for Joint Audio-Video Generation

    J. Jung, K. Rho, I. Shin, J. S. Chung

    Transactions on Machine Learning Research

  13. Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses

    S. Kim, K. Jang, S. Cho, J. S. Chung, H. Kim, S. Yun

    Findings of the Association for Computational Linguistics

  14. Probing Cross-modal Information Hubs in Audio-Visual LLMs

    J. Jung, C. Jung, J. Kim, J. S. Chung

    International Conference on Machine Learning

  15. Deep Understanding of Sign Language for Sign to Subtitle Alignment

    Y. Jang, J. Choi, J. Ahn, J. S. Chung

    IEEE Transactions on Multimedia

  16. Cinematic Audio Source Separation Using Visual Cues

    K. Zhang, S. Lee, A. Senocak, J. S. Chung

    IEEE Conference on Computer Vision and Pattern Recognition

  17. Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

    S. Kim, S. Lee, H. Ryu, J. S. Chung, A. Senocak

    IEEE Conference on Computer Vision and Pattern Recognition

  18. How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?

    A. Senocak, S. Park, T. Oh, J. S. Chung

    IEEE Conference on Computer Vision and Pattern Recognition

  19. Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial Sound

    H. Ryu, J. S. Chung, D. Harwath

    IEEE Conference on Computer Vision and Pattern Recognition

  20. Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

    S. H. Baek, J. Lee, H. Lee, J. W. Cho

    IEEE Conference on Computer Vision and Pattern Recognition

  21. DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

    N. Nguyen, T. Tran, J. Choi, H. Huynh-Nguyen, T. Hy, V. Nguyen

    IEEE Conference on Computer Vision and Pattern Recognition

  22. EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training

    D. Kwak, Y. Jang, S. Kim, J. S. Chung

    IEEE Transactions on Audio, Speech and Language Processing

  23. Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization

    S. Park, A. Senocak, J. S. Chung

    International Journal of Computer Vision

  24. LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling

    D. Kwak, Y. Jang, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  25. SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS

    T. D. Nguyen, J. Kim, J. Kim, S. Choi, Y. Lim, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  26. MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model

    T. H. Pham, T. D. Nguyen, P. T. Tran, J. S. Chung, D. D. Nguyen

    International Conference on Acoustics, Speech, and Signal Processing

  27. Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

    K. Nam, J. Choi, H. Lee, J. Heo, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  28. LAMB: LLM-Based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence

    H. Lee, J. Choi, K. Nam, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  29. UNMIXX: Untangling Highly Correlated Singing Voices Mixtures

    J. Jung, J. Kim, D. Kwak, J. Lee, J. Nam, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  30. FastAV: Efficient Token Pruning for Audio-Visual Large Language Model Inference

    C. Jung, Y. Jang, S. Lee, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

2025

  1. Toward Interactive Sound Source Localization: Better Align Sight and Sound!

    A. Senocak, H. Ryu, J. Kim, T. Oh, H. Pfister, J. S. Chung

    IEEE Transactions on Pattern Analysis and Machine Intelligence

  2. SpoofCeleb: Speech Deepfake Detection and SASV In The Wild

    J. Jung, Y. Wu, X. Wang, J. Kim, S. Maiti, Y. Matsunaga, H. Shim, J. Tian, N. Evans, J. S. Chung, W. Zhang, S. Um, S. Takamichi, S. Watanabe

    IEEE Open Journal of Signal Processing

  3. CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation

    J. Kim, H. Yang, Y. Ju, I. Kim, B. Kim, J. S. Chung

    IEEE Transactions on Audio, Speech and Language Processing

  4. AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

    C. Jung, Y. Jang, J. S. Chung

    Conference on Neural Information Processing Systems

  5. Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

    K. Zhang, T. X. Pham, S. Lee, A. Niu, A. Senocak, J. S. Chung

    Conference on Neural Information Processing Systems

  6. Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision

    C. Zhang, K. Zhang, J. S. Chung, I. S. Kweon, J. Kim, C. Mao

    Conference on Neural Information Processing Systems

  7. Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing

    J. Choi, J. Kim, J. S. Chung

    Findings of Empirical Methods in Natural Language Processing

  8. AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation

    J. Choi, J. Kim, S. Kim, T. Oh, J. S. Chung

    ACM International Conference on Multimedia

  9. VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models

    S. Kim, J. Choi, P. Peng, J. S. Chung, T. Oh, D. Harwath

    International Conference on Computer Vision

  10. MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation

    S. Cho, J. Choi, S. Kim, S. Yun

    International Conference on Computer Vision

  11. InfiniteAudio: Infinite-Length Audio Generation with Consistency

    C. Jung, H. Ki, J. Kim, J. Kim, J. S. Chung

    Interspeech

  12. SEED: Speaker Embedding Enhancement Diffusion Model

    K. Nam, J. Heo, J. Jung, G. Park, C. Jung, H. Yu, J. S. Chung

    Interspeech

  13. Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment

    J. Choi, Z. Niu, J. Kim, C. Wang, J. S. Chung, X. Chen

    Interspeech

  14. The text-to-speech in the wild (TITW) dataset

    J. Jung, W. Zhang, S. Maiti, Y. Wu, X. Wang, J. Kim, Y. Matsunaga, S. Um, J. Tian, H. Shim, N. Evans, J. S. Chung, S. Takamichi, S. Watanabe

    Interspeech

  15. Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes

    H. Ryu, S. Kim, J. S. Chung, A. Senocak

    IEEE Conference on Computer Vision and Pattern Recognition

  16. From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

    J. Kim, J. Choi, J. Kim, C. Jung, J. S. Chung

    IEEE Conference on Computer Vision and Pattern Recognition

  17. Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues

    Y. Jang, H. Raajesh, L. Momeni, G. Varol, A. Zisserman

    IEEE Conference on Computer Vision and Pattern Recognition

  18. Test-Time Augmentation for Pose-invariant Face Recognition

    J. Jung, Y. Jang, J. S. Chung

    IEEE International Conference on Automatic Face and Gesture Recognition

  19. High-Quality Joint Image and Video Compression with Causal VAE

    D. M. Argaw, X. Liu, Q. Zhang, J. S. Chung, M. Liu, F. Reda

    International Conference on Learning Representations

  20. AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

    S. Kim, H. Oh, J. Lee, A. Senocak, J. S. Chung, T. Oh

    International Conference on Learning Representations

  21. ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation

    Z. Li, S. Hu, S. Liu, L. Zhou, J. Choi, L. Meng, X. Guo, J. Li, H. Ling, F. Wei

    International Conference on Learning Representations

  22. V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow

    J. Choi, J. Kim, J. Li, J. S. Chung, S. Liu

    International Conference on Acoustics, Speech, and Signal Processing

  23. LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

    K. Rho, H. Lee, V. Iverson, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  24. VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

    J. Jung, J. Ahn, C. Jung, T. D. Nguyen, Y. Jang, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  25. Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding

    T. D. Nguyen, J. Kim, J. Choi, S. Choi, J. Park, Y. Lee, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  26. AdaptVC: High Quality Voice Conversion with Adaptive Learning

    J. Kim, J. Kim, Y. Choi, T. D. Nguyen, S. Mun, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

2024

  1. Audio Mamba: Bidirectional State Space Model for Audio Representation Learning

    M. H. Erol, A. Senocak, J. Feng, J. S. Chung

    IEEE Signal Processing Letters

  2. Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting

    Y. Kim, J. Jung, J. Park, B. Kim, J. S. Chung

    IEEE Signal Processing Letters

  3. Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding

    J. Woo, H. Ryu, Y. Jang, J. W. Cho, J. S. Chung

    ACM International Conference on Multimedia

  4. VoxSim: A perceptual voice similarity dataset

    J. Ahn, Y. Kim, Y. Choi, D. Kwak, J. Kim, S. Mun, J. S. Chung

    Interspeech

  5. Lightweight Audio Segmentation for Long-form Speech Translation

    J. Lee, S. Kim, H. Kim, J. S. Chung

    Interspeech

  6. ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions

    J. Feng, M. H. Erol, J. S. Chung, A. Senocak

    Interspeech

  7. To what extent can ASV systems naturally defend against spoofing attacks?

    J. Jung, X. Wang, N. Evans, S. Watanabe, H. Shim, H. Tak, S. Arora, J. Yamagishi, J. S. Chung

    Interspeech

  8. Disentangled Representation Learning for Environment-agnostic Speaker Recognition

    K. Nam, H. Heo, J. Jung, J. S. Chung

    Interspeech

  9. FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching

    C. Jung, S. Lee, J. Kim, J. S. Chung

    Interspeech

  10. EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning

    J. Kim, H. Lee, K. Rho, J. Kim, J. S. Chung

    International Conference on Machine Learning

  11. Faces that Speak: Jointly Synthesising Talking Face and Speech from Text

    Y. Jang, J. Kim, J. Ahn, D. Kwak, H. Yang, Y. Ju, I. Kim, B. Kim, J. S. Chung

    IEEE Conference on Computer Vision and Pattern Recognition

  12. Scaling Up Video Summarization Pretraining with Large Language Models

    D. M. Argaw, S. Yoon, F. C. Heilbron, H. Deilamsalehy, T. Bui, Z. Wang, F. Dernoncourt, J. S. Chung

    IEEE Conference on Computer Vision and Pattern Recognition

  13. Towards Automated Movie Trailer Generation

    D. M. Argaw, M. Soldan, A. Pardo, C. Zhao, F. C. Heilbron, J. S. Chung, B. Ghanem

    IEEE Conference on Computer Vision and Pattern Recognition

  14. FreGrad: Lightweight and fast frequency-aware diffusion vocoder

    T. D. Nguyen, J. Kim, Y. Jang, J. Kim, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  15. SlowFast Network for Continuous Sign Language Recognition

    J. Ahn, Y. Jang, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  16. Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification

    H. Heo, K. Nam, B. Lee, Y. Kwon, M. Lee, Y. J. Kim, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  17. Speech Guided Masked Image Modeling for Visually Grounded Speech

    J. Woo, H. Ryu, A. Senocak, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  18. VoxMM: Rich Transcription of Conversations in the Wild

    D. Kwak, J. Jung, K. Nam, Y. Jang, J. Jung, S. Watanabe, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  19. From Coarse To Fine: Efficient Training for Audio Spectrogram Transformers

    J. Feng, M. H. Erol, J. S. Chung, A. Senocak

    International Conference on Acoustics, Speech, and Signal Processing

  20. VoiceLDM: Text-to-Audio Generation with Linguistic Content

    Y. Lee, I. Yeon, J. Nam, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  21. TalkNCE: Improving Active Speaker Detection with Talking-Aware Contrastive Learning

    C. Jung, S. Lee, K. Nam, K. Rho, Y. J. Kim, Y. Jang, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  22. Seeing Through the Conversation: Audio-Visual Speech Separation based on Diffusion Model

    S. Lee, C. Jung, Y. Jang, J. Kim, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  23. Let There Be Sound: Reconstructing High Quality Speech from Silent Videos

    J. Kim, J. Kim, J. S. Chung

    AAAI Conference on Artificial Intelligence

  24. Can CLIP Help Sound Source Localization?

    S. Park, A. Senocak, J. S. Chung

    Winter Conference on Applications of Computer Vision

2023

  1. That's What I Said: Fully-Controllable Talking Face Generation

    Y. Jang, K. Rho, J. Woo, H. Lee, J. Park, Y. Lim, B. Kim, J. S. Chung

    ACM International Conference on Multimedia

  2. Sound Source Localization is All about Cross-Modal Alignment

    A. Senocak, H. Ryu, J. Kim, T. Oh, H. Pfister, J. S. Chung

    International Conference on Computer Vision

  3. FlexiAST: Flexibility is What AST Needs

    J. Feng, M. H. Erol, J. S. Chung, A. Senocak

    Interspeech

  4. Disentangled Representation Learning for Multilingual Speaker Recognition

    K. Nam, Y. Kim, J. Huh, H. Heo, J. Jung, J. S. Chung

    Interspeech

  5. Curriculum learning for self-supervised speaker verification

    H. Heo, J. Jung, J. Kang, Y. Kwon, B. Lee, Y. J. Kim, J. S. Chung

    Interspeech

  6. Self-sufficient framework for continuous sign language recognition

    Y. Jang, Y. Oh, J. W. Cho, M. Kim, D. Kim, I. S. Kweon, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

    Top 3% Paper Recognition

  7. Metric learning for user-defined keyword spotting

    J. Jung, Y. Kim, J. Park, Y. Lim, B. Kim, Y. Jang, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  8. Hindi as a second language: improving visually grounded speech with semantically similar samples

    H. Ryu, A. Senocak, I. S. Kweon, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  9. MarginNCE: Robust Sound Localization with a Negative Margin

    S. Park, A. Senocak, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  10. Advancing the dimensionality reduction of speaker embeddings for speaker diarisation: disentangling noise and informing speech activity

    Y. J. Kim, H. Heo, J. Jung, Y. Kwon, B. Lee, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  11. In search of strong embedding extractors for speaker diarisation

    J. Jung, B. Lee, J. Huh, A. Brown, Y. Kwon, S. Watanabe, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  12. Imaginary Voice: Face-styled Diffusion Model for Text-to-Speech

    J. Lee, J. S. Chung, S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

2022

  1. Signing Outside the Studio: Benchmarking Background Robustness for Continuous Sign Language Recognition

    Y. Jang, Y. Oh, J. W. Cho, D. Kim, J. S. Chung, I. S. Kweon

    British Machine Vision Conference

  2. Augmentation adversarial training for self-supervised speaker representation learning

    J. Kang, J. Huh, H. Heo, J. S. Chung

    Journal of Selected Topics in Signal Processing

  3. Pushing the limits of raw waveform speaker recognition

    J. Jung, Y. J. Kim, H. Heo, B. Lee, Y. Kwon, J. S. Chung

    Interspeech

  4. Spell my name: Keyword boosted speech recognition

    N. Jung, G. Kim, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  5. Multi-scale speaker embedding-based graph attention networks for speaker diarisation

    Y. Kwon, H. Heo, J. Jung, Y. J. Kim, B. Lee, J. S. Chung

    International Conference on Acoustics, Speech, and Signal Processing

  6. AASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks

    J. Jung, H. Heo, H. Tak, H. Shim, J. S. Chung, B. Lee, H. Yu, N. Evans

    International Conference on Acoustics, Speech, and Signal Processing