Yake WEI/卫雅珂

My research focuses on multimodal learning, with particular emphasis on understanding the internal mechanisms of heterogeneous modality integration and enhancing the external perception and cognition of multimodal intelligence.

I recently earned my Ph.D. in Artificial Intelligence from Renmin University of China (RUC) in June 2026, under the supervision of Prof. Di Hu. Before that, I received my B.E. in Computer Science and Technology from University of Electronic Science and Technology of China (UESTC) in June 2021. Grateful to everyone who supported me along the way.

Email  /  Google Scholar  /  GitHub

profile photo
🔥 Research Interests

My research focuses on how multimodal systems learn from, perceive, and reason over heterogeneous signals such as vision, audio, and language. My work spans two complementary directions:

Internal Learning Dynamics. Leveraging cognitive neuroscience and mathematical tools to investigate the coordination, integration, and interaction of heterogeneous multimodal information within unified learning systems.

External Perception & Cognition. Advancing the perception and cognition of multimodal intelligence to support coherent understanding, reasoning, and interaction across complex physical and digital environments.

Research roadmap for multimodal learning dynamics, perception, and cognition
📰 News

[2026-07] Graduated from Renmin University of China!

[2026-07] Awarded the Wu Yuzhang Scholarship!

[2026-06] Attended CVPR 2026 Doctoral Consortium in Denver!

[2026-06] Gave a talk at CVPR 2026 Sight and Sound Workshop!

[2026-03] One paper was accepted by CVPR, thanks to all co-authors!

[2025-10] Organized the first online workshop for the BML community! Thanks to TechBeat and all speakers!

[2025-09] MokA was accepted by NeurIPS as an Oral paper! Thanks to all co-authors!

[2025-06] Released our new PEFT pipeline for MLLMs, MokA! [Project page]

[2025-06] Attended student workshop @ VALSE 2025 in Zhuhai!

[2025-05] One paper was accepted by ICML, thanks to all co-authors!

[2025-03] Two paper were accepted by CVPR, thanks to all co-authors!

[2025-02] Awarded the Baidu Scholarship (10 Ph.D students worldwide)!

[2024-12] Awarded the China National Scholarship for Ph.D student!

[2024-12] Attended the Global PhD Gathering @ 2024 Pujiang AI Conference in Shanghai!

[2024-11] Gave a talk about "Balanced multimodal learning" @ Virginia Tech! Thanks to Prof. Chris Thomas for the invitation!

[2024-09] One paper was accepted by T-PAMI, thanks to all co-authors!

[2024-07] Gave a talk about "Balanced multimodal learning" @ TechBeat! [Recording]

[2024-07] One paper was accepted by ECCV, thanks to all co-authors!

[2024-05] We released a survey on multimodal fusion with low-quality multimodal data! [arXiv]

[2024-05] One paper was accepted by ICML, thanks to all co-authors!

[2024-02] One paper was accepted by CVPR, thanks to all co-authors!

[2024-01] One paper was accepted by ICLR, thanks to all co-authors!

[2023-12] Started visiting at Human Sensing Lab @ CMU!

[2023-10] One paper was accepted by Pattern Recognition, thanks to all co-authors!

[2022-08] We released a survey about recent advances in audio-visual learning! [website]

[2022-05] Gave a talk @2022 BAAI Conference. Slides are available here!

[2022-03] Two papers were accepted by CVPR, thanks to all co-authors!

[2021-12] One paper was accepted by T-PAMI, thanks to all co-authors!

[2021-06] Graduated from University of Electronic Science and Technology of China!

[Expand]

🏆 Selected Honors

2026 — Wu Yuzhang Scholarship (RUC’s highest student honor; awarded to 10 graduates annually).

2026 — Outstanding Graduate of Renmin University of China.

2026 — Selected to participate in the CVPR 2026 Doctoral Consortium.

2024 — Baidu Scholarship (10 PhD students worldwide).

2024 — National Scholarship for PhD Students (China’s highest student honor).

2021 — Outstanding Graduate of Sichuan Province (highest graduate honor awarded by Sichuan Province).

2021 — Outstanding Graduate of the University of Electronic Science and Technology of China.

Student Travel Award · NeurIPS 2025 · CVPR 2024

📑 Selected Studies (∗ equal contribution)
clean-usnob Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning

Yake Wei, Yuan Wang, Fengyun Rao, Jing Lyu, Di Hu

Preprint, 2026
arXiv

Use pixel grounding to enhance MLLM's fine-grained visual perception.

clean-usnob Information-Theoretic Decomposition for Multimodal Interaction Learning

Zequn Yang, Yake Wei, Haotian Ni, Zhihao Xu, Di Hu

CVPR, 2026
arXiv

Distinguish and strengthen multimodal interactions within the data.

clean-usnob MokA: Multimodal Low-Rank Adaptation for MLLMs

Yake Wei, Yu Miao, Dongzhan Zhou, Di Hu

NeurIPS, 2025   (Oral Presentation, 1.46% of accepted papers)
Project page

A new PEFT pipeline for MLLMs, ensuring both uni-/cross-modal adaptation.

clean-usnob RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

Haotian Ni, Yake Wei, Hang Liu, Gong Chen, Chong Peng, Hao Lin, Di Hu

ICML, 2025
arXiv / code

Rebalance and revive cooperation between modalities in Transformer.

clean-usnob Adaptive Unimodal Regulation for Balanced Multimodal Information Acquisition

Chengxiang Huang*, Yake Wei*, Zequn Yang, Di Hu

CVPR, 2025
arXiv / code

Analyze and modulate information acquisition process during multimodal training.

clean-usnob Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception

Ruotian Peng, Haiying He, Yake Wei, Yandong Wen, Di Hu

CVPR, 2025
arXiv / code

Generate high-quality image caption by divide-then-aggregate strategy.

clean-usnob On-the-fly Modulation for Balanced Multimodal Learning

Yake Wei, Di Hu, Henghui Du, Ji-Rong Wen
P.S. Thanks the valuable help from Zequn Yang

T-PAMI, 2024
arXiv / code

Analyze and modulate imbalanced unimodal learning from both feed-forward and back-propagation stage.

clean-usnob Multimodal Fusion on Low-quality Data: A Comprehensive Survey

Qingyang Zhang, Yake Wei, Zongbo Han, Huazhu Fu, Xi Peng, Cheng Deng, Qinghua Hu, Cai Xu, Jie Wen, Di Hu, Changqing Zhang

Information Fusion, 2026
arXiv / awesome list

A systematic survey about fusion of low-quality multimodal data.

clean-usnob Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning

Meng Shen, Yake Wei, Jianxiong Yin, Deepu Rajan, Di Hu, Simon See

ACM MM Asia, 2024
paper

Improve the quality of selected multimodal data pairs in active learning.

clean-usnob Diagnosing and Re-learning for Balanced Multimodal Learning

Yake Wei, Siwei Li, Ruoxuan Feng, Di Hu

ECCV, 2024
arXiv / code

Dynimically re-initialization to enhance both worse-/well-learnt modalities.

clean-usnob MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance

Yake Wei, Di Hu

ICML, 2024
arXiv / code

Solve conflicts between multimodal and unimodal gradients.

clean-usnob Enhancing Multimodal Cooperation via Sample-level Modality Valuation

Yake Wei, Ruoxuan Feng, Zihe Wang, Di Hu

CVPR, 2024
arXiv / code

Observe and improve the fine-grained cooperation between modalities.

clean-usnob Quantifying and Enhancing Multi-modal Robustness with Modality Preference

Zequn Yang, Yake Wei, Ce Liang, Di Hu

ICLR, 2024
arXiv / code

Analyze essential components for multimodal robustness and delve into the limitations imposed by modality preference.

clean-usnob Geometric-inspired graph-based Incomplete Multi-view Clustering

Zequn Yang, Han Zhang, Yake Wei, Zheng Wang, Feiping Nie, Di Hu

Pattern Recognition, 2023
paper / code

Conduct geometric analyses to mitigate missing views in weight aggregation.

clean-usnob Learning in Audio-visual Context: A Review, Analysis, and New Perspective

Yake Wei, Di Hu, Yapeng Tian, Xuelong Li

Preprint, 2022
arXiv / website / awesome list

A systematical survey about the audio-visual learning field.

clean-usnob Balanced Multimodal Learning via On-the-fly Gradient Modulation

Xiaokang Peng*, Yake Wei*, Andong Deng, Dong Wang, Di Hu

CVPR, 2022   (Oral Presentation)
arXiv / code

Alleviate imbalance in multimodal learning via dynamic gradient modulation.

clean-usnob Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Guangyao Li*, Yake Wei*, Yapeng Tian*, Chenliang Xu, Ji-Rong Wen, Di Hu

CVPR, 2022   (Oral Presentation)
arXiv / project page

Promote AVQA task and proposes MUSIC-AVQA dataset.

clean-usnob Class-aware Sounding Objects Localization via Audiovisual Correspondence

Di Hu, Yake Wei, Rui Qian, Weiyao Lin, Ruihua Song, Ji-Rong Wen

T-PAMI, 2021
arXiv / project page

Discriminative sounding objects localization.

🎙️ Selected Talks

Organized Activities

“Let's Talk about Balanced Multimodal Learning” 2025 — BML Workshop

Invited Presentations

“Vibe Coding and Anything” 2026 — Fudan University, AI Agents for Nuclear Physics Workshop
2026 — The Chinese University of Hong Kong (Shenzhen), Academic Seminar

“Unified Audio-visual Scene Understanding” 2026 — CVPR, Sight and Sound Workshop

“Vibe Coding: From Autocomplete to CLI Agent” 2026 — TechBeat
2026 — Renmin University of China, Gaoling School Academic Seminar

“Balanced Multimodal Learning” 2025 — VALSE Student Workshop
2025 — Peking University, CoRe
2024 — Global PhD Gathering, Pujiang AI Conference
2024 — Virginia Tech
2024 — TechBeat

“Exploration of Audio-visual Scene Understanding and Multimodal Learning Mechanisms” 2022 — BAAI Conference

Conference Oral Presentations

“MokA: Multimodal Low-Rank Adaptation for MLLMs” 2025 — NeurIPS — Oral Presentation

“Balanced Multimodal Learning via On-the-fly Gradient Modulation” 2022 — CVPR — Oral Presentation

“Learning to Answer Questions in Dynamic Audio-Visual Scenarios” 2022 — CVPR — Oral Presentation



Updated in July 2026