AdaptFM

Efficient Multimodal LMs Adaptive Inference

Towards Efficient Multimodal Intelligence

Hongxu (Danny) Yin

This talk will provide an overview of recent advances in multimodal foundation models. It will begin by discussing the training of large-scale multimodal models that jointly learn from audio, vision, and language, highlighting key challenges and design principles for unified representation learning. The talk will then explore the adaptation of these models to embodied settings through robotic action tuning, demonstrating how foundation models can be specialized for decision making and control. Finally, it will present recent work on efficient multimodal inference through token focusing, a technique that selectively allocates computation to the most informative visual tokens, reducing computational cost by up to two orders of magnitude while enabling reasoning over long-context visual inputs of up to 1,000 video frames and 4K-resolution imagery. Together, these topics illustrate the progression from building general-purpose multimodal models to adapting and scaling them for practical, compute-efficient real-world applications.

Overview Program