Readings
Week 2
- Paper 1: Foundations & Recent Trends in Multimodal Machine Learning Definitions, Challenges, & Open Questions - Sections 2 and 3 in particular
- Paper 2: Representation Learning: A Review and New Perspectives - Sections 1-3, 6-8, 11
- Paper 3: Experience Grounds Language
- Paper 4: Pragmatics in Language Grounding
Week 3
Unimodal representations
- Paper 1: SimCSE: Simple Contrastive Learning of Sentence Embeddings
- Paper 2: LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
- Paper 3: Visualizing and Understanding Convolutional Networks
- Paper 4: Masked Autoencoders As Spatiotemporal Learners
Multimodal Representations
- Paper 5: Learning Transferable Visual Models From Natural Language Supervision
- Paper 6: On Deep Multi-View Representation Learning: Objectives and Optimization
- Paper 7: Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Paper 8: Sigmoid Loss for Language Image Pre-Training
Week 5
- ImageBind: One Embedding Space To Bind Them All — CVPR 2023 — 2,304 citations: aligns six modalities within a shared embedding space
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models — ICML 2023 — 13,904 citations: an approach for bridging vision encoders and language models
- Understanding the Emergence of Multimodal Representation Alignment — ICML 2025 — 30 citations: examines when multimodal alignment emerges
- OneLLM: One Framework to Align All Modalities with Language — CVPR 2024 — 312 citations: extends alignment beyond image and text
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment — ICLR 2024 — 537 citations: uses language as the semantic anchor for aligning multiple modalities