ReportGem ReportGem

Academic paper

Projector Is All You Train

Authors: Nyx Iskandar, Saathvik Selvan, Slater VictoroffPublished: 2026-08-20Paper ID: 2608.19726Category: cs.CLLicense: CC BY 4.0

Abstract

The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuning the backbone of an MLLM is necessary to adapt it to a new modality. Through experiments on 3D MLLMs, we find that training only the projector is sufficient to achieve strong multimodal performance relative to existing baseline models and our jointly trained MLLMs with the same encoder and backbone. We also show that joint training leads to undesirable drift in existing capabilities of the language model, which projector-only training avoids by definition. Furthermore, projector-only training has approximately twice the training sample throughput of joint training. We validate our findings across different language model backbones via 3D classification and captioning benchmarks as well as standard benchmarks evaluating language, vision, and spatial reasoning capabilities.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader