MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding. To address this, we propose MultiModal Code-Switching (MMCS), a novel pretraining paradigm that provides explicit object-level supervision. Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding. We further develop a scalable data synthesis pipeline to generate a pretraining dataset of 773K samples with accurate object-entity correspondences. Experiments show that MMCS is highly data-efficient: with only 50K samples, it matches or surpasses models trained on 600K image-text pairs. Furthermore, MMCS consistently improves visual grounding and perception capabilities across varying model scales.
C1 reading
Select any word for its Thai meaning and pronunciation.
แปลไทยทั้งบท
โมเดลภาษาขนาดใหญ่หลายแบบ (MLLMs) ที่มีอยู่ปัจจุบันมักจะขึ้นอยู่กับคู่ภาพ-บทความสำหรับการฝึกซ้อมก่อนการสอดคล้องแบบการฝึกซ้อม, การแผนที่การแสดงภาพภาพทั่วโลกเป็นคําอธิบายบทความยาว. แต่ การสอดคล้องระดับภาพนี้ มีปัญหาความไม่ชัดเจนของการอ้างอิง: รูปแบบมีปัญหาในการสรุปความตรงกันระหว่างวัตถุวิวและองค์กรเท็กสตูลหลายอย่าง จากการแสดงภาพทั่วโลก ซึ่งนําไปสู่ความไม่ประสิทธิภาพของข้อมูล และการถอดพื้นที่ทางความหมายที่ไม่ดีที่สุด. เพื่อแก้ปัญหานี้ เราเสนอ MultiModal Code-Switching (MMCS) เป็นแนวคิดใหม่ของการฝึกอบรมก่อนที่ให้การดูแลระดับของวัตถุอย่างชัดเจน.
โดยได้รับแรงบันดาลใจจากปรากฏการณ์ทางภาษาของการเปลี่ยนรหัส MMCS ผสมสายตาและภาษา โดยเปลี่ยนองค์กรทางเท็กสต์กับวัตถุภาพที่เกี่ยวข้องกับมัน โดยบังคับการตั้งพื้นที่ของภาษาสายตาของท้องถิ่น. เราพัฒนา pipeline การสังเคราะห์ข้อมูลที่ทำได้ปรับขนาดได้ เพื่อสร้างเซ็ตข้อมูลก่อนการฝึกอบรมจากตัวอย่าง 773K ด้วยความตรงกันของสิ่งของ-องค์กรที่แม่นยํา. การทดลองแสดงให้เห็นว่า MMCS มีประสิทธิภาพในการใช้ข้อมูลสูง: ด้วยตัวอย่างเพียง 50K, มันตรงกับหรือมากกว่าแบบที่ได้รับการฝึกอบรมจากคู่ภาพ-บทความ 600K.
นอกจากนี้, MMCS ได้ปรับปรุงความทำได้ในการปรับพื้นและการรับรู้ทางการมองเห็นตลอดระดับแบบที่แตกต่างกัน.
ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้
Useful phrases from this story
โมเดลภาษาขนาดใหญ่หลายแบบที่มีอยู่.
From the storyExisting Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions.
แผนที่การแสดงภาพโลก.
From the storyExisting Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions.
ส่งผลให้ข้อมูลไม่มีประสิทธิภาพ และ.
From the storyHowever, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding.
หลักสูตรการฝึกอบรมที่ให้ความชัดเจน.
From the storyTo address this, we propose MultiModal Code-Switching (MMCS), a novel pretraining paradigm that provides explicit object-level supervision.
ได้รับแรงบันดาลใจจากปรากฏการณ์ภาษา.
From the storyInspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding.
Save & Review
Only words saved from this story appear here.