ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existing methods fail to alter the rigid, fixed computation allocation between the Vision Transformer and the Large Language Model components, limiting task-specific optimization. To address this, we introduce the Parallel Vision-Language (ParVL) scaling framework for MLLMs, which scales parallel computation by reusing the existing ViT and LLM backbone parameters across multiple vision and language branches. This framework raises a central question: given a fixed backbone parameter budget, how should additional shared-backbone computation be allocated between the vision and language modalities? We instantiate each parallel computational stream with branch-specific prefix parameters over a shared backbone, and train the entire model end-to-end via full-parameter supervised fine-tuning on roughly 13B tokens. We systematically study the computation-allocation trade-off between the ViT encoder and LLM decoder. ParVL improves overall multimodal performance over same-recipe single-branch baselines, and the best evaluated vision--language allocation varies across tasks. Code is available at https://github.com/YangYangGirl/ParVL.
C1 reading
Select any word for its Thai meaning and pronunciation.
แปลไทยทั้งบท
กลยุทธ์การปรับขนาดที่มีอยู่สำหรับโมเดลภาษาขนาดใหญ่ (MLLMs) มูลติโมเดลมวลชนโดยทั่วไปขยายตัวละเอียดแบบหรือการคํานวณการสรุปตามลําดับ โดยเกิดความจําที่สําคัญหรืออออฟเฮดแลนตี้. ที่สําคัญกว่านั้น วิธีการที่มีอยู่ส่วนใหญ่ไม่ทำได้เปลี่ยนแปลงการจัดสรรการคํานวณที่แข็งแรงและคงที่ระหว่าง Vision Transformer และองค์ประกอบ Big Language Model. เพื่อแก้ปัญหานี้ เรานําเสนอกรอบการปรับขนาด Parallel Vision-Language (ParVL) ให้กับ MLLM ซึ่งปรับระดับการคํานวณแบบร่วมกัน โดยใช้ ViT และ LLM ที่มีอยู่อีกครั้ง. ปรามาตรฐานกระดูกสันหลัง ผ่านสาขาสายตาและภาษาหลายสาขาน.
ระเบียบนี้ทำให้เกิดคําถามสําคัญอย่างหนึ่ง: เมื่อมีงบประมาณปารามีเมตรกระดูกสันหลังที่ตั้ง, การคิดคํานวณแบบแบ่งปันกระดอกสันหลังเพิ่มเติมควรถูกจํากัดอย่างไรระหว่างสายตาและภาษา. เราฉากฉากลําไหลคํานวณแบบตรงกันแต่ละสาย โดยมีปารามีเตอร์อนุพันธ์เฉพาะสาขาบนกระดูกสันหลังที่แบ่งปัน และฝึกต้นแบบทั้งมวลจากปลายไปยังปลายผ่านการปรับระดับเต็มที่. ในเทคนิคประมาณ 13B. เราศึกษาระบบการเทคโนโลยีการจําหน่ายการคํานวณ ระหว่าง ViT encoder และ LLM decoder.
พาร์วีเอล ปรับปรุงการทำงานรวมของมัลติโมเดลเหนือเส้นฐานแบบสูตรเดียวกัน และการประเมินภาพที่ดีที่สุด - การจัดสรรภาษาแตกต่างกันระหว่างงาน. โค้ดทำได้ใช้ได้ที่.
ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้
Useful phrases from this story
กลยุทธ์การปรับขนาดที่มีอยู่สําหรับ Multimodal.
From the storyExisting scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead.
เกิดความจําหรือความช้ามาก.
From the storyExisting scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead.
วิธีการที่มีอยู่ไม่เปลี่ยนแปลง.
From the storyMore importantly, most existing methods fail to alter the rigid, fixed computation allocation between the Vision Transformer and the Large Language Model components, limiting task-specific optimization.
การจําหน่ายการคํานวณที่ตั้งระหว่าง.
From the storyMore importantly, most existing methods fail to alter the rigid, fixed computation allocation between the Vision Transformer and the Large Language Model components, limiting task-specific optimization.
การจํากัดการอุดมสมรรถนะที่เฉพาะงาน.
From the storyMore importantly, most existing methods fail to alter the rigid, fixed computation allocation between the Vision Transformer and the Large Language Model components, limiting task-specific optimization.
Save & Review
Only words saved from this story appear here.