GraphVid: Interactive Graph-Controllable Video Generation
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce $\textbf{GraphVid}$, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate $\textbf{GraphVid-Bench}$, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.
C1 reading
Select any word for its Thai meaning and pronunciation.
แปลไทยทั้งบท
การสร้างวีดีโอที่ทำได้ควบคุมได้ยังคงมีความยากลําบาก เพราะความยากลําบากในการกําหนดการปฏิสัมพันธ์หลายสิ่งอย่างแม่นยํา โดยใช้คําสั่งข้อความหรือการเข้าควบคุมการเคลื่อนไหวที่จํากัดการเคลื่อนไหวของพิกเซลโดยหลัก. ในการปฏิบัติงาน การควบคุมโดยใช้เส้นทางมักจะต้องการให้ผู้ใช้วาดเส้นทางที่แม่นยําสำหรับหลายวัตถุ ซึ่งปรับขนาดไม่ดีกับความซับซ้อนของฉาก และกลายเป็นไม่ชัดเจนภายใต้การปิดหรือการเชื่อมโยง. เพื่อให้การควบคุมหลายหัวข้อมีความยืดหยุ่น แต่แม่นยํา เรานําเสนอ $\textbf{GraphVid}$ รูปแบบการสร้างภาพเป็นวีดีโอแบบอัปกรณ์ที่ทำให้การควบคุมปฏิสัมพันธ์ผ่านกราฟการปฏิสัมพันธ์ที่มีโครงสร้าง.
เรายังจัดทำ $\textbf{GraphVid-Bench}$ เป็นชุดข้อมูลวิดีโอที่เน้นการปฏิกิริยาขนาดใหญ่ พร้อมคํานวณทางสัมพันธ์ที่มีโครงสร้าง เพื่อให้ทำได้ฝึกอบรมแบบการสร้างวิดีโอที่มีความรู้เกี่ยวกับการปฏิกิริยา. แม้จะใช้ข้อมูลการฝึกอบรมน้อยกว่ามาก และมีปารามีที่ทำได้ฝึกอบรมได้น้อยกว่าวิธีการควบคุมการเคลื่อนไหวก่อนหน้านี้ GraphVid ส่งผลให้การควบคุมและคุณภาพวีดีโอที่แข็งแกร่ง. เมื่อเทียบกับ Motion-I2V, GraphVid ลด FID ถึง 39.9% และ FVD ถึง 37.6%, ขณะที่ปรับปรุง PSNR (9.87=>15.98) และ SSIM (0.38=>0.61).
ผลงานของเราแสดงให้เห็นถึงศักยภาพของอินเตอร์เฟชส์สาระที่มีโครงสร้างเป็นพาราไดม์ที่แข็งแกร่งสำหรับการสร้างวีดีโอที่ทำได้ควบคุมได้.
ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้
Useful phrases from this story
ปัญหาเพราะความยากลําบาก.
From the storyControllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement.
สpecifying precise multi-object interactions โดยใช้.
From the storyControllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement.
การควบคุมที่ขึ้นอยู่กับการใช้งานบ่อยครั้งต้องใช้ผู้ใช้งาน.
From the storyIn practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap.
กราฟการปฏิกิริยาที่มีโครงสร้าง.
From the storyTo enable flexible yet precise multi-subject control, we introduce $\textbf{GraphVid}$, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs.
แนะนําทางการสัมพันธ์ที่มีโครงสร้างเพื่อให้.
From the storyWe further curate $\textbf{GraphVid-Bench}$, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models.
Save & Review
Only words saved from this story appear here.