Pictura: Perspective-View Self-Play at Scale for Driving
Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy into a camera-input student, leaves the student imitating decisions its own view cannot justify. Instead, we establish perspective-view self-play as a practical training regime. We introduce Pictura, a GPU-accelerated multi-agent driving simulator that renders each agent's egocentric view at every step, mitigating the representation gap at its source. Pictura sustains up to 500K agent-steps/s (2M images/s) on a single H100. Using Pictura, we train Alberti by self-play with plain PPO. It is the first large-scale driving self-play policy trained directly from perspective images, without privileged observations. Training spans 50B agent steps for ~35M km of driving. It approaches the driving performance of its privileged vectorized counterpart, and transfers zero-shot to Waymo Open Motion Dataset layouts re-rendered in Pictura, where it outperforms privileged vectorized agents. Project page: https://valeoai.github.io/Pictura/
B2 reading
Select any word for its Thai meaning and pronunciation.
แปลไทยทั้งบท
การเล่นตัวตนในแบบจําลอง สร้างนโยบายขับรถที่แข็งแกร่งในระดับ. การแสดงพฤติกรรมเช่นนั้นได้ถูกทำโดยใช้การสังเกตการณ์แบบเวกเตอร์แบบมีสิทธิพิเศษ เช่น โพสและความเร็วที่แม่นยํา แม้แต่สำหรับตัวแทนที่ปิดตัว. นี่คาดว่าการรับรู้ถูกแก้ไขและแนะนําช่องว่างการแสดงด้วยการสังเกตเห็นบางส่วนของเจ้าหน้าที่ที่ใช้ในการขับรถจากมุมมองของกล้องอารมณ์.
การแก้ไขที่ทั่วไป ทำให้นักเรียนที่ได้รับสิทธิพิเศษเข้าสู่นักเรียนที่ใส่กล้อง ทำให้นักเรียนเลียนแบบการตัดสินใจที่ความคิดของเขาเองไม่ทำได้แก้ไขได้. แทนนั้น เราจัดตั้งการเล่นแบบเห็นจากมุมมองของตนเอง เป็นระบบการฝึกอบรมทางปฏิบัติ. เรานําเสนอ Pictura เป็นเครื่องจําลองขับรถแบบ GPU ที่เร่งเร่งหลายเอเจนต์ ที่ทำให้การมองของแต่ละเอเจนต์เป็นตัวตนในแต่ละขั้นตอน โดยลดช่องว่างในการแสดงตัวแทนที่แหล่งของมัน.
Pictura ทำได้รองรับขั้นตอนตัวแทนได้ถึง 500K (ภาพ 2M/s) บน H100 เดียว. โดยใช้ Pictura เราฝึกอัลเบอร์ตี้ ด้วยการเล่นด้วยตัวเอง ด้วย PPO ธรรมดา. มันเป็นนโยบายการขับขี่ของตัวเองขนาดใหญ่แรกที่ฝึกอบรมโดยตรงจากภาพมุมมอง โดยไม่มีการสังเกตเห็นพิเศษ.
ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้
Useful phrases from this story
การขับเคลื่อนนโยบายในระดับ.
From the storySelf-play in simulation produces robust driving policies at scale.
ได้ทําโดยใช้สิทธิพิเศษ.
From the storyDemonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents.
การสังเกตเห็นแบบเวกเตอร์ เช่น.
From the storyDemonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents.
ได้แก้ไข และนํามา.
From the storyThis assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras.
พนักงานที่ลงทัพขับรถจาก.
From the storyThis assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras.
Save & Review
Only words saved from this story appear here.