B2 Real News

Pictura: Perspective-View Self-Play at Scale for Driving

Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy into a camera-input student, leaves the student imitating decisions its own view cannot justify. Instead, we establish perspective-view self-play as a practical training regime. We introduce Pictura, a GPU-accelerated multi-agent driving simulator that renders each agent's egocentric view at every step, mitigating the representation gap at its source. Pictura sustains up to 500K agent-steps/s (2M images/s) on a single H100. Using Pictura, we train Alberti by self-play with plain PPO. It is the first large-scale driving self-play policy trained directly from perspective images, without privileged observations. Training spans 50B agent steps for ~35M km of driving. It approaches the driving performance of its privileged vectorized counterpart, and transfers zero-shot to Waymo Open Motion Dataset layouts re-rendered in Pictura, where it outperforms privileged vectorized agents. Project page: https://valeoai.github.io/Pictura/

Cover image for Pictura: Perspective-View Self-Play at Scale for Driving
Image: Daily English Reader / Local generated SVG (Project-owned local asset)
5 min read B2

B2 reading

Select any word for its Thai meaning and pronunciation.

0:00 0:00
Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. This assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras. A common fix, distilling the privileged policy into a camera-input student, leaves the student imitating decisions its own view cannot justify. Instead, we establish perspective-view self-play as a practical training regime. We introduce Pictura, a GPU-accelerated multi-agent driving simulator that renders each agent's egocentric view at every step, mitigating the representation gap at its source. Pictura sustains up to 500K agent-steps/s (2M images/s) on a single H100. Using Pictura, we train Alberti by self-play with plain PPO. It is the first large-scale driving self-play policy trained directly from perspective images, without privileged observations.

ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้

Useful phrases from this story

driving policies at scaleCollocation

การขับเคลื่อนนโยบายในระดับ.

From the storySelf-play in simulation produces robust driving policies at scale.

have been made using privilegedCollocation

ได้ทําโดยใช้สิทธิพิเศษ.

From the storyDemonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents.

vectorized observations such as exactCollocation

การสังเกตเห็นแบบเวกเตอร์ เช่น.

From the storyDemonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents.

is solved and introduces aCollocation

ได้แก้ไข และนํามา.

From the storyThis assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras.

deployed agent driving from theCollocation

พนักงานที่ลงทัพขับรถจาก.

From the storyThis assumes that perception is solved and introduces a representation gap with the partial observation of a deployed agent driving from the perspective view of egocentric cameras.

Save & Review

Only words saved from this story appear here.