ADAPT: Text-Promptable Perception and Diffusion-Predicted Avoidance with MPPI Control on the Polaris GEM e4
Contributed to a ROS 2 team project on the Polaris GEM e4 coupling diffusion-based pedestrian prediction, GPU-batched MPPI planning, and text-conditioned perception goals. The joint predictor was evaluated offline on Argoverse 2; the report also documents vehicle integration.

Technical work
Team project · ROS 2 integration on Polaris GEM e4; prediction evaluated on Argoverse 2
- Contributed to a ROS 2 Humble stack combining a GPU-batched MPPI planner and PACMod2 actuation; the planner optimizes steering and acceleration with a kinematic bicycle model, obstacle costs, and a TTC-based state machine.
- Contributed to single-agent and joint diffusion prediction with cross-agent attention; the joint model reported 0.302 m minADE-20 and 0.532 m minFDE-20 on Argoverse 2 validation using 20 sampled forecasts.
- The team’s predictor ran at approximately 11 ms for eight agents and 20 samples on an RTX 3060; the stack connects 10 Hz prediction to a configured 20 Hz planner with persistent trajectory selection.
- The team stack generated text-prompted object goals through camera–LiDAR projection, point filtering, and DBSCAN, supporting vehicle integration and a planned three-configuration controller/predictor comparison.
Overview
ADAPT is an autonomous driving stack on the Polaris GEM e4 that couples a learning-based pedestrian trajectory predictor with a sampling-based motion planner. The default Stanley lateral controller from the AutoShield baseline is replaced with a Model Predictive Path Integral (MPPI) planner that jointly optimizes steering and acceleration.
The planner consumes multi-step forecasts from a diffusion trajectory predictor in single-agent and joint multi-agent configurations. The system is implemented end-to-end in ROS 2 Humble and runs on the GEM e4, using its Ouster OS1-128 LiDAR, OAK-D LR RGBD camera, Septentrio GNSS/INS, and PACMod2 drive-by-wire interface. A complementary text-promptable perception module supplies open-vocabulary object goals for non-pedestrian targets.
Team Contributions
- GPU-Batched MPPI Planner: Integrated with PACMod2, sharing a high-level safety state machine with the Stanley baseline
- Diffusion Pedestrian Predictors: Single-agent and joint multi-agent diffusion models deployed as ROS 2 nodes with sticky-mode trajectory selection
- Open-Vocabulary Goal Module: Text-conditioned detection-and-LiDAR-fusion goal selection operating in both Gazebo and on the vehicle
- On-Vehicle A/B Comparison: Protocol for evaluating Stanley vs. MPPI with constant-velocity vs. joint-diffusion prediction over identical routes
System Architecture
The runtime stack is organized into four layers covering perception, fusion, prediction, and control, all running as ROS 2 nodes on the vehicle.
Perception & Fusion
- A LiDAR pipeline applies density-based clustering (DBSCAN) to the Ouster point cloud with human-geometry height/width filters, producing pedestrian detections in polar form
- In parallel, an RGBD pipeline runs YOLOv11 on the OAK-D RGB stream and back-projects the bounding-box centroid using the stereo depth image
- LiDAR and camera detections are fused via approximate time synchronization — range weighted toward LiDAR (superior depth accuracy), bearing weighted toward camera (higher angular resolution)
Prediction & Control
- A single inference node consumes the fused pedestrian track and publishes future trajectories plus per-step covariance and time-to-collision estimates; the active prediction model is selected at launch so downstream nodes remain mode-agnostic
- The MPPI node subscribes to either the fused pedestrian estimate (treated as a static obstacle) or the predicted trajectories with covariance, and issues PACMod2 commands
- A high-level state machine emits
CRUISE,SLOW_CAUTION,STOP_YIELD, andCREEP_PASSstates based on time-to-collision and pedestrian motion — a coarse safety overlay independent of the chosen predictor
Diffusion Trajectory Predictor
The trajectory predictor is a joint multi-agent denoiser (JointTrajectoryDenoiser) that adapts the Motion Indeterminacy Diffusion (MID) formulation to a ROS-deployable, scene-level conditional model. All pedestrians in a scene are denoised jointly so each agent's future is conditioned on the presence and history of the others through cross-agent attention.
- Architecture: Six-layer Transformer history encoder (shared weights, joint conditioning on diffusion timestep + ego velocity), three-layer cross-agent interaction encoder, four-layer Transformer decoder with positional embeddings — d=256, 8 heads, FFN 512, dropout 0.1, supports up to 16 agents/scene, 8.18M total parameters
- Training: DDPM noise-prediction objective (100 diffusion timesteps, cosine β schedule) with DDIM inference (10 steps, K=20 samples/scene); AdamW, cosine LR 2×10-4→10-5, 200 epochs, batch size 16, mixed precision (AMP fp16), EMA decay 0.999
- Data: Pre-trained on the pedestrian split of Argoverse 2 Motion Forecasting (49,202 training / 6,100 validation scenes, resampled to 4 Hz), then fine-tuned on ETH/UCY pedestrian benchmark scenes plus a procedurally generated corpus of synthetic pedestrian arcs (S-bends, spirals, U-turns, accelerating arcs) to densify long-tail tight-turn motions
- Augmentation: Per-scene random rotation (±15°), per-agent history dropout (0–20%), per-scene translation jitter (±0.1m), plus observation-noise augmentation (σpos=0.05m, σvel=0.2 m/s) during fine-tuning
MPPI Motion Planner
Sampling-based model predictive control reformulates receding-horizon trajectory optimization as a Monte Carlo estimation problem: at each control tick the planner draws K perturbed control sequences, forward-simulates each through a rear-axle kinematic bicycle model, scores them with a running + terminal cost, and returns a cost-weighted (softmax) average as the nominal control.
- Running cost decomposes into path-tracking (cross-track to nearest reference waypoint), velocity tracking, steering-stability, pedestrian-avoidance, and static-cone-avoidance terms
- Trajectory-conditioned pedestrian cost: when the diffusion predictor is active, each of the M predicted pedestrian trajectories contributes an isotropic Gaussian repulsion around its position at the corresponding rollout step, plus a finite clearance penalty; in legacy constant-velocity mode the predicted position propagates linearly with growing positional covariance over the lead time
- Defaults: K=100 rollouts, H=100-step horizon, Δt=0.1s, 20 Hz control rate, vref=4 m/s, a∈[-1.0, 2.0] m/s², δmax=0.61 rad
- The effective sample size (ESS) is monitored at runtime as a diagnostic for noise-scale tuning
Text-Promptable Perception
For non-pedestrian targets (e.g., "stop sign," "red cone"), a text-conditioned camera+LiDAR fusion module emits a single navigation goal pose from a runtime-supplied prompt, using either a YOLO-World bounding box or a LangSAM (Grounding-DINO + SAM 2.1 Hiera-Small) dense mask as a proxy region.
- Pipeline: projective frustum clip of the LiDAR cloud into the prompted mask → height-band filtering & statistical outlier rejection → DBSCAN clustering → cluster selection by 2D reprojection distance to the detection target
- Ray-cast fallback: when no cluster survives filtering, the system unprojects the detection centroid to a fixed estimated distance (15m) and flags the goal as estimated, to be replaced once the vehicle closes the distance
- Goal latching: a 3.0s goal-hold buffer accumulates candidate poses and emits the elementwise median once ≥5 samples are collected, suppressing per-frame jitter and short detection dropouts
- Validated in both the Gazebo simulator and on the vehicle's RTX 3060 within the runtime latency budget
Results & Performance
| Metric | Value |
|---|---|
| minADE-20 (minimum average displacement error over 20 samples) | 0.302 m |
| minFDE-20 (minimum final displacement error over 20 samples) | 0.532 m |
| Miss rate @ 2m | 5.66% |
| Inference latency (M=8, K=20) | ~11 ms (on RTX 3060) |
| Latency budget | < 30 ms |
Evaluated on the Argoverse 2 validation split with K=20 DDIM samples per scene. The full MPPI + joint-diffusion stack runs end-to-end on the GEM e4, with hardware bring-up (Ouster OS1-128, OAK-D LR, Septentrio GNSS/INS, PACMod2) integrated under a unified sensor-initialization launch.
The displacement and miss-rate metrics evaluate sampled forecasts on Argoverse 2, not driving success. Approximately 11 ms measures predictor inference for eight agents and 20 samples; it excludes sensing and vehicle actuation. Prediction runs at 10 Hz, while the planner is configured at 20 Hz. A planned three-configuration vehicle comparison covers Stanley with constant velocity, MPPI with constant velocity, and MPPI with joint diffusion.
Technologies Used
Team
Polaris GEM e4 Autonomy Team — CS 588: Autonomous Vehicle Systems, Group 10
University of Illinois Urbana-Champaign — Spring 2026
The architecture and results described here are team-level contributions.