← All projects
ManipulationSpring 2026

Contact-Gated Residual RL for Precision USB Insertion with VLA-Guided Trajectory Planning

Developed the MuJoCo environment, SAC residual training, curriculum, and ROS 2 deployment package for contact-gated USB insertion. Training spans approximately 21.6 million steps; the UR5e demonstration uses a scripted approach, while OpenVLA was evaluated in simulation.

Residual RLSACVLAOpenVLAMuJoCoROS2UR5eSim-to-RealContact-RichDomain Randomization

Residual RL / USB insertion
Residual RL / USB insertion · Watch on YouTube

Technical work

MuJoCo training · OpenVLA in simulation · UR5e demo with scripted approach

  • Trained a six-joint Soft Actor-Critic policy for USB insertion in MuJoCo over approximately 21.6 million environment steps, tightening the cavity success criterion from 1.80× to 1.0625× through a staged curriculum.
  • Implemented a latched contact handoff to prevent repeated controller switching when the plug briefly loses contact; revised the depth reward to address policies that hovered at the port rim.
  • Built a 48-dimensional observation with joint state, port-relative error, force/torque history, and previous actions; randomized port pose, friction, sensor noise, and control latency during training.
  • Developed the ROS 2 deployment package for UR5e and Robotiq 2F-85, supporting a 100 Hz command loop and a real-arm insertion demonstration with a scripted approach.

Overview

Vision-Language-Action (VLA) models like OpenVLA generalize remarkably well across tabletop manipulation tasks, but their joint-target outputs sit several millimeters off the goal axis once a peg touches its mating surface — roughly an order of magnitude larger than the 0.3 mm clearance budget of a USB-A connector.

This project bridges that gap with a contact-gated residual SAC policy: the VLA (or scripted base) owns the free-space approach, and the residual takes over completely at first contact, handling the sub-millimeter compliance needed to slip the plug past the chamfer and seat it fully. The residual is small, trained entirely in simulation, and deploys as a ROS 2 package on a UR5e + Robotiq 2F-85 platform.

A separate stereo CV port-pose tracker (pink-tape fiducials, epipolar pairing, one-shot ChArUco camera-to-base calibration) feeds a pinocchio QP-IK pick-and-transport pipeline, decoupled from but fully compatible with the residual deploy path.

The Core Idea: Latching Contact Gate

The key design insight is a one-way latching handoff. Earlier approaches tracked contact force in real time and toggled the residual on/off, producing a bang-bang contact-bounce loop: residual pushes → plug briefly loses contact → gate snaps shut → base policy drifts arm upward → contact resumes → repeat.

Switching to a latch that fires once and never releases ends the loop. Once &lVert;F&rVert; > 2 N, the base policy is dropped for the rest of the episode and the residual drives alone from the static pre-insertion pose:

  • g = 0 (approach): q_cmd = q_base  —  VLA / scripted trajectory drives free-space approach
  • g = 1 (insertion): q_cmd = q_pre_insert + q_res  —  residual-only, latched for the rest of the episode

This separation makes the residual's training distribution clean: it never has to learn to share control mid-episode, and the upstream VLA is free to use any joint-target interface without disturbing the downstream residual.

Control Pipeline / 100 Hz Command Loop

The command loop is distinct from model inference. OpenVLA is a simulation-only base-policy option; the hardware demonstration uses the scripted base.

  • 1. Perception: 48-D observation packed from joint state (q, q̇), bias-corrected wrist F/T (F, T), TCP pose error vs. port (Δxyz, Δrpy), previous raw action, previous residual, and a 2-step F/T history (+12 dims)
  • 2. Base Policy: OpenVLA-7B (sim only) or scripted 3-waypoint joint-space trajectory (q_home → q_pre_insert → q_descent_end) at 100 Hz
  • 3. Jacobian DLS-IK: Maps VLA end-effector delta to joint space via damped least squares (λ = 0.1); no-op for the scripted base which already emits joint targets
  • 4. Residual SAC: MLP [256, 256, 256] → 6-D action in [−1,1], scaled by per-joint clip [0.025, 0.025, 0.030, 0.040, 0.040, 0.050] rad, EMA-smoothed (α = 0.8)
  • 5. Latching Handoff: Contact-force / range trigger flips the one-way gate; fused command pushed through 1-step latency FIFO to MuJoCo or ROS 2 /forward_position_controller

Training: Two-Stage Curriculum & Domain Randomization

Stage 1 — Cavity Tightening (S1 → S5a, ~20.7M steps): Starts with --disable_base_policy (residual learns the full descent) and a generous 1.80× cavity inflation. Each phase tightens the success criterion toward a near-flush 1.0625×, progressively adding scripted-base noise (σ from 0 → 0.01 rad/joint). This forces the residual to commit deeper rather than hovering at the rim.

Stage 2 — DR Augmentation (DRA2–DRA3, ~1M steps): Holds cavity tight and augments with F/T sensor noise and increased base-policy noise (σ = 0.015 rad/joint), targeting the noisy upstream stream a real VLA would produce.

Domain randomization axes active throughout: port pose (±3 mm xy, ±3° yaw), joint friction (×[0.70, 1.30]), position-gain (×[0.95, 1.05]), F/T bias (±0.5 N / ±0.05 Nm), plug spawn pose (xy ±15 mm, z ±7.5 mm), joint encoder noise, F/T sensor noise, 1-step action latency.

The total is approximate: one roughly 500,000-step phase is estimated because its logs were not retained. Training steps count simulation interactions, not successful insertion trials.

Key Contributions

  • Latching Activation Gate: One-way handoff that prevents switching back to the approach controller during an episode; base policy fully owns approach, residual fully owns insertion — no shared mid-episode control
  • Per-Joint Action Clip: Bounded per-joint residual offsets, with tighter limits on proximal joints (0.025 rad) and larger distal offsets (0.050 rad), narrowing the SAC action range
  • F/T History Augmentation: Added two prior force/torque samples (+12 dimensions); qualitative rollouts showed fewer under-corrections with noisy approach commands
  • Base-Policy Noise Curriculum: Scheduled joint-target Gaussian noise (σ from 0 to 0.015 rad/joint) during training to expose the residual to imperfect approach commands
  • Stereo CV Tracker: Pink-tape fiducials, epipolar pairing, one-shot ChArUco camera-to-base calibration; shared between QP-IK pick-and-transport and residual deploy path

Reward Design

11 default-on reward terms cover pose proximity, orientation, depth, force shaping, action smoothness, residual magnitude, time cost, and terminal success bonus:

  • Pose Proximity: Exponential bonus exp(−|e|/τ) with τ = 2 mm — steep gradient near the port, flat tail outside
  • Depth Bonus: R_depth × (d/d*)^β with β = 1.5 — superlinear exponent critical to avoid the "hover at the rim" local minimum that β = 1 produces
  • Force Shaping: Linear penalty above 10 N deadband + tent-shaped bonus around 3 N target contact force + per-step bonus for port-wall contact (prevents lateral skating)
  • Action Smoothness: Penalties on 1-step delta (spikes) and windowed std-dev over 5 steps (sustained chatter)
  • Terminal Bonus: +5×10⁴ on successful seating (plug tip at depth ≥ 10.5 mm inside inflated cavity)

Technologies Used

MuJoCo Stable-Baselines3 SAC OpenVLA-7B ROS 2 UR5e Robotiq 2F-85 Gymnasium Pinocchio QP-IK ChArUco Calibration Intel D435i Stereo CV Domain Randomization PyTorch Python

Results & Qualitative Findings

  • Residual closes the last-millimeter gap: The report documents a representative real-arm seating demonstration with a scripted approach and learned insertion controller; the base brings the plug into the contact-onset zone, and the residual handles the sub-millimeter compliance needed to slip past the chamfer
  • F/T history matters: Without the 2-step F/T history, the policy under-corrects when the upstream prior is noisy
  • Base-noise curriculum matters: Without it, the policy overfits to the clean scripted base and degrades immediately when noise is injected at inference
  • OpenVLA simulation integration: The residual SAC trained against the scripted base transferred directly when OpenVLA replaced the scripted base in simulation, because both share the same handoff contract: base delivers plug to contact-onset zone, residual takes over
  • CV tracker shared cleanly: Same stereo triangulated port pose feeds both the QP-IK pick-and-transport pipeline and the residual's latched handoff reference

Full bucketed port-distance evaluation (100 episodes per bucket, Close ≤ 5 mm / Mid 5–10 mm / Far 10–15 mm) and quantitative CV tracker accuracy are deferred to the camera-ready version.

The 0.3 mm figure describes connector clearance. The demonstration does not establish a measured positioning-accuracy distribution, repeatable deployment success rate, or electrical connection verification.

Team

MEng Autonomy and Robotics project at University of Illinois Urbana-Champaign (AE598 Advanced Robotic Planning, Prof. Timothy Bretl):

  • Het Patel — Residual SAC training, MuJoCo environment, curriculum design, ROS 2 deploy package
  • Sunny Deshpande — System architecture, OpenVLA integration, IK pipeline
  • Keisuke Ogawa — Stereo CV tracker, hand-eye calibration, hardware bring-up

Limitations & Future Work

  • OpenVLA is sim-only: Live VLA inference at 100 Hz on a single GPU is non-trivial; the on-hardware path uses the scripted base for safety and latency reasons
  • Sim-to-real F/T mismatch: MuJoCo wrist F/T noise is rougher than what the real UR5e reports; contact gate threshold was hand-tuned on hardware
  • Legacy ROS 1 trajectory replay: IK-pipeline trajectory currently replayed on real arm via ROS 1 publisher; residual node is ROS 2
  • Learned port-pose substitution: A learned port-pose head on synthetic + real RGB would generalize beyond pink-tape fiducials in cluttered scenes