Contact-Gated Residual RL for Precision USB Insertion with VLA-Guided Trajectory Planning
Developed the MuJoCo environment, SAC residual training, curriculum, and ROS 2 deployment package for contact-gated USB insertion. Training spans approximately 21.6 million steps; the UR5e demonstration uses a scripted approach, while OpenVLA was evaluated in simulation.

Technical work
MuJoCo training · OpenVLA in simulation · UR5e demo with scripted approach
- Trained a six-joint Soft Actor-Critic policy for USB insertion in MuJoCo over approximately 21.6 million environment steps, tightening the cavity success criterion from 1.80× to 1.0625× through a staged curriculum.
- Implemented a latched contact handoff to prevent repeated controller switching when the plug briefly loses contact; revised the depth reward to address policies that hovered at the port rim.
- Built a 48-dimensional observation with joint state, port-relative error, force/torque history, and previous actions; randomized port pose, friction, sensor noise, and control latency during training.
- Developed the ROS 2 deployment package for UR5e and Robotiq 2F-85, supporting a 100 Hz command loop and a real-arm insertion demonstration with a scripted approach.
Overview
Vision-Language-Action (VLA) models like OpenVLA generalize remarkably well across tabletop manipulation tasks, but their joint-target outputs sit several millimeters off the goal axis once a peg touches its mating surface — roughly an order of magnitude larger than the 0.3 mm clearance budget of a USB-A connector.
This project bridges that gap with a contact-gated residual SAC policy: the VLA (or scripted base) owns the free-space approach, and the residual takes over completely at first contact, handling the sub-millimeter compliance needed to slip the plug past the chamfer and seat it fully. The residual is small, trained entirely in simulation, and deploys as a ROS 2 package on a UR5e + Robotiq 2F-85 platform.
A separate stereo CV port-pose tracker (pink-tape fiducials, epipolar pairing, one-shot ChArUco camera-to-base calibration) feeds a pinocchio QP-IK pick-and-transport pipeline, decoupled from but fully compatible with the residual deploy path.
The Core Idea: Latching Contact Gate
The key design insight is a one-way latching handoff. Earlier approaches tracked contact force in real time and toggled the residual on/off, producing a bang-bang contact-bounce loop: residual pushes → plug briefly loses contact → gate snaps shut → base policy drifts arm upward → contact resumes → repeat.
Switching to a latch that fires once and never releases ends the loop. Once &lVert;F&rVert; > 2 N, the base policy is dropped for the rest of the episode and the residual drives alone from the static pre-insertion pose:
- g = 0 (approach): q_cmd = q_base — VLA / scripted trajectory drives free-space approach
- g = 1 (insertion): q_cmd = q_pre_insert + q_res — residual-only, latched for the rest of the episode
This separation makes the residual's training distribution clean: it never has to learn to share control mid-episode, and the upstream VLA is free to use any joint-target interface without disturbing the downstream residual.
Control Pipeline / 100 Hz Command Loop
The command loop is distinct from model inference. OpenVLA is a simulation-only base-policy option; the hardware demonstration uses the scripted base.
- 1. Perception: 48-D observation packed from joint state (q, q̇), bias-corrected wrist F/T (F, T), TCP pose error vs. port (Δxyz, Δrpy), previous raw action, previous residual, and a 2-step F/T history (+12 dims)
- 2. Base Policy: OpenVLA-7B (sim only) or scripted 3-waypoint joint-space trajectory (q_home → q_pre_insert → q_descent_end) at 100 Hz
- 3. Jacobian DLS-IK: Maps VLA end-effector delta to joint space via damped least squares (λ = 0.1); no-op for the scripted base which already emits joint targets
- 4. Residual SAC: MLP [256, 256, 256] → 6-D action in [−1,1], scaled by per-joint clip [0.025, 0.025, 0.030, 0.040, 0.040, 0.050] rad, EMA-smoothed (α = 0.8)
- 5. Latching Handoff: Contact-force / range trigger flips the one-way gate; fused command pushed through 1-step latency FIFO to MuJoCo or ROS 2 /forward_position_controller
Training: Two-Stage Curriculum & Domain Randomization
Stage 1 — Cavity Tightening (S1 → S5a, ~20.7M steps): Starts with --disable_base_policy (residual learns the full descent) and a generous 1.80× cavity inflation. Each phase tightens the success criterion toward a near-flush 1.0625×, progressively adding scripted-base noise (σ from 0 → 0.01 rad/joint). This forces the residual to commit deeper rather than hovering at the rim.
Stage 2 — DR Augmentation (DRA2–DRA3, ~1M steps): Holds cavity tight and augments with F/T sensor noise and increased base-policy noise (σ = 0.015 rad/joint), targeting the noisy upstream stream a real VLA would produce.
Domain randomization axes active throughout: port pose (±3 mm xy, ±3° yaw), joint friction (×[0.70, 1.30]), position-gain (×[0.95, 1.05]), F/T bias (±0.5 N / ±0.05 Nm), plug spawn pose (xy ±15 mm, z ±7.5 mm), joint encoder noise, F/T sensor noise, 1-step action latency.
The total is approximate: one roughly 500,000-step phase is estimated because its logs were not retained. Training steps count simulation interactions, not successful insertion trials.
Key Contributions
- Latching Activation Gate: One-way handoff that prevents switching back to the approach controller during an episode; base policy fully owns approach, residual fully owns insertion — no shared mid-episode control
- Per-Joint Action Clip: Bounded per-joint residual offsets, with tighter limits on proximal joints (0.025 rad) and larger distal offsets (0.050 rad), narrowing the SAC action range
- F/T History Augmentation: Added two prior force/torque samples (+12 dimensions); qualitative rollouts showed fewer under-corrections with noisy approach commands
- Base-Policy Noise Curriculum: Scheduled joint-target Gaussian noise (σ from 0 to 0.015 rad/joint) during training to expose the residual to imperfect approach commands
- Stereo CV Tracker: Pink-tape fiducials, epipolar pairing, one-shot ChArUco camera-to-base calibration; shared between QP-IK pick-and-transport and residual deploy path
Reward Design
11 default-on reward terms cover pose proximity, orientation, depth, force shaping, action smoothness, residual magnitude, time cost, and terminal success bonus:
- Pose Proximity: Exponential bonus exp(−|e|/τ) with τ = 2 mm — steep gradient near the port, flat tail outside
- Depth Bonus: R_depth × (d/d*)^β with β = 1.5 — superlinear exponent critical to avoid the "hover at the rim" local minimum that β = 1 produces
- Force Shaping: Linear penalty above 10 N deadband + tent-shaped bonus around 3 N target contact force + per-step bonus for port-wall contact (prevents lateral skating)
- Action Smoothness: Penalties on 1-step delta (spikes) and windowed std-dev over 5 steps (sustained chatter)
- Terminal Bonus: +5×10⁴ on successful seating (plug tip at depth ≥ 10.5 mm inside inflated cavity)
Technologies Used
Results & Qualitative Findings
- Residual closes the last-millimeter gap: The report documents a representative real-arm seating demonstration with a scripted approach and learned insertion controller; the base brings the plug into the contact-onset zone, and the residual handles the sub-millimeter compliance needed to slip past the chamfer
- F/T history matters: Without the 2-step F/T history, the policy under-corrects when the upstream prior is noisy
- Base-noise curriculum matters: Without it, the policy overfits to the clean scripted base and degrades immediately when noise is injected at inference
- OpenVLA simulation integration: The residual SAC trained against the scripted base transferred directly when OpenVLA replaced the scripted base in simulation, because both share the same handoff contract: base delivers plug to contact-onset zone, residual takes over
- CV tracker shared cleanly: Same stereo triangulated port pose feeds both the QP-IK pick-and-transport pipeline and the residual's latched handoff reference
Full bucketed port-distance evaluation (100 episodes per bucket, Close ≤ 5 mm / Mid 5–10 mm / Far 10–15 mm) and quantitative CV tracker accuracy are deferred to the camera-ready version.
The 0.3 mm figure describes connector clearance. The demonstration does not establish a measured positioning-accuracy distribution, repeatable deployment success rate, or electrical connection verification.
Team
MEng Autonomy and Robotics project at University of Illinois Urbana-Champaign (AE598 Advanced Robotic Planning, Prof. Timothy Bretl):
- Het Patel — Residual SAC training, MuJoCo environment, curriculum design, ROS 2 deploy package
- Sunny Deshpande — System architecture, OpenVLA integration, IK pipeline
- Keisuke Ogawa — Stereo CV tracker, hand-eye calibration, hardware bring-up
Limitations & Future Work
- OpenVLA is sim-only: Live VLA inference at 100 Hz on a single GPU is non-trivial; the on-hardware path uses the scripted base for safety and latency reasons
- Sim-to-real F/T mismatch: MuJoCo wrist F/T noise is rougher than what the real UR5e reports; contact gate threshold was hand-tuned on hardware
- Legacy ROS 1 trajectory replay: IK-pipeline trajectory currently replayed on real arm via ROS 1 publisher; residual node is ROS 2
- Learned port-pose substitution: A learned port-pose head on synthetic + real RGB would generalize beyond pink-tape fiducials in cluttered scenes