← All projects
PerceptionDecember 2025

Open-World Semantic-Based Zero-Shot 6D Pose Estimation Using SAM3 And FoundationPose

Led integration of SAM-3 segmentation, FoundationPose tracking, and dynamic language-guided target switching in a four-person computer-vision project. The 44-evaluation YCB-Video subset records 88.31% mean ADD-S AUC at 13.08 seconds per processed frame.

FoundationPoseSAM-3Moondream2Zero-ShotVLMObjaverse-XLPyTorch6D Pose

Open-world 6D pose estimation
Open-world 6D pose estimation · Open animation

Technical work

Four-person computer vision project · Offline YCB-Video evaluation

  • Led integration of SAM-3 segmentation with FoundationPose, coordinating mask generation, pose estimation, and dynamic language-guided target switching across the team’s perception pipeline.
  • The integrated pipeline recorded 88.31% mean ADD-S AUC across a 44-evaluation YCB-Video subset covering 18 objects and 787 frames; mean end-to-end processing time was 13.08 seconds per frame.
  • Integrated tracking with the team’s mesh-acquisition workflow, supporting benchmark CAD, Objaverse-XL retrieval, and experiments with single-image reconstruction.
  • Analyzed rotational symmetry, proxy-mesh mismatch, and segmentation latency; documented 44.23° mean rotation error and retained per-object results to expose failures hidden by the aggregate score.

Overview

The team combined semantic scene analysis, SAM-3 masks, FoundationPose tracking, and mesh acquisition to explore language-guided object pose estimation. My contribution was pipeline integration, SAM-3/FoundationPose coordination, and dynamic target switching.

The evaluation is offline. It permits benchmark CAD models, while the mesh-acquisition workflow also explores Objaverse-XL retrieval and single-image reconstruction. The aggregate metric does not isolate the performance of generated or retrieved meshes.

Problem Statement

Object-pose pipelines need an object representation, a segmentation mask, and an initialized pose. The project explores how language can select the target and how modular mask, mesh, and tracking components can reduce manual setup.

Retrieval introduces shape mismatch and segmentation introduces substantial latency. Both matter when moving from an offline demonstration toward responsive manipulation.

Key Features

  • Open-Vocabulary Detection: Moondream2 VLM for semantic scene analysis in the team pipeline
  • Mesh Acquisition: On-the-fly 3D proxy generation via Objaverse-XL retrieval (10M+ assets) and TripoSR
  • Language-Driven Segmentation: SAM-3 integration for text-prompt-based target segmentation
  • Dynamic Target Switching: Language-guided target updates coordinated across mask, mesh, and pose-tracking components
  • Hierarchical Mesh Acquisition: Three-tier system (benchmark CAD → Objaverse retrieval → TripoSR generation)
  • Pose Tracking: Render-and-compare architecture with transformer-based pose refinement
  • Mesh Dictionary Caching: Asynchronous mesh fetching and caching to eliminate redundant queries
  • Multi-Object Support: Semantic object inventory and target selection; throughput is evaluated on the processed target frames

Technologies Used

FoundationPose (CVPR 2024) SAM-3 Moondream2 VLM Objaverse-XL TripoSR YOLOv8 PyTorch 2.0/2.7 CUDA 11.8/12.6 NVDiffRast Gemini API Python

Technical Architecture

Stage 1: Semantic Scene Analysis

  • Moondream2 VLM generates comprehensive object inventory from RGB stream
  • Produces discrete candidate list (e.g., "red bottle," "blue cup," "black keyboard")
  • Gemini API fallback for semantic query enhancement when detection fails
  • Enables prompt-based object specification even for BOP unseen objects

Stage 2: On-the-Fly 3D Mesh Generation

  • Primary: Load ground-truth CAD models when available (benchmark datasets)
  • Retrieval: Query Objaverse-XL database via language-guided similarity search
  • Generation: TripoSR generates candidate mesh from single observed image
  • Selection: Mesh manager scores candidates based on silhouette, depth, and IoU alignment

Stage 3: Language-Driven Segmentation

  • SAM-3 accepts natural language prompts to output pixel-level masks
  • Produces prompted masks; a controlled detector comparison is not reported
  • Temporal consistency for video tracking with frame-to-frame coherence

Stage 4: Unified 6D Pose Estimation & Tracking

  • Render-and-compare: FoundationPose aligns retrieved mesh with video observation
  • Pose scoring: Uniform sampling, composite scoring (IoU + Depth + Silhouette)
  • Iterative refinement: Transformer-based pose refinement of the pose hypothesis
  • Dynamic switching: Coordinated mask and mesh updates change the tracked target

Results & Performance

Unweighted means over the 44 object-scene rows in the linked detailed evaluation CSV. This subset is not a full YCB-Video benchmark run. The batch log lists 55 completed jobs, while the summary includes 44 evaluations; the difference is not reconciled in the supplied artifacts.

44 Evaluations • 12 Scenes • 18 Objects • 787 Frames

Metric Mean ± Std across evaluations
ADD AUC 71.61% ± 39.10%
ADD-S AUC 88.31% ± 28.56%
Rotation Error 44.23° ± 58.38°
Translation Error 2.44cm ± 6.13cm
Processing Time 13.08s ± 0.15s

Selected Object Results

Object ADD AUC Rotation Translation
Power Drill (4 scenes) 100% 2.0° ± 0.4° 0.24cm
Bleach Cleanser (3 scenes) 100% 2.4° ± 0.8° 0.34cm
Banana (2 scenes) 100% 5.3° ± 0.4° 0.40cm
Mustard Bottle (2 scenes) 90% 19.7° ± 25.6° 0.22cm
Pudding Box (1 scene) 100% 2.7° 0.24cm

Runtime Breakdown

  • Average Processing Time: 13.08 seconds per frame
  • SAM-3 Segmentation: ~12.5 seconds (95.6% of total time)
  • FoundationPose Tracking: ~0.58 seconds for the tracking component
  • First Frame Registration: 13–16 seconds (includes mask generation + initial pose alignment)

Comparison Scope

The project extends FoundationPose with language-guided target selection, segmentation, and mesh-acquisition interfaces. PoseCNN and DenseFusion provide related-work context in the report.

The 44-evaluation subset uses a different coverage and aggregation from published full-benchmark scores. It therefore does not establish a ranking against those methods or a state-of-the-art result. A fair comparison would run all methods on the same frames, object models, and initialization protocol.

Challenges & Solutions

  • Challenge: Direct 3D reconstruction produced low-quality meshes with artifacts
    Solution: Shifted to retrieval-based strategy leveraging the Objaverse-XL asset collection
  • Challenge: Symmetric objects exhibited rotational ambiguity
    Evaluation treatment: Reported ADD-S alongside raw rotation error to account for symmetric geometry. This changes evaluation, not the estimated orientation; rotation failures remain visible in the per-object results.
  • Challenge: SAM-3 mask generation dominated processing time (~12.5 sec/frame)
    Solution: Separate segmentation and tracking timings expose the bottleneck; the full pipeline averages 13.08 seconds per processed frame
  • Challenge: Retrieved meshes may differ in exact proportions from real instances
    Solution: Depth-based scale estimator adjusts mesh dimensions; composite scoring selects best candidate

Team

Het Patel • Sunny Deshpande • Ansh Bhansali • Keisuke Ogawa

CS543 Computer Vision — University of Illinois Urbana-Champaign — December 2025

My contribution: overall pipeline integration, SAM-3/FoundationPose coordination, and dynamic target switching. Semantic scene analysis and mesh-acquisition implementation were shared team work.