← All projects
ManipulationOctober 2025

HumanVLA: Shared Vision–Language Representations for Humanoid Control

Co-developed a course-project study of CLIP representations in HumanVLA for language-conditioned humanoid rearrangement. The report specifies a teacher–student design and evaluation on the HITR task setting in Isaac Gym; its CLIP performance values are hypothetical estimates.

HumanVLACLIPIsaac GymHITRBehavior CloningDAggerPyTorch

HumanVLA / Humanoid representations
HumanVLA / Humanoid representations · Open animation

Technical work

Three-author research design study · CLIP performance values are hypothetical estimates

  • Co-developed a course-project study of CLIP image and text representations within HumanVLA for language-conditioned humanoid object rearrangement in the HITR task setting.
  • Designed a student-policy variant with 128-dimensional adapters for visual and language embeddings, combining those representations with proprioception and previous actions for control.
  • Documented a privileged teacher–student design using behavior cloning and DAgger, with a proposed comparison of frozen and trainable visual encoders under the same action decoder.
  • Defined task-completion, placement-error, and execution-time metrics for evaluating the proposed encoder variants.

Overview

The study asks whether a shared visual–language representation can replace HumanVLA’s separate encoders while preserving a policy’s ability to connect task instructions, object observations, and humanoid state.

The completed report centers on HumanVLA, CLIP, HITR, and Isaac Gym. The authors contributed equally. Table I explicitly identifies the CLIP values as hypothetical estimates, so this portfolio does not present them as measured training or task-success results.

Research Question

Language-conditioned rearrangement requires a controller to relate object appearance and language to physical action. The proposed encoder change studies that interface while retaining the established teacher–student control structure.

  • Can shared CLIP embeddings provide useful alignment between the image and instruction?
  • How should those embeddings be adapted to the student policy’s input representation?
  • What changes when the visual encoder is frozen versus trainable?
  • How should task completion, placement error, and execution time be evaluated?

Study Contributions

  • Co-developed the CLIP-based student representation and documented its relationship to the HumanVLA baseline.
  • Specified 128-dimensional adapters for image and text embeddings.
  • Combined semantic inputs with proprioception and previous actions in the student policy design.
  • Compared frozen and trainable encoder configurations as a proposed ablation.
  • Documented task and evaluation definitions for the HITR setting.

Technologies and Methods

  • HumanVLA for the baseline humanoid control architecture.
  • CLIP for shared image/text representations.
  • Isaac Gym and HITR for the simulation/task setting.
  • Privileged teacher, behavior cloning, DAgger, and an MLP action decoder.
  • GR00T and OpenVLA appear as related work, not implemented control stacks in this study.

Technical Architecture

Student representation

  • Encode visual observations and language instructions with CLIP.
  • Project image/text features through 128-dimensional adapters.
  • Combine adapted features with proprioception and previous actions.
  • Retain the student action decoder to focus the study on the representation change.

Teacher–student learning design

  • A privileged teacher supplies action supervision from simulation state.
  • Behavior cloning initializes the student from teacher demonstrations.
  • DAgger collects corrections on states visited by the student to address distribution shift.
  • Frozen and trainable visual-encoder variants test sensitivity to representation updates.

Evaluation Scope

The report describes a 615-task HITR setting. This is the task setting, not a measured count of tasks solved by the project.

Table I reproduces paper baseline results and labels the CLIP values as hypothetical estimates. These values are not empirical results of the proposed CLIP variants.

  • Evaluation dimensions: task completion, object placement error, and execution time.
  • A measured comparison would require checkpoints, seeds, task splits, and evaluation logs.
  • Physical-robot transfer and generalization to household scenes were not established.

Limitations & Further Work

  • Measure frozen versus fine-tuned CLIP variants under a common protocol before drawing encoder-performance conclusions.
  • Inspect whether the policy uses language or relies mainly on visual and proprioceptive cues.
  • Test robustness to appearance changes, contact variation, and instruction ambiguity.
  • SLAM-based navigation, BEHAVIOR-1K/OmniGibson scenes, waste sorting, and hardware transfer are possible extensions. They are distinct from the documented HumanVLA study.

Team

Het Patel • Vardhan Dongre • Sunny Deshpande

Equal-contribution course project. The report is retained as the primary description of the proposed architecture and its evaluation status.