HumanVLA: Shared Vision–Language Representations for Humanoid Control
Co-developed a course-project study of CLIP representations in HumanVLA for language-conditioned humanoid rearrangement. The report specifies a teacher–student design and evaluation on the HITR task setting in Isaac Gym; its CLIP performance values are hypothetical estimates.

Technical work
Three-author research design study · CLIP performance values are hypothetical estimates
- Co-developed a course-project study of CLIP image and text representations within HumanVLA for language-conditioned humanoid object rearrangement in the HITR task setting.
- Designed a student-policy variant with 128-dimensional adapters for visual and language embeddings, combining those representations with proprioception and previous actions for control.
- Documented a privileged teacher–student design using behavior cloning and DAgger, with a proposed comparison of frozen and trainable visual encoders under the same action decoder.
- Defined task-completion, placement-error, and execution-time metrics for evaluating the proposed encoder variants.
Overview
The study asks whether a shared visual–language representation can replace HumanVLA’s separate encoders while preserving a policy’s ability to connect task instructions, object observations, and humanoid state.
The completed report centers on HumanVLA, CLIP, HITR, and Isaac Gym. The authors contributed equally. Table I explicitly identifies the CLIP values as hypothetical estimates, so this portfolio does not present them as measured training or task-success results.
Research Question
Language-conditioned rearrangement requires a controller to relate object appearance and language to physical action. The proposed encoder change studies that interface while retaining the established teacher–student control structure.
- Can shared CLIP embeddings provide useful alignment between the image and instruction?
- How should those embeddings be adapted to the student policy’s input representation?
- What changes when the visual encoder is frozen versus trainable?
- How should task completion, placement error, and execution time be evaluated?
Study Contributions
- Co-developed the CLIP-based student representation and documented its relationship to the HumanVLA baseline.
- Specified 128-dimensional adapters for image and text embeddings.
- Combined semantic inputs with proprioception and previous actions in the student policy design.
- Compared frozen and trainable encoder configurations as a proposed ablation.
- Documented task and evaluation definitions for the HITR setting.
Technologies and Methods
- HumanVLA for the baseline humanoid control architecture.
- CLIP for shared image/text representations.
- Isaac Gym and HITR for the simulation/task setting.
- Privileged teacher, behavior cloning, DAgger, and an MLP action decoder.
- GR00T and OpenVLA appear as related work, not implemented control stacks in this study.
Technical Architecture
Student representation
- Encode visual observations and language instructions with CLIP.
- Project image/text features through 128-dimensional adapters.
- Combine adapted features with proprioception and previous actions.
- Retain the student action decoder to focus the study on the representation change.
Teacher–student learning design
- A privileged teacher supplies action supervision from simulation state.
- Behavior cloning initializes the student from teacher demonstrations.
- DAgger collects corrections on states visited by the student to address distribution shift.
- Frozen and trainable visual-encoder variants test sensitivity to representation updates.
Evaluation Scope
The report describes a 615-task HITR setting. This is the task setting, not a measured count of tasks solved by the project.
Table I reproduces paper baseline results and labels the CLIP values as hypothetical estimates. These values are not empirical results of the proposed CLIP variants.
- Evaluation dimensions: task completion, object placement error, and execution time.
- A measured comparison would require checkpoints, seeds, task splits, and evaluation logs.
- Physical-robot transfer and generalization to household scenes were not established.
Limitations & Further Work
- Measure frozen versus fine-tuned CLIP variants under a common protocol before drawing encoder-performance conclusions.
- Inspect whether the policy uses language or relies mainly on visual and proprioceptive cues.
- Test robustness to appearance changes, contact variation, and instruction ambiguity.
- SLAM-based navigation, BEHAVIOR-1K/OmniGibson scenes, waste sorting, and hardware transfer are possible extensions. They are distinct from the documented HumanVLA study.
Team
Het Patel • Vardhan Dongre • Sunny Deshpande
Equal-contribution course project. The report is retained as the primary description of the proposed architecture and its evaluation status.