Homanga Bharadhwaj

I am an Assistant Professor of Computer Science at Johns Hopkins University. I am also a core faculty member in the Data Science and AI Institute and the Laboratory for Computational Sensing and Robotics .

At Hopkins, I lead the Brains, Bots, and Behavior Lab (website launching soon!).

I am engaged in the quest for understanding embodied intelligence by trying to simulate it. Although this quest has kept me fully occupied for the past several years, I also paint and write poems, and have a Bachelor of Arts degree in Fine Arts. Some of my paintings can be found here, and poems here.

Refresh page to see Homangas in different environments.
Education / Employment

News / Highlights

If you have any questions / want to collaborate, feel free to send me an email! I am always excited to learn more by talking with people. Please include "HELLOHOMANGA" in the subject of the email so that I don't miss it.

Research

I'm interested in developing embodied AI systems capable of helping us in the humdrum of everyday activities within messy rooms, offices, and kitchens, in a reliable, compliant, and scalable manner without requiring significant embodiment-specific data collection and task-specific heuristics. A major thrust of my research is on combining robot-specific data with predictive planning from diverse web videos such as YouTube clips of humans doing daily chores and human interaction data obtained through wearables, for developing robust robot learning algorithms deployable in the real-world and compliant with human preferences. I have eclectic research interests, and have also worked on robustness in machine learning, and improving sample efficiency, and representations in reinforcement learning.

In my research, I conduct experiments across robot embodiments for demonstrating generalization of policies to unseen tasks including those involving manipulation of completely unseen object types with novel motions. Here are some glimpses of results:

Research highlights

Recent frameworks that learn dexterous behavior from human video and forecast how interactions unfold in 3D.

ICRA 2026

AINA

Multi-fingered robot manipulation learned from in-the-wild human demonstrations.

IROS 2026

SPIDER

Physics-informed retargeting of human motion to dexterous robot embodiments.

arXiv 2026

MotionForesight

Future object-centered 3D scene-flow forecasting from monocular human videos.

ECCV 2024

Track2Act

Point-track prediction from internet video for diverse zero-shot robot manipulation.

Most significant bits
MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction
Homanga Bharadhwaj*, Yash Jangir*
arXiv 2026
paper   website
We demonstrate a very simple approach for repurposing video models to forecast future 3D scene flow in daily scenarios from short monocular human-interaction videos. Trained on just 40k videos, it generalizes to diverse real-world scenarios.
Functional Force-Aware Retargeting from Virtual Human Demos to Soft Robot Policies
Uksang Yoo, Mengjia Zhu, Evan Pezent, Jom Preechayasomboon, Jean Oh, Jeffrey Ichnowski, Amir Memar, Ben Abbatematteo, Homanga Bharadhwaj, Ashish Deshpande, Harsha Prahlad
RSS 2026  
paper website
SoftAct enables functional retargeting from virtual human demos to soft robot hands by reasoning about contact geometry and force distribution.
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
Hongyi Chen, Tony Dong, Tiancheng Wu, Liquan Wang, Yash Jangir, Yaru Niu, Yufei Ye, Homanga Bharadhwaj, Zackory Erickson*, Jeffrey Ichnowski*
arXiv 2026  
paper website
We propose VideoManip: a device-free framework that learns dexterous manipulation directly from RGB human videos by reconstructing 3D hand-object interaction trajectories for policy learning.
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
Rustin Soraki, Homanga Bharadhwaj*, Ali Farhadi*, Roozbeh Mottaghi*
ECCV 2026  
paper website

We tackle the problem of predicting future 3D object trajectories given a past context, purely from human videos of everyday interactions.

Walk through Paintings: Ego-centric World Models from Internet Priors
Anurag Bagchi, Zhipeng Bao, Homanga Bharadhwaj, Yu-Xiong Wang, Pavel Tokmakov, Martial Hebert
ECCV 2026  
paper website

We develop a scalable recipe for converting pre-trained video models to ego-centric world models, capable of diverse OOD manipulation and navigation.

Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction Videos
Mingfei Chen, Yifan Wang, Zhengqin Li, Homanga Bharadhwaj, Yujin Chen, Chuan Qin, Ziyi Kou, Yuan Tian, Eric Whitmire, Rajinder Sodhi, Hrvoje Benko, Eli Shlizerman, Yue Liu
ECCV 2026  
paper website

EgoMan: Motion cues from human videos + Reasoning from VLMs enables future 3D hand trajectory prediction in-the-wild for novel tasks in novel scenes.

Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations
Irmak Guzey, Haozhi Qi, Julen Urain, Changhao Wang, Jessica Yin, Krishna Bodduluri, Mike Lambeta, Lerrel Pinto, Akshara Rai, Jitendra Malik, Tingfan Wu, Akash Sharma, Homanga Bharadhwaj
ICRA 2026  
paper website  

AINA is a framework for building multi-fingered robot manipulation policies directly by watching videos of humans with Aria glasses on, without any robot interaction/tele-operation/simulation data.

SPIDER: Scalable Physics-Informed DExterous Retargeting
Chaoyi Pan, Changhao Wang, Haozhi Qi, Zixi Liu, Homanga Bharadhwaj, Akash Sharma, Tingfan Wu, Guanya Shi, Jitendra Malik, Francois Hogan
IROS 2026  
paper website  

SPIDER is a physics-based retargeting framework to transform and augment kinematic-only human demonstrations to dynamically feasible robot trajectories at scale. By aligning human motion and robot feasibility at scale, SPIDER offers a general, embodiment-agnostic foundation for humanoid and dexterous hand control.

DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy
Sungjae Park, Homanga Bharadhwaj, Shubham Tulsiani
ICRA 2026  
paper website  

DemoDiffusion is a simple and scalable method for enabling robots to perform manipulation tasks by imitating a single human demonstration, without requiring any paired human-robot data or reinforcement learning. The key idea is to refine a re-targeted human trajectory using a pre-trained generalist diffusion policy.

Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
Hongyi Chen, Yunchao Yao*, Yufei Ye, Homanga Bharadhwaj, Jiashun Wang, Shubham Tulsiani, Zackory Erickson, Jeffrey Ichnowski
IROS 2026  
paper website  

Humans grasp objects with a purpose! Web2Grasp enables such functional grasping for dexterous robot hands via hand-object reconstruction from web images - without requiring any robot teleop data collection for imitation learning.

HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
Chen Bao, Jiarui Xu, Xiaolong Wang*, Abhinav Gupta*, Homanga Bharadhwaj*
TMLR, 2025  
paper website  

We develop an in-context action prediction assistant for daily activities. HandsOnVLM enables predicting future interaction trajectories of human hands in a scene given high-level colloquial task specifications in the form of natural language.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation
Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta, Shubham Tulsiani, Carl Doersch, Ted Xiao, Dhruv Shah, Fei Xia, Dorsa Sadigh, Sean Kirmani
CoRL 2025  
paper website video  

Casting language-conditioned manipulation as human video generation followed by closed-loop policy execution conditioned on the generated video enables solving diverse real-world tasks involving object/motion types unseen in the robot dataset.

Track2Act: Predicting Point Tracks from Internet Videos Enables Diverse Zero-shot Manipulation
Homanga Bharadhwaj, Roozbeh Mottaghi*, Abhinav Gupta*, Shubham Tulsiani*
ECCV 2024  
paper website  

We can train a model for embodiment-agnostic point track prediction from web videos combined with embodiment-specific residual policy learning for diverse real-world manipulation in everyday office and kitchen scenes. The resulting goal-conditioned policy can be zero-shot deployed in unseen scenarios.

RoboAgent: Towards Sample Efficient Robot Manipulation with Semantic Augmentations and Action Chunking
Homanga Bharadhwaj*, Jay Vakil*, Mohit Sharma*, Abhinav Gupta, Shubham Tulsiani, Vikash Kumar
ICRA 2024  
Robot Learning Workshop, NeurIPS 2023 (Outstanding Presentation Award)  
paper website data  

We can develop a single robot manipulation agent capable of over 38 tasks across 100s of scenes, through semantic augmentations for multiplying data, and action chunking transformers for fitting the multi-modal data distribution.

Towards Generalizable Zero-Shot Manipulation via Translating Human Interaction Plans
Homanga Bharadhwaj, Abhinav Gupta*, Vikash Kumar*, Shubham Tulsiani*
ICRA 2024 (Best Paper in Robot Manipulation Finalist)  
paper website video  

Learning interaction plans from diverse passive human videos on the web, followed by translation to robotic embodiments can help develop a single goal-conditioned policy that scales to over 100 diverse tasks in unseen scenarios, including real kitchens and offices.

CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning
Mandi Zhao, Homanga Bharadhwaj, Vincent Moens, Shuran Song, Aravind Rajeswaran, Vikash Kumar
Pretraining for Robotics Workshop, CoRL 2022 (Spotlight Talk)  
paper website

Through effective augmentations enabled by recent advances in generative modeling, we can develop a framework for learning robust manipulation policies capable of solving multiple tasks in diverse real-world scenes.

For the full paper list, check my Google Scholar.


I love his website design.

Visitor Hit Counter