At Hopkins, I lead the Brains, Bots, and Behavior Lab (website launching soon!).
I am engaged in the quest for understanding embodied intelligence by trying to simulate it. Although this quest has kept me fully occupied for the past several years, I also paint and write poems, and have a Bachelor of Arts degree in Fine Arts. Some of my paintings can be found here, and poems here.
[2024] Open-X: Best Conference Paper Award at ICRA 2024
[2024] HOPMan: Best Paper in Robot Manipulation Finalist at ICRA 2024
[2023] RoboAgent: Outstanding Presentation Award at the Robot Learning Workshop, NeurIPS 2023
[2023] Our sample-efficient universal manipulation research was covered by TechCrunch, ACM, IEEE
[2022] Selected as a AI Mentorship (AIM) Fellow by Meta
[2021] Our research on safe exploration for robotics was convered by VentureBeat
If you have any questions / want to collaborate, feel free to send me an email! I am always excited to learn more by talking with people. Please include "HELLOHOMANGA" in the subject of the email so that I don't miss it.
Research
I'm interested in developing embodied AI systems capable of helping us in the humdrum of everyday activities within messy rooms, offices, and kitchens, in a reliable, compliant, and scalable manner without requiring significant embodiment-specific data collection and task-specific heuristics. A major thrust of my research is on combining robot-specific data with predictive planning from diverse web videos such as YouTube clips of humans doing daily chores and human interaction data obtained through wearables, for developing robust robot learning algorithms deployable in the real-world and compliant with human preferences. I have eclectic research interests, and have also worked on robustness in machine learning, and improving sample efficiency, and representations in reinforcement learning.
In my research, I conduct experiments across robot embodiments for demonstrating generalization of policies to unseen tasks including those involving manipulation of completely unseen object types with novel motions. Here are some glimpses of results:
Research highlights
Recent frameworks that learn dexterous behavior from human video and forecast how interactions unfold in 3D.
AINA is a framework for building multi-fingered robot manipulation policies directly by watching videos of humans with Aria glasses on, without any robot interaction/tele-operation/simulation data.
SPIDER is a physics-based retargeting framework to transform and augment kinematic-only human demonstrations to dynamically feasible robot trajectories at scale. By aligning human motion and robot feasibility at scale, SPIDER offers a general, embodiment-agnostic foundation for humanoid and dexterous hand control.
DemoDiffusion is a simple and scalable method for enabling robots to perform manipulation tasks by imitating a single human demonstration, without requiring any paired human-robot data or reinforcement learning. The key idea is to refine a re-targeted human trajectory using a pre-trained generalist diffusion policy.
Humans grasp objects with a purpose! Web2Grasp enables such functional grasping for dexterous robot hands via hand-object reconstruction from web images - without requiring any robot teleop data collection for imitation learning.
We develop an in-context action prediction assistant for daily activities. HandsOnVLM enables predicting future interaction trajectories of human hands in a scene given high-level colloquial task specifications in the form of natural language.
Casting language-conditioned manipulation as human video generation followed by closed-loop policy execution conditioned on the generated video enables solving diverse real-world tasks involving object/motion types unseen in the robot dataset.
We can train a model for embodiment-agnostic point track prediction from web videos combined with embodiment-specific residual policy learning for diverse real-world manipulation in everyday office and kitchen scenes. The resulting goal-conditioned policy can be zero-shot deployed in unseen scenarios.
We can develop a single robot manipulation agent capable of over 38 tasks across 100s of scenes, through semantic augmentations for multiplying data, and action chunking transformers for fitting the multi-modal data distribution.
Learning interaction plans from diverse passive human videos on the web, followed by translation to robotic embodiments can help develop a single goal-conditioned policy that scales to over 100 diverse tasks in unseen scenarios, including real kitchens and offices.
Through effective augmentations enabled by recent advances in generative modeling, we can develop a framework for learning robust manipulation policies capable of solving multiple tasks in diverse real-world scenes.