Research
I approach this question from two connected directions. In visual navigation and tracking, I study how a
policy can adapt its behavior to the spatial context instead of following a fixed notion of safety, and
how the dynamics of observer, target and environment can be modeled with explicit causal structure rather
than as one entangled system. In human motion, I study how continuous movement aligns with language, and
what a representation can learn from the temporal structure of movement without labels. The two
directions meet when a robot learns from a person: a human demonstration can be read as a statement of
how the world should change, which the robot then carries out with its own body. In the long term, I want
to make interaction an organizing principle for visual understanding, and to build models around that
idea.
DIRECTION I
Visual navigation and tracking
Junyi Hu, Shuaihang Yuan, Geeta Chandra Raju Bethala, Anthony Tzes, Yi Fang
IROS 2026
First Author
Agent β Environment
How much room a robot should leave depends on where it is, and a margin fixed in advance is too
cautious in one place and too tight in another. We let the robot read it from what it sees: a diffusion
policy proposes paths from RGB-D, and a critic trained with map geometry picks the one whose clearance
fits the scene, while the deployed system needs no map. On HM3D it reaches 0.783 success against
0.710 for NavDP, and it runs on a Unitree G1.
Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang
Under Review
First Author
Causal Structure
A robot that follows a person sees the world move whenever it moves, and a world model trained on that
footage takes a shortcut: it comes to believe its own steering is what moves the person. CST-WM gives
the model the causal structure of the scene, with separate branches for target, robot and observation
and no direct path from the action to the target, so an action can change what the robot sees of the
target only by moving it. On a Unitree Go2 it follows a person through occlusion, distractors and fast
motion in 20 of 30 trials, against 14 for TrackVLA.
DIRECTION II
Human motion
Junyi Hu, Zhewen He, Haomian Huang, Aoxiang Yang, Yi Fang
ECCV 2026
First Author
Motion β Language
VTaMo achieves unprecedented alignment between video and text. Rather than leaving the
correspondence between a continuous stream and a sentence to attention, it matches the two explicitly
with optimal transport, word by word; because each word is assigned a stretch of frames, the alignment
also cuts the stream into the pieces that carry each word, with no frame-level labels. Experiments are
on four sign-language translation benchmarks.
Junyi Hu, Zhewen He, Haomian Huang, Zhenhua Li, Zhifei Li, Yi Fang
Under Review
First Author
Self-Supervised Learning
Image self-supervision learns by comparing a whole image with small crops taken from it. SignDino
applies the same idea along time: it follows each hand and the face separately through a video
and compares a short window of one part with that partβs whole sequence, so a representation of motion
is learned from how the body moves, with no labels at all. Experiments are on sign-language video,
where the hands and the face carry the signal.
Junyi Hu, Zhewen He, Haomian Huang, Zhenhua Li, Yi Fang
Under Review
First Author
Real-Time Streaming
The first real-time translation pipeline built for the real world, and both fast and accurate.
Online systems so far simply run a recognizer over sliding windows; this one is designed for noisy live
input: it trusts each joint only as far as the pose detector does, recognizes from a short window of
motion, and commits each word once, as soon as it is sure, without ever reading future frames.
Experiments are on five sign-language translation benchmarks.
Zhewen He, Junyi Hu, Haomian Huang, Zhenhua Li, Yu-Shen Liu, Yi Fang
ECCV 2026
Co-Author
~1M Clips Β· 2,058 Hours
Recorded motion video comes from a few cameras, rooms and people, so a model trained on it breaks when
any of them changes, and a benchmark score does not say which one did it. SignNet-1M re-renders
existing video with new viewpoints, backgrounds and performers while keeping the motion, and tags every
clip with what was changed and by how much: about one million clips (2,058 hours), built from seven
sign-language corpora.
WHERE THEY MEET
From a human demonstration to a robotβs execution
Junyi Hu, Zhewen He, Yi Fang
Under Review
First Author
Human β Robot
A demonstration shows two things at once: what happened to the objects, and how one particular arm made
it happen. PEWAM keeps only the first. It compiles a demonstration into an effect program, which
records where each object travels, where it is held and how the scene ends, with the demonstratorβs
body removed, and a world-action model works out, closed loop, how its own body can produce that
effect. A person can edit the program by hand, the same program runs on four different arms without any
retraining, and on a real Franka arm programs taken from single human videos complete 36 of 40
trials, against 22 for the same model given a goal image.
OTHER WORK
Surface representation
Zhewen He, Junyi Hu, Yi Fang
Under Review
Co-Author
Surface Representation
Heat-kernel textures describe appearance on a surface through how heat spreads over it, but evaluate
every kernel through a global spectral basis of the whole mesh. HeatAtlas evaluates each kernel where
it lives, on a local chart unfolded from the mesh, and makes those charts the representation itself,
which removes all global preprocessing. On a set of 32 structurally hard objects it reaches 32.29
dB, against 26.82 dB for the strongest baseline.