Junyi Hu

Undergraduate, School of Mathematical Sciences, Fudan University

Research Assistant, New York University

Applying for PhD study, Fall 2027

I study how visual systems can understand interaction in the physical world: how people, agents, objects and environments influence one another over time.

Junyi Hu

News

Research

I approach this question from two connected directions. In visual navigation and tracking, I study how a policy can adapt its behavior to the spatial context instead of following a fixed notion of safety, and how the dynamics of observer, target and environment can be modeled with explicit causal structure rather than as one entangled system. In human motion, I study how continuous movement aligns with language, and what a representation can learn from the temporal structure of movement without labels. The two directions meet when a robot learns from a person: a human demonstration can be read as a statement of how the world should change, which the robot then carries out with its own body. In the long term, I want to make interaction an organizing principle for visual understanding, and to build models around that idea.

DIRECTION I

Visual navigation and tracking

A diffusion policy proposes candidate trajectories from RGB-D; a learned safety critic scores and selects one.

Learning Adaptive Safety Margins for Visual Navigation

Junyi Hu, Shuaihang Yuan, Geeta Chandra Raju Bethala, Anthony Tzes, Yi Fang

IROS 2026 First Author Agent ↔ Environment

How much room a robot should leave depends on where it is, and a margin fixed in advance is too cautious in one place and too tight in another. We let the robot read it from what it sees: a diffusion policy proposes paths from RGB-D, and a critic trained with map geometry picks the one whose clearance fits the scene, while the deployed system needs no map. On HM3D it reaches 0.783 success against 0.710 for NavDP, and it runs on a Unitree G1.

CST-WM separates target evidence, robot and observation, and removes the direct edge from action to target evidence.

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang

Under Review First Author Causal Structure

A robot that follows a person sees the world move whenever it moves, and a world model trained on that footage takes a shortcut: it comes to believe its own steering is what moves the person. CST-WM gives the model the causal structure of the scene, with separate branches for target, robot and observation and no direct path from the action to the target, so an action can change what the robot sees of the target only by moving it. On a Unitree Go2 it follows a person through occlusion, distractors and fast motion in 20 of 30 trials, against 14 for TrackVLA.

DIRECTION II

Human motion

The VTaMo pipeline: visual encoder, optimal-transport local alignment, orthogonal global alignment, and the language decoder.

VTaMo: Video–Text Alignment Model for Sign Language Translation

Junyi Hu, Zhewen He, Haomian Huang, Aoxiang Yang, Yi Fang

ECCV 2026 First Author Motion ↔ Language

VTaMo achieves unprecedented alignment between video and text. Rather than leaving the correspondence between a continuous stream and a sentence to attention, it matches the two explicitly with optimal transport, word by word; because each word is assigned a stretch of frames, the alignment also cuts the stream into the pieces that carry each word, with no frame-level labels. Experiments are on four sign-language translation benchmarks.

Left: image self-supervision samples global and local crops inside one image. Right: SignDino samples global and local views along the time axis of one tracked body part.

SignDino: Self-Supervised Sign Language Representation Learning via Temporal-Axis Self-Distillation

Junyi Hu, Zhewen He, Haomian Huang, Zhenhua Li, Zhifei Li, Yi Fang

Under Review First Author Self-Supervised Learning

Image self-supervision learns by comparing a whole image with small crops taken from it. SignDino applies the same idea along time: it follows each hand and the face separately through a video and compares a short window of one part with that part’s whole sequence, so a representation of motion is learned from how the body moves, with no labels at all. Experiments are on sign-language video, where the hands and the face carry the signal.

Three-stage pipeline: confidence-aware masked pose encoding, discriminative VQ tokenization, and a strictly causal streaming pipeline of retrieval, past-only voting and translation.

Real-Time Sign Language Translation by Prototype-Assisted Causal Streaming

Junyi Hu, Zhewen He, Haomian Huang, Zhenhua Li, Yi Fang

Under Review First Author Real-Time Streaming

The first real-time translation pipeline built for the real world, and both fast and accurate. Online systems so far simply run a recognizer over sliding windows; this one is designed for noisy live input: it trusts each joint only as far as the pose detector does, recognizes from a short window of motion, and commits each word once, as soon as it is sure, without ever reading future frames. Experiments are on five sign-language translation benchmarks.

Overview of the SignNet-1M pipeline: novel-view rendering, scene editing and performer substitution.

SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks

Zhewen He, Junyi Hu, Haomian Huang, Zhenhua Li, Yu-Shen Liu, Yi Fang

ECCV 2026 Co-Author ~1M Clips Β· 2,058 Hours

Recorded motion video comes from a few cameras, rooms and people, so a model trained on it breaks when any of them changes, and a benchmark score does not say which one did it. SignNet-1M re-renders existing video with new viewpoints, backgrounds and performers while keeping the motion, and tags every clip with what was changed and by how much: about one million clips (2,058 hours), built from seven sign-language corpora.

WHERE THEY MEET

From a human demonstration to a robot’s execution

Copying the demonstrator fails on a held-out task, after a push, on another arm and under a new grasp; programming the effect succeeds in all four.

Keep the Effect, Drop the Actor: Programmable Effect-to-Execution World-Action Models

Junyi Hu, Zhewen He, Yi Fang

Under Review First Author Human β†’ Robot

A demonstration shows two things at once: what happened to the objects, and how one particular arm made it happen. PEWAM keeps only the first. It compiles a demonstration into an effect program, which records where each object travels, where it is held and how the scene ends, with the demonstrator’s body removed, and a world-action model works out, closed loop, how its own body can produce that effect. A person can edit the program by hand, the same program runs on four different arms without any retraining, and on a real Franka arm programs taken from single human videos complete 36 of 40 trials, against 22 for the same model given a goal image.

OTHER WORK

Surface representation

HeatAtlas overview: local charts unfolded from the mesh, short-time heat kernel weights, chart content, and composition.

HeatAtlas: Surface Textures as Learnable Atlases of Local Heat Kernels

Zhewen He, Junyi Hu, Yi Fang

Under Review Co-Author Surface Representation

Heat-kernel textures describe appearance on a surface through how heat spreads over it, but evaluate every kernel through a global spectral basis of the whole mesh. HeatAtlas evaluates each kernel where it lives, on a local chart unfolded from the mesh, and makes those charts the representation itself, which removes all global preprocessing. On a set of 32 structurally hard objects it reaches 32.29 dB, against 26.82 dB for the strongest baseline.

Future Work

A human demonstration is usually consumed as a trajectory to imitate. PEWAM reads it instead as a statement of how the world should change, but I chose the terms of that statement: rigid boxes, two contact points, an end state. Effects in the physical world say more than that. Objects bend, open, pour and break, tasks have an order and a phase, and one change sets up the next. I want effect programs that can express these things, and world-action models that find their own way to carry them out, even on bodies they were never trained on.

The same holds for the structure my models rely on. The causal factorization in CST-WM, the body-part streams in SignDino and the effect program in PEWAM were all mine to design, which does not scale past problems I already understand. Can such structure be discovered from data rather than designed by hand, and what evidence would show that a discovered structure is real, rather than merely convenient for the task?

I also want to explore new foundational representations. Heat diffusion is one example I have started to work with: it describes how a quantity spreads from each point to its neighbours over time using nothing but local geometry, which makes it a natural way to write down how local influence builds up into global structure.

About

I am an undergraduate in the School of Mathematical Sciences at Fudan University, on a First-Class Undergraduate Scholarship (top 5%), and I am now applying for PhD study to begin in Fall 2027.

Since July 2025 I have been a research assistant at New York University Abu Dhabi, working with Prof. Yi Fang, where the work on this page was done. I lead each project end to end, from the question to the experiments, the writing and the rebuttal, and I build the systems myself: multi-GPU training, simulation environments, and deployment on real robots, a Unitree G1 and Go2, an xArm7 and a Franka FR3. My training in mathematics shows in the tools I reach for: optimal transport to match two sequences, factorizations chosen for what they rule out.

Alongside this, I intern on site at Liqing Intelligence, a world-model startup, and run an NSFC undergraduate project on grading disc degeneration from MRI in collaboration with Shanghai Ninth People’s Hospital.

Contact

NYU  jh10472@nyu.edu
Fudan  23307110151@m.fudan.edu.cn
Phone  +86 18155912005 (same as WeChat)