Junyi Hu

Undergraduate, School of Mathematical Sciences, Fudan University

Research Assistant, New York University


Research

Aligning what a model sees with what it must say or do

I work on multimodal alignment. Across quite different problems, the same question keeps returning: what is the right correspondence between a continuous visual stream and the discrete structure a downstream model actually needs β€” a sentence, or an action? Two threads follow from it.

🀟 Sign language understanding

Gloss-free translation, where the decoder is normally left to discover a latent cross-modal permutation on its own β€” and the large-scale, structure-preserving data needed to make any of it hold up outside near-frontal studio conditions.

πŸ€– Embodied navigation

Learned, context-dependent safety criteria for robots acting from raw egocentric RGB-D. Privileged geometry supervises training offline; the deployed policy builds no map and needs no per-scene tuning.

Publications

Selected work

The VTaMo pipeline: frozen CLIP-ViT encoder, Sinkhorn optimal-transport local alignment, orthogonal global alignment, and a LoRA-adapted Flan-T5 decoder.

VTaMo: Video–Text Alignment Model for Sign Language Translation

Junyi Hu, Zhewen He, Haomian Huang, Aoxiang Yang, Yi Fang

ECCV 2026 First author Gloss-free SLT

Gloss-free translation asks the decoder to learn translation and silently discover a latent cross-modal permutation at once β€” fragile, because sign languages do not follow spoken word order. VTaMo makes the alignment explicit at three granularities: entropy-regularized optimal transport with a learnable null token for local frame↔token correspondence, a learnable orthogonal transform over a memory queue for global calibration, and position-aligned contrastive learning over reordered features. A text-only recovery model restores fluent sentences from the decoded pseudo-gloss.

Overview of the SignNet-1M augmentation pipeline: novel-view rendering, scene editing and identity reenactment.

SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks

Zhewen He, Junyi Hu, Haomian Huang, Zhenhua Li, Yu-Shen Liu, Yi Fang

ECCV 2026 Co-author ~1M clips Β· 2,058 h

Existing SLT benchmarks are shot near-frontal, in studios, with a handful of signers, so state-of-the-art models break under viewpoint, background and signer shift. SignNet-1M is ~1M augmented clips across ASL, DGS and CSL from seven source corpora, varied along three structure-preserving axes β€” 3DGS novel-view rendering, diffusion-based scene editing, and cross-reenactment identity substitution β€” each tagged with its factor axis and severity. The paired Orig / Zero-shot / Trained protocol separates the coverage blind spot of existing benchmarks from the training value of augmented data; training on SignNet-1M recovers up to +14.71 BLEU-4.

A diffusion policy proposes candidate trajectories from RGB-D; a learned safety critic scores and selects one.

Learning Adaptive Safety Margins for Visual Navigation

Junyi Hu*, Shuaihang Yuan*, Geeta Chandra Raju Bethala, Yi Fang, Anthony Tzes

IROS 2026 Co-first author* Real-robot

A fixed clearance margin is wrong in both directions: set it high and the robot detours in open space, set it low and it cuts corners in clutter. A diffusion policy proposes K = 16 candidate trajectories from egocentric RGB-D, and a learned critic carrying a trajectory-dependent clearance budget scores them. ESDF geometry is privileged information used only offline to supervise the teacher β€” the deployed selector runs from RGB-D alone, with no map building. 0.783 SR / 0.611 SPL on HM3D against NavDP's 0.710 / 0.529, and 8/10 on the hardest real-world scene with a Unitree G1 trained purely in simulation.

Contact

Get in touch

NYU  jh10472@nyu.edu
Fudan  23307110151@m.fudan.edu.cn