Zhewen He, Junyi Hu, Haomian Huang, Zhenhua Li, Yu-Shen Liu, Yi Fang
ECCV 2026
Co-author
~1M clips Β· 2,058 h
Existing SLT benchmarks are shot near-frontal, in studios, with a handful of signers, so
state-of-the-art models break under viewpoint, background and signer shift. SignNet-1M is ~1M
augmented clips across ASL, DGS and CSL from seven source corpora, varied along three
structure-preserving axes β 3DGS novel-view rendering, diffusion-based scene editing, and
cross-reenactment identity substitution β each tagged with its factor axis and severity. The paired
Orig / Zero-shot / Trained protocol separates the coverage blind spot of existing benchmarks
from the training value of augmented data; training on SignNet-1M recovers up to +14.71 BLEU-4.