SignDino

Self-Supervised Sign Language Representation Learning
via Temporal-Axis Self-Distillation

Junyi Hu  ·  Zhewen He  ·  Haomian Huang  ·  Zhenhua Li  ·  Zhifei Li  ·  Yi Fang†

New York University Abu Dhabi

† corresponding author  ·  jh10472@nyu.edu

Left: image DINOv3 samples spatial global and local crops inside a single image. Right: SignDino samples temporal global and local crops along the tracked left-hand, right-hand and face rows of a single anatomical stream.
The local–global distillation game, moved from image space to the time axis. Image DINOv3 (left) samples spatial global and local crops inside one image. SignDino (right) samples temporal global (Tg ∈ [64, 96]) and local (Tl ∈ [10, 32]) crops along the tracked hand and face rows of a single anatomical stream, with iBOT masks on a few frame slots inside the local crop.

Abstract

Self-supervised vision learns structure by playing a local–global game across the spatial axis of an image: a crop and the whole picture are asked to agree. Structured motion has an axis of its own, and the same game should be playable along it — where the counterpart of a local crop is not another view of the same frame, but a shorter temporal view of the same tracked part.

We introduce SignDino, which moves the DINOv3 student–teacher recipe from the spatial domain of image crops to the temporal domain of tracked streams. Two properties of a moving body drive the design, and neither is specific to sign: motion is produced by a small set of anatomically distinct parts, and meaning lives in how those parts are organised over time. Each video is therefore decomposed into left-hand, right-hand and face streams by a detector-first YOLOv8n + ByteTrack pipeline; a frozen DINOv3 ViT-B/16 embeds each per-frame crop once; and the only trained components are lightweight temporal Transformers forming the student and EMA teacher.

Every term in the objective — temporal DINO self-distillation, iBOT-style masked-frame prediction, KoLeo feature spreading, Gram anchoring of the frame-to-frame similarity structure — is inherited unchanged from image self-supervision. What changed is the axis the local–global game is played on. Strong image-level primitives stay fixed and the model learns only how articulator states evolve across time. Sign language is where we evaluate rather than what the method was built to fit: across sign-to-English translation, isolated sign recognition and fingerspelling detection, SignDino provides a strong public self-supervised representation and is competitive with or ahead of the state of the art under matched downstream evaluation — and the phonological probe, where all sixteen ASL-Lex features improve, is the result that says what the representation actually learned.

01

Temporal-axis DINO

Per-frame articulator tokens replace spatial patch tokens; multi-temporal-crop replaces multi-spatial-crop; iBOT masks frames rather than patches.

02

Three anatomical streams

Left hand, right hand and face are encoded independently over time and fused only at the downstream head, so each stream learns its own dynamics.

03

Detector-first crops

YOLOv8n + ByteTrack recovers hands that per-frame detection misses under motion blur and self-occlusion, instead of relying on pose-driven crops.

Method

SignDino is three independently trained per-stream SSL encoders sharing one upstream anatomical-crop pipeline. A frozen DINOv3 ViT-B/16 turns every per-frame crop into a 768-d CLS vector once, and the vectors are cached; the student and teacher are small temporal Transformers on top of that cache. Only ≈30M parameters ever receive an SSL gradient, and the image backbone receives none.

SignDino pipeline: video to YOLOv8n plus ByteTrack to per-frame anatomical streams, then a frozen DINOv3 encoder to per-frame embeddings, then per-stream temporal SSL, then downstream heads.
Full pipeline. Each source video passes through the detector-tracker, yielding synchronised per-frame left-hand, right-hand and face streams. A frozen DINOv3 ViT-B/16 encodes every crop into a 768-d CLS token, giving three per-frame embedding sequences that each feed one independent temporal SSL trainer. The stage-2 teachers become the per-stream sign encoders that the translation, ISLR and fingerspelling heads consume as fused per-frame features.

The temporal student–teacher game

The teacher sees only long temporal crops of a stream; the student sees both long and short crops, with a random fraction of frame positions replaced by a learnable [MASK] token. Stage 1 combines the crop-level DINO cross-entropy, the frame-level iBOT cross-entropy and the DKoleo spreading term. Stage 2 freezes a snapshot of the teacher as a Gram teacher and adds a loss that matches the student's frame-to-frame similarity matrix to the snapshot's — preserving which frames of a clip are alike, which is the structure that fingerspelling and continuous translation rely on.

Diagram of the temporal student/teacher game: global and local crops, EMA teacher with Sinkhorn-Knopp, student with softmax, stage 1 losses and stage 2 Gram anchoring.
Temporal student/teacher game for one anatomical stream. The frozen DINOv3 backbone supplies per-frame embeddings once; the SSL game runs entirely on top of them. The teacher is the EMA of the student.

Why the detector-tracker matters

Single-frame detection drops hands under fast motion and self-occlusion. ByteTrack's Kalman fill recovers them, reaching ~95% per-frame hand detection on How2Sign. Removing it costs −1.5 BLEU on How2Sign and −0.094 R@1 on ASL Citizen.

Six OpenASL frames shown twice: YOLOv8n alone on top misses hands that YOLOv8n plus ByteTrack recovers on the bottom.
Detector-first crops recover hands YOLOv8n alone misses. Same frame under YOLOv8n only (top) and YOLOv8n + ByteTrack with Kalman fill (bottom). Green boxes are active detections; magenta are tracker predictions for hands the detector missed at that frame.

Results

Everything below follows the SHuBERT public-data protocol: the same 984 source hours of decontaminated YouTube-ASL, the same downstream datasets, task heads, schedules and metrics. Only the upstream representation changes.

Sign language translation

MethodPT hrs How2Sign BLEUHow2Sign BLEURT OpenASL BLEUFLEURS-ASL BLEU
YouTube-ASL fine-tune98412.446.6——
YouTube-SL-25320715.447.9—4.4
SSVP-SLT105415.549.6——
Uni-Sign98414.949.423.1—
SHuBERT98416.249.923.24.7
SignDino (frozen)98416.850.423.64.9
SignDino (live fine-tune)98417.951.324.55.3
Corpus BLEU-4 and BLEURT-20. FLEURS-ASL is zero-shot. The frozen row already beats the previous live-fine-tuned state of the art, so the temporal-axis features carry the signal without any encoder adaptation.

Isolated sign recognition & fingerspelling

MethodTrainable ASL Citizen R@1Sem-Lex R@1 WLASL2000 P-IASL-STEM mIoU
I3D25M0.63———
SignCLIP217M0.600.30——
Uni-Sign580M——63.52—
SHuBERT (rank-1 LoRA)0.17M0.650.5460.900.40
SignDino (rank-1 LoRA)0.17M0.670.5666.90.41
SignDino (full fine-tune)all0.7040.59368.50.43
At the matched 0.17M-parameter LoRA operating point SignDino beats Uni-Sign on WLASL2000 by +3.3 P-I with roughly one one-thousandth of its trainable parameters.

Phonological feature probing

Sixteen ASL-Lex 2.0 feature classifiers, trained on frozen features, ask whether the representation preserves sub-lexical structure rather than only whole-sign identity. SignDino improves on every one of the sixteen features, with the largest gains on the temporally articulated ones — path movement, wrist twist, repeated movement.

Average over 16 featuresSem-Lex R@1ASL Citizen R@1
SHuBERT79.3885.02
SignDino87.64 +8.2691.85 +6.83

What the encoder learns

Five k-means clusters of right-hand features, each row showing ten nearest-to-centroid crops sharing a handshape across different signers and backgrounds.
Zero-shot clusters on held-out frames. Each hand stream surfaces handshape clusters — pointing, spread fingers, thumb-up, two-hand contact — that recur across many signers, backgrounds and skin tones, evidence that the encoder internalises articulator structure rather than identity or background.
Two t-SNE panels: family-relation nouns cluster centrally with weather nouns far away; perception verbs, motion verbs and sleep occupy distinct regions.
Cross-sentence t-SNE after translation fine-tuning. Family-relation nouns concentrate while weather nouns sit far away (left); perception verbs, motion verbs and sleep occupy distinct regions (right).

Code

The full implementation — SSL losses, the temporal encoder, the multi-temporal-crop sampler, the embedding cache, and all three downstream heads — is on GitHub under the MIT licence. The YAML configs match the paper's hyperparameter card line by line.

git clone https://github.com/junyi2005/signdino.git
cd signdino && pip install -r requirements.txt

python smoke_test.py                       # synthetic end-to-end check, no data needed

python preprocess/cache_embeddings.py --crop-root /path/to/CROPS   # one-off frozen-DINOv3 cache
python train.py --config configs/pretrain.yaml --stream lh --exp-name pretrain_lh   # stage 1
python train.py --config configs/refine.yaml   --stream lh --exp-name refine_lh \
    --init-ckpt runs/pretrain_lh/ckpt_e100.pt                                       # stage 2

Datasets and third-party checkpoints (DINOv3, ByT5-Base, BLEURT-20, the YOLOv8n detectors) are not redistributed here; obtain each from its official release under its own terms.

Citation

@article{hu2026signdino,
  title   = {SignDino: Self-Supervised Sign Language Representation Learning
             via Temporal-Axis Self-Distillation},
  author  = {Hu, Junyi and He, Zhewen and Huang, Haomian and Li, Zhenhua and
             Li, Zhifei and Fang, Yi},
  journal = {arXiv preprint},
  year    = {2026}
}