WING: World Action Learning viaINteraction-Centric Spectral Latent Guidance

From human experience to robot intelligence.

Zhiming Liu†Yikun Miao†Ying Chen†Hongrui YinFangqi ZhuXiaoyi PangQuanxin ShouZhengyang YanHaodong WangSong Guo*
The Hong Kong University of Science and Technology, Hong Kong SAR, China

† Equal contribution* Corresponding author

SCROLL TO EXPLORE

HUMAN EXPERIENCE. ROBOT ACTION.

At a Glance

WING overview: transferring interaction knowledge from egocentric videos to robot actions through interaction-centric spectral latent guidance, with results on LIBERO, RoboTwin 2.0, RoboCasa and real-world tasks.

Learning what transfers across embodiments.

01

Focus on interaction

Decouple interaction-relevant signals from task-irrelevant motion in egocentric videos, and distill them into latent actions.

02

Find shared dynamics

Identify slowly varying temporal structures shared by human interactions and robot behaviors in the spectral domain.

03

Guide robot actions

Use the shared spectral latent action to guide robot action generation at inference time.

99.20%LIBERO
93.80%RoboTwin 2.0
57.7%RoboCasa–GR1

Average success rates reported in the paper.

WING simulation performance

Success rate (%) · 0–100 scale

LIBERO

Four-suite average

+0.70 pp over the strongest shown baseline

RoboTwin 2.0

Clean + randomized average

+0.20 pp over the strongest shown baseline

RoboCasa–GR1

24-task tabletop average

+2.30 pp over the strongest shown baseline

Hover, focus, or tap a bar for details. All bars use the same 0–100% scale.

HOW IT WORKS

Methodology of WING

Learn interactions from human videos. Turn shared dynamics into robot actions.

MOTIVATION 1

Head motion in egocentric videos

Egocentric videos capture rich hand–object interactions, but the camera also moves with the wearer's head. Looking around or leaning forward shifts the entire scene, mixing observer motion with the motion of the hands and objects.

Egocentric view 01
Egocentric view 02

A latent action model trained to reconstruct the next frame must explain both kinds of change. Its encoding can therefore absorb head motion alongside the interaction itself. This observer-specific motion does not describe the manipulation skill we want a robot to learn, and can obscure the action information that should transfer.

Our goal: preserve hand–object interactions while suppressing camera motion.

OUR SOLUTION

WING-LAM: interaction-centric latent actions

WING-LAM explicitly separates camera motion from interaction motion, then distills the interaction representation into a latent action model that takes only RGB frame pairs.

WING-LAM: geometric routing separates interaction and background tracks into interaction and camera latent branches; the interaction representation is distilled into a model trained on original and camera-warped RGB frame pairs.
Geometric routing separates motion; distillation transfers interaction cues to an RGB-only model.

1. Separate the motion signals. Point tracks and hand–object region masks distinguish background motion from interaction motion. Background tracks estimate a global camera warp. Compensating for this warp reveals the residual motion associated with hand–object interactions.

2. Learn two latent branches. A teacher model encodes camera motion and interaction motion separately. The camera branch predicts the global warp, while the interaction branch predicts the residual displacements. Together, the decoded motions reconstruct observed trajectories.

3. Distill the interaction representation. An RGB-only student learns the teacher's interaction latent from frame pairs. We also apply camera-like warps that preserve the interaction target, encouraging the student to produce consistent latent actions despite changes in viewpoint. The resulting model extracts interaction-centric latent actions without requiring point tracks or region masks as inputs.

These latent actions preserve manipulation structure with less sensitivity to observer motion, providing the foundation for WING's spectral latent guidance.

MOTIVATION 2

Shared interactions, different execution rhythms

Removing camera motion does not eliminate the gap between human and robot behavior. Even when both perform the same interaction, they can differ in speed, pauses, and the timing of individual movements. Interaction-centric latent actions still encode these temporal differences, making direct transfer of the full latent sequence challenging.

We therefore examine latent action sequences in the frequency domain. Our analysis finds stronger correspondence between human and robot interactions in their low-frequency components: the slowly varying temporal structure is more consistent across embodiments than the full trajectory.

Our goal: transfer the shared temporal structure while allowing the robot to generate fine-grained actions suited to its own embodiment.

OUR SOLUTION

Spectral latent action guidance

WING transforms interaction-centric latent action sequences into the frequency domain and retains their lowest-frequency components as a structured guidance target for robot action generation.

INTERACTION → SPECTRAL GUIDANCE

Encode interactions over time

WING-LAM encodes successive frame pairs. The resulting sequence contains both slowly varying structure and rapid temporal changes.

Illustrative signal · Select a stage or adjust K to explore. At inference, WING predicts the retained coefficients to guide robot actions.

1. Extract the shared temporal structure. We apply a discrete cosine transform (DCT) along the time axis of the latent sequence. Keeping the lowest K frequencies suppresses rapid variations and preserves shared human–robot dynamics.

2. Predict guidance from the current context. The spectral target is constructed from future frame pairs during training. At inference time, a lightweight predictor estimates it from the current image, language instruction, and robot state, so future observations are not required.

3. Guide robot action generation. The predicted frequency coefficients are projected into guidance tokens. The action model attends to these tokens through gated residual updates, using the shared spectral latent action to guide the generation of executable robot action chunks.

Low-frequency truncation is applied to the latent guidance. The robot's actions retain the fine-grained control needed to complete the task.

FROM REPRESENTATIONS TO ROBOT CONTROL

What WING does

Four questions about learning from human experience—and turning it into robot actions.

  1. Q1What structures in egocentric latent actions reveal transferable interaction knowledge?
  2. Q2Can WING effectively transfer latent action knowledge to improve robot performance?
  3. Q3How do interaction-centric latent actions and spectral latent action guidance contribute?
  4. Q4How effective is WING’s pretraining strategy for downstream robot policy learning?

QUESTION 01 · TRANSFERABLE STRUCTURE

What is worth transferring?

Useful latent actions should describe the interaction, rather than the observer’s movement. WING-LAM preserves action semantics while reducing camera-motion information and sensitivity to camera perturbations. Its frozen representations achieve the highest average action-classification accuracy in this comparison.

Across human–robot pairs, low-frequency latent components separate matching interactions more clearly from mismatched ones. The shared slow structure provides a stronger transfer signal than the full trajectory or its high-frequency components.

QUESTION 02 · ROBOT PERFORMANCE

Does human experience improve robot control?

WING achieves the highest average success rate among the methods reported in the paper on LIBERO, RoboTwin 2.0, and RoboCasa–GR1. The tables below compare every method and evaluation setting reported in the main simulation results.

From simulation to four real-world tasks

On the dual-arm robot, WING reaches 75.0% average success in the standard setting. Under generalization, it remains competitive with π₀.₅ in average success and attains a slightly higher average progress score. Success measures completion; progress credits partial execution.

QUESTION 03 · CORE COMPONENTS

Why do both components matter?

Interaction-centric encoding and spectral guidance make complementary contributions. These controlled ablations exclude egocentric WAM pretraining. Without WING-LAM, we use DreamDojo’s latent encoder; without DCT, we guide the policy with the full time-domain latent trajectory.

Enough temporal structure, without the full spectrum

Keeping four DCT coefficients performs best across all four LIBERO suites. Two coefficients lose useful dynamics; all eight add high-frequency variation without gains.

Success rate (%) · K = 8 retains the full spectrum. K = 4 performs best in every suite. Figure 5(a) ↗

QUESTION 04 · EGO PRETRAINING

How should the model learn from egocentric video?

All strategies use the same 200-hour egocentric pretraining budget. Predicting spectral guidance works better than video-only learning or treating latent actions as direct action targets. Removing interaction-centric debiasing also reduces performance in most settings.

Human latent actions are most useful as guidance for action generation, while the robot policy retains its own action-prediction objective.

QUALITATIVE RESULTS

Additional experiments

What does WING-LAM encode? Across different scenes, decoded motion concentrates on hands and their interactions with objects.

BEYOND THE TRAINING SETTING

Real-world generalization

THE RESEARCH

Abstract

Learning general-purpose robot policies requires large-scale real-world interaction data, yet collecting robot demonstrations through teleoperation remains expensive and difficult to scale. Egocentric videos provide a rich source of human interaction experience that shares task-relevant semantics with robotic manipulation, creating an opportunity to align human and robot actions in a shared latent action space for cross-embodiment knowledge transfer. However, existing latent action approaches typically infer actions through reconstruction between consecutive frames, which are not inherently interaction-centric and can be dominated by nuisance variations, such as ego-camera motion. Moreover, although human interactions and robot actions share interaction semantics, they often exhibit substantially different temporal dynamics, making direct transfer to robot policies challenging. To address these challenges, we propose WING (World Action Learning via INteraction-Centric Spectral Latent Guidance), a framework for transferring interaction knowledge from egocentric videos to robot policies. First, we introduce an interaction–motion decoupling mechanism that separates interaction-relevant signals from task-irrelevant motion and selectively distills the interaction-centric components into latent actions. Second, motivated by the observation that cross-embodiment task semantics are primarily encoded in slowly varying temporal structures, we identify shared components between egocentric latent actions and robot behaviors in the spectral domain and use them as guidance for action generation at inference time. With these designs, WING achieves average success rates of 99.20% on LIBERO, 93.80% on RoboTwin 2.0, and 57.7% on RoboCasa–GR1, while also demonstrating strong performance across four real-world tasks under diverse generalization settings. These results demonstrate that WING can effectively distill embodied interaction knowledge from large-scale egocentric videos and transfer it to robot control, providing a scalable pathway for acquiring physical interaction knowledge from human experience.

Read the full paper

CITE THIS WORK

BibTeX

Manuscript citation. Publication details will be added when available.

@misc{liu_wing,
  title = {{WING}: World Action Learning via {INteraction-Centric} Spectral Latent Guidance},
  author = {Liu, Zhiming and Miao, Yikun and Chen, Ying and Yin, Hongrui and Zhu, Fangqi and Pang, Xiaoyi and Shou, Quanxin and Yan, Zhengyang and Wang, Haodong and Guo, Song},
  note = {Manuscript}
}