FINGR: Learning Dexterous Hand Control for Real-World Rubik’s Cube Solving

Yutong Liang*, Quanquan Peng*, Matthew Kim*, Xiaolong Wang

University of California San Diego

*Equal contribution

Video

Abstract

Manipulating a Rubik's Cube with one single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next move. To learn this coordination, we introduce FINGR (Future-supervised Interaction Network with Geometric Representations), a policy that combines finger-relative geometry with future interaction prediction. A shared point encoder expresses the cube relative to each fingertip and aggregates its points without depending on cubie indexing. Learned future tokens share the observation encoder and receive supervision for contact-force changes, layer-turn progress, and finger joint displacement at multiple time scales. The resulting representation conditions a flow policy that directly generates finger actions. On a real dexterous hand, our policy achieves around 99.0% success over 300 turn attempts, compared with 79.7% for the base flow policy. Integrated with grasping and table-assisted regrasping, the policy solves all ten scrambled 2×2×2 cubes in a mean complete-system time of approximately 137 seconds.

Method

FINGR pipeline Finger-relative cube geometry, finger state, tactile images, and cube tokens enter a transformer encoder. Future queries learn contact-force change, layer-turn progress, and finger motion. The encoded representation conditions the action head.

Finger-relative geometry describes the cube from each fingertip. Predicting future contact, turn progress, and finger motion shapes the representation used by the flow policy.

Experiment

Real Robot Setup

2×2 Single World-Record ScrambleAutonomous solve · Original speed
Human Random ScrambleShuffled by a person, solved autonomously · Original speed

Complete Solves Trajectory Visualization

ScrambleR' F U2 R2 U2 R' F U' F R2 U2

Ready to explore0:00 / 0:00
PickupLayer turnsAlignmentRegraspPlace / returnObserve
Execution

Results

Layer-turn Success

300 attempts / method

FINGR · 99.0% success · 5.25 s per attempt

Complete Solves

PickupLayer turnsAlignmentRegraspPlace / returnObserve
050100150200 s
Solve 01 · 140.87 s
Pickup
8.72 s
Layer turns
90.20 s
Alignment
1.80 s
Regrasp
24.95 s
Place / return
11.40 s
Observe
3.80 s

BibTeX

@article{fingr2026,
  title={FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving},
  author={Yutong Liang and Quanquan Peng and Matthew Kim and Xiaolong Wang},
  journal={arXiv preprint arXiv:2609.33973},
  year={2026},
  url={https://arxiv.org/abs/2609.33973}
}