VISTA: VIsually Inferred Spatial ConTact Attention
for Contact-Rich Manipulation
VISTA: VIsually Inferred Spatial ConTact Attention
for Contact-Rich Manipulation
Abstract: Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provide rich contact information but introduce additional hardware complexity, calibration requirements, and deployment costs. To bridge this gap, we propose VISTA-Policy, an imitation learning paradigm that utilizes the Visual Deformation Field (VDF), a 3D displacement representation of a compliant gripper, as high-dimensional visuo-physical feedback. The framework integrates: 1) a Physics-Aware Encoding Engine for real-time VDF decoding; 2) an Energy Aggregation Denoising Mechanism to isolate true interaction signals; and 3) a Deformation-Augmented Policy Network with incremental gripper actions for precise closed-loop correction. Extensive evaluations on Cross-Scale Object Grasping, Cap Unscrewing, and Calligraphy Writing demonstrate that VISTA-Policy outperforms the strong pure-vision baseline 3D Diffusion Policy and the tactile baseline. VISTA-Policy further demonstrates substantial out-of-distribution generalization to unseen object scales and robustness against dynamic disturbances, offering a durable and cost-effective route toward general-purpose fine-grained manipulation in unstructured environments.
Video
Pipeline
Architecture of VISTA-Policy, comprising three core components: a Physics-Aware Encoding Engine for 3D VDF extraction, a spatial energy aggregation mechanism for contact denoising, and a multi-modal policy backbone driven by a relative increment gripper action space.
Task1: Cross-Scale Object Grasping
Action Space Representation Ablation
Training Set: Objects with ~4 cm width only. The different deployment results of DP3-Wrist-Abs and VISTA-Abs on 2 cm object exposes a fundamental conflict between absolute gripper action spaces and VDF feedback. Constrained by dataset boundaries (>4cm), the absolute policy struggles to breach this numerical boundary to generate smaller apertures during deployment. However, lacking physical contact, the VDF feedback remains zero, deviating significantly from the successful grasp profile learned during training. Consequently, the policy refuses to proceed to the lifting phase, inducing a deadlock stall. Reconfiguring the action space into relative increments fundamentally liberates the system from numerical boundaries, enabling the policy to utilize the VDF as the sole criterion to drive the gripper to any required width for robust grasping.
DP3-Wrist-Abs (3cm)
Barely lifting (unstable)❌
DP3-Wrist-Abs (2cm)
Lift directly without grasp ❌
VISTA-Abs (6cm)
Severe over-grasping ✅
VISTA-Abs (2cm)
Oscillatory Stall without grasp✅
🌟 Robustness and Resiliency Evaluation
During disturbance trials, the DP3-Wrist baseline shows no awareness after object dislodgement and continues elevating the arm, whereas VISTA immediately triggers a self-recovery regrasp (succeeding in 4/5 attempts)—highlighting the advantage of deformation fields for contact state perception. When handling fragile objects like soft tofu and playing cards, the baseline suffers from precise grasping regulation issues, causing over-gripping, lifting failure, or structural crushing. VISTA, by contrast, utilizes deformation field feedback to perform damage-free, adaptive grasps, confirming its superior robustness in delicate manipulation tasks.
DP3-Wrist (baseline)
Failure under Disturbance❌
Card: Over-Gripping❌
Tofu: Failure to Lift❌
Tofu: Structural Crushing❌
VISTA-Policy (ours)
Self-Recovery under Disturbance✅
Card: Adaptive Horizontal Grasp✅
Card: Adaptive Vertical Grasp✅
Tofu: Damage-Free Grasp✅
Task2: Cap Unscrewing
Multi-Object Training
The performance of the DP3-Wrist does not improve despite being exposed to all cap sizes during training; issues such as over-gripping and bottle dislocation due to asymmetric forces persist. During the re-alignment process, the policy frequently suffers from severe axial deviation, rendering subsequent actions completely ineffective. In comparison, VISTA maintains smoother overall motions and successfully performs secondary attempts to open the caps.
DP3-Wrist (baseline)
Asymmetric Force❌
Dislocation❌
Small Scale Failure❌
Secondary Twisting Failure❌
VISTA-Policy (ours)
Bilateral Force Balance✅
Large Scale Success✅
Small Scale Success✅
Secondary Re-Twisting✅
Task3: Calligraphy
The calligraphy task demands stringent z-axis compliance and real-time contact perception. Without the deformation field serving as an explicit physical state to indicate the writing phase, baselines often fail by either floating above the paper, oscillating, or over-pressing. Conversely, VISTA demonstrates robust performance, maintaining reliable adaptability even under out-of-distribution (OOD) paper heights and dynamic disturbances.
Main Experiment
DP3-Wrist
Vertical Jitter❌
DP3-Wrist
Excessive Pressing❌
VISTA
Stable Writing✅
VISTA
Stable Writing✅
🌟 Robustness and Resiliency Evaluation
DP3-Wrist (OOD)
Failure❌
DP3-Wrist (Dyn. Disturbance)
Failure❌
VISTA (OOD)
Success ✅
VISTA (Dyn. Disturbance)
Success ✅
Conclusion
Summary: VISTA introduces a novel imitation learning paradigm that extracts the observable visual deformation of passive compliant grippers as visuo-physical feedback. Evaluated across challenging contact-rich tasks—including Cross-Scale Object Grasping, Cap Unscrewing, and Calligraphy Writing—VISTA-Policy achieves superior OOD generalization and robustness, significantly outperforming both pure-vision and hardware-tactile baselines.
Key Insight: Simply scaling demonstration data is insufficient for policies to implicitly master contact dynamics. Embedding Visual Deformation Fields (VDF) as a low-cost physical prior provides an effective and scalable pathway toward general contact-rich robotic manipulation.