VISTA: VIsually Inferred Spatial ConTact Attention
for Contact-Rich Manipulation
VISTA: VIsually Inferred Spatial ConTact Attention
for Contact-Rich Manipulation
Abstract: Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object-gripper interactions; dedicated tactile or force sensors can provide rich contact information but introduce additional hardware complexity, calibration requirements, and deployment costs. To bridge this gap, we propose VISTA-Policy, an imitation learning paradigm that utilizes the Visual Deformation Field (VDF), a 3D displacement representation of a compliant gripper, as high-dimensional visuo-physical feedback. The framework integrates: 1) a Physics-Aware Encoding Engine for real-time VDF decoding; 2) an Energy Aggregation Denoising Mechanism to isolate true interaction signals; and 3) a Deformation-Augmented Policy Network with incremental gripper actions for precise closed-loop correction. Extensive evaluations on Cross-Scale Object Grasping, Cap Unscrewing, and Calligraphy Writing demonstrate that VISTA-Policy outperforms the strong pure-vision baseline 3D Diffusion Policy and the tactile baseline. VISTA-Policy further demonstrates substantial out-of-distribution generalization to unseen object scales and robustness against dynamic disturbances, offering a durable and cost-effective route toward general-purpose fine-grained manipulation in unstructured environments.
Video
Pipeline
Architecture of VISTA-Policy, comprising three core components: a Physics-Aware Encoding Engine for 3D VDF extraction, an Energy Aggregation Denoising Mechanism for contact denoising, and a Deformation-Augmented Policy Network driven by a relative incremental gripper action space.
An Intuitive Real-Time VDF Visualization Interface
Building upon the VDF representation, we developed an intuitive real-time interface that visualizes gripper deformation and the underlying contact state. During expert demonstration collection, the interface provides operators with clear feedback on physical interactions, allowing them to adapt their actions accordingly. During policy deployment, it also provides an intuitive visualization of the model’s performance.
Task1: Cross-Scale Object Grasping
Gripper Action Representation Ablation (Single-Object Training)
DP3-Wrist-Abs
Barely Lifting (unstable) (3cm)❌
DP3-Wrist-Abs
Premature Lift without Grasp (2cm)❌
VISTA-Abs
Severe Over-Grasping (6cm)❌
VISTA-Abs
Oscillatory Stall without Grasp (2cm)❌
DP3-Wrist
Severe Over-Grasping (5cm) ❌
DP3-Wrist
Premature Lift without Grasp (0.7cm)❌
VISTA (ours)
Adaptive Grasp (5.5cm, Tissue) ✅
VISTA (ours)
Extreme-Scale Zero-Shot Success (0.7cm) ✅
Training Set: Objects with ~4 cm width only. The different deployment results of DP3-Wrist-Abs and VISTA-Abs on a 2 cm object expose a fundamental conflict between absolute gripper action spaces and VDF feedback. Constrained by dataset boundaries (>4cm), the absolute policy struggles to breach this numerical boundary to generate smaller apertures during deployment. However, lacking physical contact, the VDF feedback remains zero, deviating significantly from the successful grasp profile learned during training. Consequently, the policy refuses to proceed to the lifting phase, inducing a deadlock stall. Reconfiguring the action space into relative increments fundamentally liberates the system from numerical boundaries, enabling the policy to utilize the VDF as the key criterion to drive the gripper to any required width for robust grasping.
Multi-Object Training
DP3
Premature Lift without Grasp (0.7cm)❌
TDF-DM
Over-Grasping with Delayed Adjustment (5.5cm, Tissue) ❌
TDF-DM
Premature Lift without Grasp (0.7cm)❌
VISTA (ours)
🌟Extreme-Scale Zero-Shot Success (0.7cm) ✅
Despite multi-object training, the baselines still struggle with narrow targets and over-grip larger objects. VISTA instead uses VDF-based contact-state guidance to adaptively regulate grasping across both seen and unseen scales. As highlighted in the rightmost 🌟 video, VISTA can still successfully grasp an unseen and extremely narrow object (which means the global point-cloud representation is highly sparse) through continuous gripper adjustment.
🌟 Robustness and Recovery Evaluation
DP3-Wrist
Failure under Disturbance❌
Card: Over-Gripping❌
Tofu: Failure to Lift❌
Tofu: Structural Crushing❌
VISTA (ours)
Self-Recovery under Disturbance✅
Card: Adaptive Horizontal Grasp✅
Card: Adaptive Vertical Grasp✅
Tofu: Damage-Free Grasp✅
During disturbance trials, the DP3-Wrist baseline shows no awareness after object dislodgement and continues elevating the arm, whereas VISTA immediately triggers a self-recovery regrasp (succeeding in 4/5 attempts)—highlighting the advantage of VDF for contact state perception. When handling fragile objects like soft tofu and playing cards, the baseline suffers from precise grasping regulation issues, causing over-gripping, lifting failure, or structural crushing. VISTA, by contrast, utilizes VDF feedback to perform damage-free, adaptive grasps, confirming its superior robustness in delicate manipulation tasks.
Task2: Cap Unscrewing
Training on Minimum Scale (2 cm)
Training on Maximum Scale (6 cm)
DP3-Wrist
Zero-Shot Failure on Large Cap ❌
VISTA (ours)
Zero-Shot Success on Large Cap ✅
DP3-Wrist
Zero-Shot Failure on Small Cap ❌
VISTA (ours)
Zero-Shot Success on Small Cap ✅
Boundary-Scale Generalization: Each policy is trained on only one boundary-scale cap (2 cm or 6 cm) and evaluated across a range of cap sizes. DP3-Wrist shows clear performance degradation on unseen scales, while VISTA dynamically regulates its grasp based on VDF feedback and maintains more stable cross-scale performance.
Multi-Object Training
DP3-Wrist
Asymmetric Force❌
Dislocation❌
Small Scale Failure❌
Secondary Twisting Failure❌
VISTA (ours)
Bilateral Force Balance✅
Large Scale Success✅
Small Scale Success✅
Secondary Re-Twisting✅
TDF-DM
Successful Unscrewing ✅
Unscrewed but Dislocated❌
Dislocation (Small Cap) ❌
Dislocation (Large Cap) ❌
The performance of the DP3-Wrist does not improve despite being exposed to all cap sizes during training; issues such as over-gripping and bottle dislocation due to asymmetric contact persist. During re-alignment, severe axial deviation can render subsequent twisting ineffective. TDF-DM also exhibits frequent bottle dislocation, as its rigid tactile end-effector is sensitive to small pose errors and can generate excessive interaction during twisting. In comparison, VISTA maintains smoother overall motions and successfully performs secondary attempts to open the caps.
Task3: Calligraphy Writing
Main Experiment
DP3-Wrist
Vertical Jitter❌
DP3-Wrist
Excessive Pressing❌
VISTA
Stable Writing✅
VISTA
Stable Writing✅
TDF-DM
Failed to Hold the Brush ❌
TDF-DM
Excessive Pressing❌
TDF-DM
Swept Above the Paper ❌
TDF-DM
Barely Contacted the Paper ❌
🌟 Robustness and Recovery Evaluation
DP3-Wrist
OOD Failure❌
DP3-Wrist
Dyn. Disturbance Failure❌
VISTA
OOD Success ✅
VISTA
Dyn. Disturbance Success ✅
The calligraphy writing task requires precise z-axis compliance and reliable contact perception. DP3-Wrist suffers from unstable paper contact, while TDF-DM suffers from modality inconsistency, making it difficult to extract unified contact-state regularities. Conversely, VISTA demonstrates robust performance, maintaining reliable adaptability even under out-of-distribution (OOD) paper heights and dynamic disturbances.
Perception Module Effectiveness Validation
Canny
Raw
Ours
We compare Canny, Raw (our ablated baseline), and Ours (the full perception pipeline). Our method maintains more accurate and continuous edge extraction under illumination changes and partial occlusions, outperforming the baselines both visually and quantitatively.
Robust Edge Extraction During Contact
Our full perception pipeline maintains accurate and continuous extraction of the gripper’s anterior inner edges under illumination changes and partial occlusions, ensuring reliable VDF reconstruction.
Contact State Switching Sensitivity Test
The contact confidence stays near zero during free motion and rapidly switches to one upon valid contact. EMA smoothing suppresses transient noise while preserving sensitive and stable contact-state transitions.
Conclusion
Summary: VISTA introduces a novel imitation learning paradigm that extracts the observable visual deformation of the passive compliant gripper as visuo-physical feedback. Evaluated across challenging contact-rich tasks—including Cross-Scale Object Grasping, Cap Unscrewing, and Calligraphy Writing—VISTA-Policy achieves superior OOD generalization and robustness, significantly outperforming both pure-vision and hardware-tactile baselines.
Key Insight: Simply scaling demonstration data is insufficient for policies to implicitly master contact dynamics. Embedding Visual Deformation Fields (VDF) as a low-cost physical prior provides an effective and scalable pathway toward general contact-rich robotic manipulation.