Anonymous author(s)
Accurate final grasp alignment remains challenging for edge-prominent objects such as thin plates, discs, and rods, whose sparse contacts are easily occluded and poorly resolved by depth sensing. We present TacRefineNet, a tactile-only, goal-conditioned framework for local refinement along tactilely observable pose dimensions. Given current and target multi-finger tactile images and their corresponding hand-joint configurations, a Siamese policy network directly predicts corrective wrist pose increments. The hand iteratively opens, moves, and regrasps, forming an external-dexterity tactile servoing loop. Cross-combination training pairs current and target samples, allowing targets within the sampled pose range to be specified without retraining. We collect 156,007 simulated samples from 15 plates, discs, and rods and train the policy entirely in MuJoCo before zero-shot deployment to an 11-DoF five-fingered hand with piezoresistive sensors. On seen objects, the real system achieves 80.7% and 59.3% success under the 10∘/10 mm criterion for fixed and random targets, respectively; after five steps, the mean errors are approximately 5.2 mm and 3.5∘. Experiments further show continuous correction under long-horizon perturbations and limited within-category transfer to unseen objects, with reduced performance for symmetric or weakly discriminative contacts.
We utilize a piezoresistive tactile sensor array integrated into an 11-DoF dexterous hand. Each fingertip sensor consists of an 11 * 9 taxel grid, where each taxel measures normal contact force. The physical spacing between the taxels on the real tactile fingertip is approximately 1.1 mm. The raw taxel outputs are transformed into tactile images, enabling the use of vision-based encoders for feature extraction.
To evaluate the robustness of our method in dynamic scenarios, we conduct a long-horizon object tracking experiment. A fixed tactile image is provided as the target, while the object’s pose and position are continuously perturbed throughout the sequence. The goal is to assess whether our system can consistently adjust to maintain the desired grasp. The results demonstrate that our method can reliably perform fine-grained grasping toward a specified target pose, even under continuous variations in object pose.