Simulate the gel, then the picture it makes
An optical tactile sensor is a soft gel dome with a camera inside: touch deforms the gel, and the camera sees the result. Simulating one is hard twice over, because the gel deforms a lot and non-linearly, and because its patterned, internally lit surface is hard to render. DOT-Sim splits the problem. The gel is an elastic material simulated with the Material Point Method, whose stiffness parameters are calibrated in minutes by differentiating through the simulator. The image is then predicted as a residual on top of the real sensor's idle image. Classifiers and a control policy trained only on these simulated images work on a real DenseTact sensor without any real training data for the task.
PSNR of simulated tactile images over the strongest baseline, averaged over three settings: 31.3 dB vs 26.7 dB for calibrated Tacto
accuracy on real images for a classifier trained only in simulation, on indenters the renderer never saw in real images (best baseline 52.9%, chance 50%)
average error when a real robot follows a trajectory with a policy trained only in simulation (0.896 ± 0.031 mm over 10 trials)
to calibrate the gel's physical parameters from 19 recordings, on a single GPU

Two stages: physics, then optics
- PhysicsSimulate the gel with MPMThe sensor is a set of particles carrying mass, velocity and deformation, moved through a background grid. A rigid indenter follows its recorded path and pushes the particles aside.
- PhysicsCalibrate by gradient descentYoung's modulus and Poisson's ratio are optimized so that the simulated surface matches the reference deformation (Chamfer distance), with gradients from the differentiable simulator.
- OpticsLook from insideA virtual camera at the centre of the sensor's base casts rays through the gel and records a depth map and a surface-normal map, the way the real camera looks at the gel from inside.
- OpticsPredict the residualA network maps depth and normals to the change of the image relative to the idle frame. Adding it to the real idle image gives the simulated tactile image.

Calibration in detail
The indenter's pose is tracked with a motion-capture system in 19 recordings, poking the sensor at different angles and depths. Measuring the deformed gel directly is impractical (the indenter hides it), so reference deformations come from the Abaqus finite-element solver, which is accurate but too slow to use as the simulator itself.
The MPM simulations of all recordings run in parallel; log E and ν are optimized for 30 iterations, and the median over the recordings is kept. This takes a few minutes on one RTX A5000.
Why a residual
Most of a tactile image does not change during contact: the pattern, the lighting and the colour gradients stay where they are, and the signal is a local change. Predicting only that change, and adding it to a real idle frame, keeps the sensor's own look for free and leaves the network a much easier problem. The network is a DeepLabV3-ResNet50 trained with a plain pixel-wise L2 loss.
How it compares to other tactile simulators
| Simulator | Physics | Backend | Optical simulation | Sim-to-real |
|---|---|---|---|---|
| Tacto | PyBullet | PyBullet | OpenGL | None |
| Taxim | FEM | N/A | Calibrated LUT | None |
| DiffTactile | FEM | Taichi | Learned reflectance | Marker only |
| DOT-Sim (ours) | MPM | Warp | ResNet-based model | Full optical |
A soft gel you can press, in your browser
A 2D toy of the physics stage, not the paper's simulator: a cross-section of the gel dome, simulated live with the Material Point Method. Drag the indenter into the gel. The strips underneath show what a camera at the centre of the base sees along its rays: depth, surface normal, and how much each ray changed from the idle state. Depth and normal maps like these are the input of DOT-Sim's rendering network.
Indenter
Poisson's ratio ν
Higher ν keeps the gel's volume, so it bulges sideways when pressed.
Real-to-sim calibration, in miniature
A “recording” is a press simulated with a hidden ν. The fit starts from ν = 0.12 and follows the gradient of the Chamfer distance between its surface and the recorded one.
The toy fits one parameter and estimates its gradient by finite differences. DOT-Sim fits Young's modulus and Poisson's ratio of a 3D simulator, with gradients from automatic differentiation, on 19 real recordings.
Simulated images against the real sensor
Six indenters are pressed into a DenseTact 2.0, a soft hemispherical sensor with a random surface pattern. Drag the divider to compare the real image with each simulator's image of the same contact. Indenters #1 and #3 are the two held out in the hardest setting: the renderer never saw a real image of them.
Contact
Simulated by
The panels are taken from the paper's figures, at their resolution. The number is computed between the two panels shown; the benchmark numbers are below.
Three settings
Easy: indenters #1 and #3, frames split 80/20 at random. Medium: all six indenters, split 80/20. Hard: train on #2, #4, #5 and #6, test on #1 and #3, which are never seen. DOT-Sim has the best score in every setting and on every metric. From Easy to Hard it loses 1.6 dB of PSNR, where DiffTactile loses 7.4 dB; Tacto, which has no learned rendering, hardly changes.

Does the residual matter?
Regressing the whole image from depth and normals, without the idle frame, gives blurrier images that lose the fine surface pattern. In the Hard setting it costs 1.6 dB of PSNR.

Does the simulated gel deform like the real one?
The simulated surface is compared with the reference surface of the sensor for all indenters (2,048 points sampled on each). DOT-Sim is best on all four metrics. Chamfer distance and Earth Mover's Distance average over the whole dome, most of which barely moves, so they differ little between methods. The two metrics that focus on the deformed region separate them more: the Chamfer distance over the worst 1% of points, and the F-score at 1 mm (69.9 vs 64.7).
The reference deformations come from finite-element analysis of the gel, as in the calibration. DiffTactile is missing here: with its released code the authors could not obtain a comparable sensor deformation (see the paper).
Trained in simulation, used on the real sensor
If the simulated images are close enough to real ones, a model trained only on them should work on the real sensor as it is. Three tests, none of which uses real images for training the task.
Which indenter is it?
A classifier learns to tell indenter #1 from #3 on simulated images and is tested on real ones. In the harder case, the renderer itself was trained without any real image of these two indenters; only their meshes are known. DOT-Sim reaches 81.2% there, and 90.5% when the renderer has seen them.
Is there a lump under the skin?
The sensor presses on foam “skin” lying over a tumor phantom, a bump (present) or a dent (absent). The classifier is trained on DOT-Sim images only and tested on real images for three skins of different softness and thickness. It is right in 80.6% to 96.5% of the cases; with the other simulators' images, the classifier stays near chance (at most 52.8%).

Following a trajectory by touch
A ResNet-18 policy is trained by behaviour cloning on simulated demonstrations. From each tactile image it outputs a 6-DoF velocity, with no force measurement. On a real xArm 7, running at 25 Hz, it tracks the demonstrated trajectory with an average error of 0.896 ± 0.031 mm over 10 trials.
Reinforcement learning in the simulator
In simulation, a PPO agent that sees only the simulated tactile image learns to turn a box by 10° while keeping contact, choosing among nine planar motions. Training converges within 15 minutes.

What it costs and where it struggles
With the settings used for the results, the simulation runs at 3.6 frames per second on an RTX A6000, too slow for real-time control. Most of that is the number of MPM substeps: with a fifth of them it reaches 17.1 FPS and loses 1.2 dB of PSNR. A coarser grid barely changes the speed.
It also generalizes poorly to shapes far from the training indenters, in particular sharp edges and fine surface detail, where simulated and real contacts no longer line up. Denser particles, a learned correction of the local deformation and a larger, more varied set of indenters are the paper's suggested remedies.
All experiments use one sensor, a DenseTact 2.0, chosen because its large deformations and patterned surface are hard to simulate.
| Voxel (mm) | Softness | Substeps | FPS | PSNR (dB) |
|---|---|---|---|---|
| 1.2 | 15 | 100 | 3.6 | 31.39 |
| 1.2 | 15 | 20 | 17.1 | 30.17 |
| 2.4 | 30 | 100 | 3.8 | 30.98 |
| 2.4 | 30 | 20 | 17.2 | 29.79 |
Read the full abstract
Simulating optical tactile sensors presents significant challenges due to their high deformability and intricate optical properties. To address these issues and enable a physically accurate simulation, we propose DOT-Sim: Differentiable Optical Tactile Simulation. Unlike prior simulators that rely on simplified models of deformable sensors, DOT-Sim accurately captures the physical behavior of soft sensors by modeling them as elastic materials using the Material Point Method (MPM). DOT-Sim enables rapid calibration of optical tactile sensor simulation using a small number of demonstrations within minutes, which is substantially faster than existing methods. Compared to current baselines, our approach supports much larger and non-linear deformations. To handle the optical aspect, we propose a novel approach to simulating optical responses by learning a residual image relative to the real-world idle state. We validate the physical and visual realism of our method through a series of zero-shot sim-to-real tasks. Our experiments show that DOT-Sim (1) accurately replicates the physical dynamics of a DenseTact optical tactile sensor in reality, (2) generates realistic optical outputs in contact-rich scenarios, and (3) enables direct deployment of simulation-trained classifiers in the real world, achieving 85% classification accuracy on challenging objects and 90% accuracy in embedded tumor-type detection, and (4) allows precise trajectory following with policy trained from demonstrations in simulation with an average error of less than 0.9 mm.
BibTeX
@inproceedings{you2026dotsim,
title={DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-to-Sim Physical Calibration},
author={You, Yang and Do, Won Kyung and Swann, Aiden and Antonova, Rika and Kennedy, Monroe and Guibas, Leonidas},
booktitle={IEEE International Conference on Robotics and Automation (ICRA)},
year={2026},
note={arXiv:2604.27367}
}