Siltframe · Robustness audit · SAMPLE
Sample report. Prepared on public RELLIS-3D data to show the format of a customer audit.
mIoU lost under night / low light, the worst condition (clean 38.1 → 13.8).
Drivable-surface IoU under night / low light. The model stops finding the ground it should drive on.
mIoU recovered on held-out degradations after a hard-case fine-tune, vs. an equal-length clean fine-tune.
The model is solid in clear conditions (38.1 mIoU, 87 drivable-surface IoU) but has one critical weakness and two serious ones. Night is critical: it removes most of the signal the model relies on to separate terrain types, and drivable-surface detection collapses. Dust and lens contamination are serious: distant structure (trees, vehicles, obstacles) disappears first. Fog and motion blur cost less, but thin classes (fences, barriers) suffer in all of them.
| Condition | mIoU | vs clean | Drivable IoU | Sev 1 / 2 / 3 |
|---|---|---|---|---|
| Clear (reference) | 38.1 | — | 87 | — |
| Night / low light held-out | 13.8 | −64% | 36 | 20 / 14 / 8 |
| Dust Siltframe | 21.4 | −44% | 74 | 25 / 22 / 17 |
| Mud & water on lens held-out | 21.6 | −43% | 76 | 29 / 21 / 14 |
| Motion blur held-out | 26.3 | −31% | 79 | 31 / 26 / 22 |
| Rain on lens Siltframe | 28.1 | −26% | 74 | 36 / 31 / 18 |
| Fog held-out | 29.3 | −23% | 81 | 32 / 30 / 26 |
“held-out” = an independent corruption generator (ImageNet-C or a literature low-light model) that is never used to create training data. Those numbers are the conservative ones.
IoU per class, averaged over severities. Darker cells = larger loss relative to clear weather. Classes shown are those relevant to driving with a clear-weather IoU of at least 15.
| Class | Clear | Night | Dust | Lens mud | Fog | Lens rain | Blur |
|---|---|---|---|---|---|---|---|
| grass | 86 | 29 | 76 | 78 | 80 | 74 | 78 |
| concrete | 79 | 12 | 26 | 27 | 65 | 59 | 64 |
| puddle | 70 | 35 | 50 | 37 | 58 | 59 | 39 |
| mud | 36 | 4 | 16 | 13 | 19 | 24 | 17 |
| tree | 76 | 31 | 32 | 46 | 57 | 59 | 63 |
| bush | 66 | 32 | 53 | 42 | 58 | 55 | 35 |
| person | 84 | 48 | 73 | 70 | 69 | 57 | 61 |
| barrier | 25 | 2 | 7 | 10 | 19 | 7 | 19 |
| rubble | 28 | 0 | 0 | 1 | 5 | 16 | 5 |
One real test frame, degraded four ways. Colours follow the RELLIS-3D ontology; ground truth for reference:













Fine-tuned from the baseline with on-the-fly Siltframe degradations (50 % of samples), compared with an identical fine-tune on clean data. Mean mIoU over severities.
| Condition | Control | + hard cases | Δ |
|---|---|---|---|
| Night / low light | 12.9 | 18.3 | +5.4 |
| Dust | 21.1 | 28.4 | +7.3 |
| Mud & water on lens | 22.1 | 26.0 | +3.9 |
| Fog | 29.2 | 32.7 | +3.5 |
| Rain on lens | 28.3 | 32.2 | +3.9 |
| Motion blur | 26.4 | 26.6 | +0.3 |
At night the largest per-class gain is puddle (+17.5 IoU); mud does not improve (-1.9). Clean-weather mIoU is unchanged within one point.
Baseline trained on 1,200 clear RELLIS-3D frames at 640×400. Degraded test sets: 300 real test frames × 6 conditions × 3 severities, rendered with fixed seeds. Metrics: mIoU over classes present in the test split; drivable surface = dirt, grass, asphalt, concrete. Fine-tune comparison: control arm with identical steps on clean data. Limits: degraded inputs are synthetic (RELLIS-3D has no real night or dust frames); a customer audit adds your own real adverse-condition frames as the final check.