Fixes and improvements
April 20, 2026 · View on GitHub
This note tracks fixes and improvements we are rolling into the released code and models.
2026-04-16
The changes below make the released models more robust for in-the-wild
use. Please use interactvlm-3d-hcontact-damon-fix for DAMON 3D
human-contact evaluation from this point on; the numbers below are
produced by this model. An arXiv update with the new numbers and
re-releases of the other Model Zoo entries are on the way.
-
Validation inference mode.
evaluate.pynow defaults to autoregressivegeneratemode. The previousforwardpath (from LISA) used teacher forcing, where body-part tokens were included in the prompt while producing[SEG]. Ingeneratemode, the model must generate body parts before[SEG], matching test-time usage. Thanks to Ha Linh Nguyen for first reporting this issue. -
Body-part dropout during training (
hC_body_part_dropout_prob). Training still uses teacher forcing, which creates a mismatch withgenerateinference. To reduce this gap, we drop body-part tokens with probabilitypby switching from thepartstemplate to thesimpletemplate. This forces the model to predict masks from visual evidence alone in a subset of updates, improving generalization. -
Improved per-view GT contact masks (
mv2).generate_damon_human_mask.pynow supports--min_vertices 2(previously 3), and DAMON training uses the4MV-Z_Vitru_mv2view set. A contact triangle is retained if at least two vertices project inside the silhouette, instead of all three. This stabilizes boundary regions. The overall gain is modest; masks from the previous version remain usable. -
3D contact predictor.
HumanContact3DPredictornow aggregates views with a soft sigmoid and barycentric-weighted scatter, restoring gradient flow through the 2D→3D step. The previous hard-threshold version was effectively detached. -
Loss cleanups.
compute_dice_lossno longer returns early on empty-GT views, andHumanContact3DLossclamps its inputs before BCE. -
Binary contact metric threshold.
get_damon_binary_contactnow thresholds predictions at0.5before forming the per-image union.
Updated DAMON numbers
Two models are released with this update, both evaluated on the full DAMON
test split (1370 samples) with inference_type=generate and threshold 0.5:
interactvlm-3d-hcontact-damon-fix—partsanswer template ("The contacting body parts are {body_parts}, and the contact region is [SEG]."). Use this for the strongest numbers.interactvlm-3d-hcontact-damon-noParts—simpleanswer template ("Sure, [SEG].").
Binary contact (per-image)
| Model | F1 | Precision | Recall |
|---|---|---|---|
interactvlm-3d-hcontact-damon-fix | 70.32 | 67.91 | 78.91 |
interactvlm-3d-hcontact-damon-noParts | 64.46 | 67.34 | 68.31 |
Semantic contact (per-object)
| Category | # samples | with body parts (damon-fix) |
without body parts (damon-noParts) |
||||
|---|---|---|---|---|---|---|---|
| F1 | Precision | Recall | F1 | Precision | Recall | ||
| transport | 87 | 72.25 | 67.84 | 82.36 | 70.35 | 70.49 | 76.97 |
| sports | 305 | 72.44 | 71.20 | 81.47 | 66.64 | 69.16 | 71.84 |
| kitchen | 38 | 65.35 | 62.12 | 75.31 | 57.32 | 57.04 | 64.85 |
| food | 32 | 62.05 | 57.61 | 76.67 | 56.08 | 52.35 | 70.82 |
| accessory | 47 | 60.34 | 56.15 | 70.05 | 46.45 | 47.65 | 51.32 |
| furniture | 146 | 58.90 | 54.04 | 72.87 | 45.04 | 57.59 | 42.78 |
| everyday-objects | 174 | 54.98 | 53.62 | 63.42 | 52.17 | 50.45 | 64.20 |
| supporting | 541 | 68.84 | 67.43 | 76.78 | 65.23 | 67.14 | 70.17 |
| DAMON (weighted) | 1370 | 66.49 | 64.35 | 75.79 | 60.98 | 63.37 | 66.52 |
Please use these numbers for any comparison against InteractVLM on DAMON.