EMBODIED REASONING SANDBOX

Molmo2-ER Image Inspector

Upload an image, ask a spatial or pointing question, and inspect the model's response and grounded coordinates.

32 512
0 1.5
Prompt ideas

Model: allenai/Molmo2-ER. Point prompts return <points coords=...> tokens, visualized above. Molmo2-ER is a perception backbone and does not produce robot-control action tokens.