Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Upload a scene image and ask the model about robot task planning, affordance grounding,
spatial reasoning, or object pointing. The model reasons about the visual observation
and emits its final answer inside an <answer> tag. If the answer contains 2D points
(normalized to [0, 1000]), they are visualized on the image.
128 2048
Examples