Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Upload a scene image and ask the model about robot task planning, affordance grounding, spatial reasoning, or object pointing. The model reasons about the visual observation and emits its final answer inside an <answer> tag. If the answer contains 2D points (normalized to [0, 1000]), they are visualized on the image.

128 2048
Examples