Vision-language model for embodied reasoning: spatial understanding, task planning, and physical-world agentic robotics
Vision-language model for embodied reasoning: spatial understanding, task planning, and physical-world agentic robotics. Supports a context window of 131,072 tokens. Supports extended reasoning. Supports tool calling. Input priced at $1 per million tokens. Weights are not publicly released; access is via the provider API.
api
paid
No benchmark results have been added yet.