Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
4.00T1 sourcearXiv cs.RO
Source record
Published by arXiv cs.RO (T1 source). The original is at https://arxiv.org/abs/2607.26148.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryResearch paper studies agentic embodied control for vision-and-language navigation. Using three software-engineering agent harnesses with only monocular RGB input and discrete actions, the authors show zero-shot agents reach 70.7-78% success. Adding a trained waypoint tool reduces wall time. Model choice dominates over harness design, though longer-horizon tasks and latency remain weak points.
Why it mattersControlled ablations isolating model versus harness effects in agentic robot control. Useful baseline for practitioners deciding where to invest effort when wiring LLM agents to physical or simulated embodiments.
Cited by
No citations on record.
