Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline
3.60T1 sourcearXiv cs.RO
Source record
Published by arXiv cs.RO (T1 source). The original is at https://arxiv.org/abs/2607.22997.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryThis paper presents an end-to-end AMD ROCm-accelerated pipeline for training and deploying vision-language-action manipulation policies without CUDA. It demonstrates four tasks: Sim-to-Real on a Franka arm using SmolVLA, language-grounded object selection, Real2Sim synthetic data via 3D Gaussian Splatting with Genesis, and large-scale RL for quadruped and humanoid locomotion. All run on RDNA4/RDNA3.5 hardware and are reproducible on the free Radeon Cloud Platform.
Why it mattersDemonstrates a reproducible AMD-only alternative to CUDA for embodied AI workloads, with free cloud access. Useful for anyone evaluating hardware independence for VLA training and deployment.
Cited by
No citations on record.
