Our family of vision-language-action models is built for real-world deployment. It delivers better-than-Gemini Flash perception with up to 10x lower cost, while running on edge hardware (e.g., Jetson) or scaling through OpenAI-compatible cloud APIs.