OhhO Serve
Robot AI inference, as an API. Deploy a Vision-Language-Action model behind a REST endpoint in one command — pluggable model backends, batching, and Prometheus metrics built in.
Robot AI inference, as an API.
Modern robots run on large Vision-Language-Action (VLA) models, but getting one into production means GPU servers, model loading, batching, versioning and monitoring. OhhO Serve turns all of that into a single endpoint.
Point Serve at a model — OpenVLA, SmolVLA, ACT, a diffusion policy, or your own fine-tune — and it exposes a clean REST API your robot calls with an image and an instruction, and gets back an action. Swap models without touching robot code.
Serve is built for the real world: health checks, hot model loading, 4-bit quantization for tighter VRAM budgets, and first-class Prometheus metrics so you can see latency and throughput at a glance.
- Model-agnostic backend. OpenVLA, SmolVLA, ACT, diffusion, or a custom class — selected by config, not code.
- Clean predict API. POST an image + instruction, get back an action vector. Your robot doesn't need to know which model is behind it.
- Hot loading & health. /health, /load_model and /predict endpoints let you swap or reload models with zero downtime.
- Fits your hardware. Optional 4-bit quantization runs large VLA models on modest GPUs. Bring your own — we help you size it, or recommend hardware that fits your budget.
- Observability first. Latency, throughput and GPU metrics export to Prometheus and the OhhO Fleet dashboards.
Related: OhhO Train, OhhO Data, OhhO View