Concept render — not footage of OmniBot

Prototype

OhhO Serve

Robot AI inference, as an API. Deploy a Vision-Language-Action model behind a REST endpoint in one command — pluggable model backends, batching, and Prometheus metrics built in.

Robot AI inference, as an API.

Modern robots run on large Vision-Language-Action (VLA) models, but getting one into production means GPU servers, model loading, batching, versioning and monitoring. OhhO Serve turns all of that into a single endpoint.

Point Serve at a model — OpenVLA, SmolVLA, ACT, a diffusion policy, or your own fine-tune — and it exposes a clean REST API your robot calls with an image and an instruction, and gets back an action. Swap models without touching robot code.

Serve is built for the real world: health checks, hot model loading, 4-bit quantization for tighter VRAM budgets, and first-class Prometheus metrics so you can see latency and throughput at a glance.

  • Model-agnostic backend. OpenVLA, SmolVLA, ACT, diffusion, or a custom class — selected by config, not code.
  • Clean predict API. POST an image + instruction, get back an action vector. Your robot doesn't need to know which model is behind it.
  • Hot loading & health. /health, /load_model and /predict endpoints let you swap or reload models with zero downtime.
  • Fits your hardware. Optional 4-bit quantization runs large VLA models on modest GPUs. Bring your own — we help you size it, or recommend hardware that fits your budget.
  • Observability first. Latency, throughput and GPU metrics export to Prometheus and the OhhO Fleet dashboards.

Related: OhhO Train, OhhO Data, OhhO View