Inference for robot fleets

Robot policies, served in 22.7 ms. Under the deadline. Off the robot.

QuadShadow runs your vision-language-action models — SmolVLA, π0-class, GR00T-class, and your fine-tunes — on cloud GPUs with a control-loop SLO. No GPU on the robot. One push updates the whole fleet.

Measured action-chunk latency · 0–60 ms scale

22.7 msH100 · full quality
30 mscontrol deadline
48.3 mscommodity L4 · full quality

Measured, not promised

22.7 ms
Full-quality chunk latency, p95 22.8 ms, H100
✓ measured 2026-08-18
1.0 ms
Network + serialization, robot→cloud→robot, in-region
✓ measured, 280 samples
0.45 ms
Effective latency per action (50-action chunks)
✓ derived from p50
~30 robots
Served per commodity GPU at full quality
✓ measured throughput

How it works

  1. Point your policy at us

    Bring a LeRobot-compatible checkpoint — ours or your fine-tune. One config file: quadshadow connect --config robot.yaml

  2. Stream observations, receive action chunks

    Camera frames and state go up; 50-action chunks come back in ~23–48 ms. Session-affine, region-pinned, jitter under 0.2 ms after warmup.

  3. Fail safe, update in one push

    Motors halt on disconnect by protocol. New checkpoint versions roll out to the whole fleet — or roll back — in a single command.

Models

ModelParamsChunk latency · H100Chunk latency · L4Status
SmolVLA (base + fine-tunes)450 M 22.7 ms48.3 ms Serving
π0 / π0.5-class~3.3 B Benchmarking
GR00T N1.73 B Integrating
Your fine-tuned checkpoints≤7 B per-embodiment registry, one endpoint per fleet Design partners

Design-partner program — 3 fleets, 90 days, free.

We integrate your embodiment, serve up to 25 robots at a p95 < 75 ms in-region SLO, and do the plumbing with you. In return: your feedback and a case study.

Read the program