Physical AI · warehouse and data center robot fleets · Nebius Token Factory + NVIDIA Nemotron

Robot fleets that stop paying for the same decision twice.

Nemotron Nano on every robot. Peer handshakes in the aisle. Nemotron Ultra on Nebius Token Factory only for what's genuinely new. Every night, the day's incidents become tomorrow's weights.

FIG 01 · Decision tiers

Three speeds of fleet cognition

Known problems are settled on the robot in under a second. Peers spread what one robot learned. Nebius compute goes where only the core can help.

SYSTEM 1 · REFLEX

Nemotron Nano

Returns an action and a confidence score for every event.

Runs on
vLLM Serverless Endpoint
Latency
under 1 s
Settles
routine events
SYSTEM 1.5 · PEER EXCHANGE

Handshakes + pairing

Robots swap bounded payloads when they pass. No model call.

Runs on
the robots
Latency
under 50 ms
Settles
problems another robot already solved
SYSTEM 2 · DELIBERATION

Nemotron Ultra + Cosmos

Novel or safety-critical events, and first-sighting perception.

Runs on
Nebius Token Factory
Latency
5–30 s
Settles
new situations, people in lanes

Every model output is validated against a typed schema, and a deterministic A* + CBS planner owns every trajectory.

FIG 02 · Antennation handshakeRobots swap a small, bounded payload as they pass. Confidence decays with every hop, tags expire, and safety events always escalate.
System 1.5 · runtime

Peer exchange: knowledge moves at walking speed

When two robots pass, each hands the other what it knows about the aisle behind it. Only the first robot to see a hazard pays for a perception call.

  • Payload: hazard precedent, source, hop count, time to live
  • Confidence decays per hop; tags expire on their own
  • Conflicting precedents trigger an escalation

Named after how ants exchange information on contact.

FIG 03 · Pairing schedulerEach robot meets a peer slightly ahead of it, so every pass carries a small, usable difference. Random pairing plateaus.
System 1.5 · learning

Pairing scheduler: learn from the robot just ahead

The scheduler adds pass-by waypoints so a new robot meets peers in order of competence, keeping each exchange small enough to absorb within one pass.

  • Competence score per robot: precedent coverage and recent escalation rate
  • Evaluated against random pairing and no peers on Serverless Jobs

Named after countercurrent exchange: in a duck's leg, opposite flows keep a small gradient along the whole length.

FIG 04 · Night Shift

Night Shift: replay, dream, gate, consolidate

One Serverless Jobs pipeline per night. The edge experiences an incident once; Nebius lives through hundreds of versions of it, then fine-tunes the fleet's reflex model before the morning shift.

01 REPLAY

Re-judge the day

Token Factory batch · Data Lab

Ultra re-judges every Nano decision and flags confident mistakes. The disagreement rate sets tomorrow's threshold τ.

02 DREAM

Simulate variants

Serverless Jobs fan-out

Ultra writes counterfactual variants of the hardest episodes; dozens of jobs run them through the simulator and planner.

03 GATE

Keep only safe lessons

Serverless Jobs

A lesson must fix its variants without regressing normal operations. Most are rejected.

04 CONSOLIDATE

Train and promote

H200 job · MLflow · vLLM endpoint

LoRA fine-tune of Nano, tracked in MLflow, promoted only if it beats today's model.

FIG 05 · Workloads

Where the compute goes

We don't cut compute. We move it from repeating answers to deliberating, stress-testing and learning.

WorkloadRuns onModelGrows with
Deliberation on novel eventsToken FactoryNemotron Ultranew sites, new hazards, fleet size
First-sighting perceptionToken Factory dedicated endpointCosmos3 Super Reasonernew hazards
Reflex inferenceServerless Endpoint (vLLM)Nemotron Nano + nightly adapterrobots × events
Peer exchangethe robotsnonepasses per hour
Night ShiftToken Factory batch + Serverless Jobs + MLflowUltra, Nano (LoRA)incidents × variants × fleets
FIG 06 · Invariants

Safety is never delegated

People escalate

A person in a lane always goes to Ultra, whatever the peers say.

Schema first

Every model output is validated; malformed actions are retried, never executed.

Planner owns motion

A* + CBS computes every trajectory, collision-free by construction.

Gated learning

No adapter ships without passing the red-team and normal-operations suites.

Source

Run it yourself

Architecture, implementation plan, demo script and platform feedback are in the repository.