Robot fleets that stop paying for the same decision twice.
Nemotron Nano on every robot. Peer handshakes in the aisle. Nemotron Ultra on Nebius Token Factory only for what's genuinely new. Every night, the day's incidents become tomorrow's weights.
Three speeds of fleet cognition
Known problems are settled on the robot in under a second. Peers spread what one robot learned. Nebius compute goes where only the core can help.
Nemotron Nano
Returns an action and a confidence score for every event.
- Runs on
- vLLM Serverless Endpoint
- Latency
- under 1 s
- Settles
- routine events
Handshakes + pairing
Robots swap bounded payloads when they pass. No model call.
- Runs on
- the robots
- Latency
- under 50 ms
- Settles
- problems another robot already solved
Nemotron Ultra + Cosmos
Novel or safety-critical events, and first-sighting perception.
- Runs on
- Nebius Token Factory
- Latency
- 5–30 s
- Settles
- new situations, people in lanes
Every model output is validated against a typed schema, and a deterministic A* + CBS planner owns every trajectory.
Peer exchange: knowledge moves at walking speed
When two robots pass, each hands the other what it knows about the aisle behind it. Only the first robot to see a hazard pays for a perception call.
- Payload: hazard precedent, source, hop count, time to live
- Confidence decays per hop; tags expire on their own
- Conflicting precedents trigger an escalation
Named after how ants exchange information on contact.
Pairing scheduler: learn from the robot just ahead
The scheduler adds pass-by waypoints so a new robot meets peers in order of competence, keeping each exchange small enough to absorb within one pass.
- Competence score per robot: precedent coverage and recent escalation rate
- Evaluated against random pairing and no peers on Serverless Jobs
Named after countercurrent exchange: in a duck's leg, opposite flows keep a small gradient along the whole length.
Night Shift: replay, dream, gate, consolidate
One Serverless Jobs pipeline per night. The edge experiences an incident once; Nebius lives through hundreds of versions of it, then fine-tunes the fleet's reflex model before the morning shift.
Re-judge the day
Token Factory batch · Data LabUltra re-judges every Nano decision and flags confident mistakes. The disagreement rate sets tomorrow's threshold τ.
Simulate variants
Serverless Jobs fan-outUltra writes counterfactual variants of the hardest episodes; dozens of jobs run them through the simulator and planner.
Keep only safe lessons
Serverless JobsA lesson must fix its variants without regressing normal operations. Most are rejected.
Train and promote
H200 job · MLflow · vLLM endpointLoRA fine-tune of Nano, tracked in MLflow, promoted only if it beats today's model.
Where the compute goes
We don't cut compute. We move it from repeating answers to deliberating, stress-testing and learning.
| Workload | Runs on | Model | Grows with |
|---|---|---|---|
| Deliberation on novel events | Token Factory | Nemotron Ultra | new sites, new hazards, fleet size |
| First-sighting perception | Token Factory dedicated endpoint | Cosmos3 Super Reasoner | new hazards |
| Reflex inference | Serverless Endpoint (vLLM) | Nemotron Nano + nightly adapter | robots × events |
| Peer exchange | the robots | none | passes per hour |
| Night Shift | Token Factory batch + Serverless Jobs + MLflow | Ultra, Nano (LoRA) | incidents × variants × fleets |
Safety is never delegated
A person in a lane always goes to Ultra, whatever the peers say.
Every model output is validated; malformed actions are retried, never executed.
A* + CBS computes every trajectory, collision-free by construction.
No adapter ships without passing the red-team and normal-operations suites.
Run it yourself
Architecture, implementation plan, demo script and platform feedback are in the repository.