One hazard. One Ultra call. The whole fleet knows.
When a robot meets something new, like a spill in aisle 4 or a coolant leak at the end of row C2, Nemotron Ultra works out what to do. NemoFleet stores that decision against the map cell. Every robot that reaches the cell afterwards reuses it through Nemotron Nano instead of paying for the same reasoning again. A person in the lane is never reused. That always goes back to Ultra.
What the fleet has done so far
Real counts from real Token Factory calls, across both sites. Open the console and click Run scenario to add to them.
Every robot meets the same hazard
In a warehouse or a data-center hall, one spill is met by robot after robot. Today a fleet has two bad options. It can send every sighting to a frontier reasoning model, which is slow and costs more as the fleet grows. Or it can let a small model guess, which is fast but wrong on things it has never seen. In practice, operators halt the whole zone and wait.
See, decide, remember
Each event goes through three steps. Ultra handles what is new, unsafe or uncertain. Fleet memory makes sure it only handles each hazard once.
Perception
The robot's camera frame becomes structured JSON: hazards, whether a person is present, whether the path is clear, and a recommended action.
- Model
- Gemma 3 27B today · Cosmos3 Super Reasoner ready (one-command switch)
- Runs on
- Nebius Token Factory
Nano, then Ultra if needed
Nano returns typed fleet actions with a confidence score. New hazards, people in a lane and low confidence go to Ultra.
- Models
- Nemotron 3 Nano 30B · Nemotron 3 Ultra 550B
- Runs on
- Token Factory Responses API, strict JSON schema
Fleet memory
Ultra's decision is stored against the map cell, with an expiry. The next robot at that cell gets it in Nano's prompt and no Ultra call is made. If the scene has changed, Nano lowers its confidence and Ultra decides again.
- Stored in
- D1 precedents · KV hazard tags · R2 raw calls
- Never reused
- safety events, people in a lane
Every model output is checked against a schema before it runs, and a deterministic A* + CBS planner owns every trajectory. Models choose actions, never paths.
Safety is never delegated
A person in a lane always goes to Ultra. Fleet memory and peers never settle it.
Outputs are constrained by a strict JSON schema and validated again. An invalid answer is retried once, then the robot stops.
A precedent is context, not a command. If what the robot sees differs, the decision goes back to Ultra.
A* + CBS computes every trajectory, collision-free by construction.
Knowledge that still moves in a dead zone
Wi-Fi fades between tall metal racks and at the end of cold aisles. When two robots pass, each hands the other the precedents it holds, with a hop count and an expiry. A robot that has lost its link still knows about the spill.
- Today: handshakes copy precedents in fleet memory with hop counts and expiring tags
- Next: offline mode, where a robot knows only what its peers handed it, measured against the same run with no handshakes
What runs where
Every model call in the console is real and logged, with latency and tokens per call.
| Role | Model | Runs on | Status |
|---|---|---|---|
| Deliberation on new or unsafe events | Nemotron 3 Ultra 550B | Token Factory · Responses API · strict schema | Live |
| Reflex decisions, reuse of precedents | Nemotron 3 Nano 30B | Token Factory · Responses API · strict schema | Live |
| Camera perception | Gemma 3 27B · Cosmos3 Super Reasoner | Token Factory · serverless / dedicated endpoint | Gemma live · Cosmos wired, waiting on its endpoint |
| Night Shift replay: Ultra re-judges Nano | Nemotron 3 Ultra 550B | Token Factory | Live |
| Night Shift dream, gate, LoRA consolidate | Ultra · Nano + LoRA | Nebius Serverless Jobs · MLflow | Planned |
| Fleet memory and call log | none | Cloudflare D1 · KV · R2 | Live |
The fleet learns overnight
Fleet memory helps the next robot within minutes. Night Shift turns the day's decisions into a better reflex model for tomorrow.
Re-judge the day
Token FactoryUltra re-judges Nano's decisions and flags disagreements. The disagreement rate guides tomorrow's threshold τ.
Simulate variants
Serverless Jobs fan-outUltra writes variants of the hardest episodes, and jobs run them through the planner.
Keep only safe lessons
Serverless JobsA lesson must fix its variants without making normal operations worse.
Train and promote
H200 job · MLflowLoRA fine-tune of Nano, promoted only if it beats today's model.
Pairing: learn from the robot just ahead
A scheduler adds pass-by waypoints so a new robot meets peers in order of competence. Each exchange stays small enough to take in during one pass.
Named after countercurrent exchange in a duck's leg, where opposite flows keep a small gradient along the whole length.
Run it yourself
Apache 2.0. The repository has setup steps, the architecture, the QA suite and our platform feedback.