Dataleo Insight · 2026-09-01· Research
OR-Transformer tests real-time AI replenishment at 1,024-item scale
A research team from MIT, Purdue, Caltech and the University of Virginia proposes OR-Transformer, a deep reinforcement learning framework for large-scale stochastic joint replenishment. The model targets a classic inventory problem where many SKUs share ordering costs, lead times and correlated demand patterns.
The paper matters because it attacks a planning-system bottleneck directly: replenishment decisions where the SKU/action space makes solve-at-decision-time optimization too slow. In experiments up to 1,024 items, OR-Transformer remains stable where several learning baselines degrade, and beats rolling-horizon MILP baselines on decision speed and reported inventory cost.
For Supply Chain leaders, the implication is architectural. Some inventory policies may need to move from repeated online optimization toward trained decision policies when the decision frequency and SKU coupling exceed what classical solvers can handle in operational time. That does not remove the need for optimization discipline. It shifts the control question toward where the policy is trained, how it is stress-tested, and when planners should reject or override it.
Mots-clés
