Skip to main content
Bibha home

Research digest,

MilleMiglia: realistic middle-mile logistics benchmarks without private data

Middle-mile freight moves between distribution centres on fixed timetables, yet researchers have had very little public data to test solvers against. MilleMiglia, an open-source generator from researchers at Google and academic partners, creates synthetic middle-mile logistics benchmark networks that keep the hard constraints and leave out the commercially sensitive detail.

By Vijay Yadav, Associate Technical Architect, Bibha AI Labs

PaperLotfi, A., Petris, M., Cuvelier, T., De Backer, B., & Archetti, C. (2024). A novel instance generator for simulating middle-mile logistics networks [Conference presentation]. 33rd European Conference on Operational Research (EURO 2024), Copenhagen, Denmark. HAL: hal-04755189. (external site)

Aerial view at sunrise of a freight distribution centre with trucks at its loading docks and a motorway leading to a second logistics hub.
AI-generated illustration: trucks leaving a regional distribution centre on the motorway towards the next hub.

Why does middle-mile logistics lack public benchmarks?

Companies that run middle-mile networks treat their network layouts and freight volumes as confidential, so researchers have had very little realistic data to work with. In the Google Research post introducing MilleMiglia (18 September 2026), Aymane Lotfi and Thibaut Cuvelier argue that this shortage has held back academic work on the stretch of the supply chain that spans the longest distances and accounts for a large share of logistics spending.

Logistics research has long concentrated on the two ends of a shipment's journey. The first mile collects goods from producers, the last mile hands them to customers, and both are usually framed as vehicle routing problems, a field with well-established public benchmark libraries. The middle mile sits between them: bulk freight moving between distribution centres at regional or continental scale, for e-commerce, retail replenishment, automotive parts and temperature-controlled medicines bound for hospitals.

Without shared instances, published methods are hard to compare and hard to reproduce. MilleMiglia is the authors' response: a generator of networks that behave like real ones statistically while containing nothing that belongs to any operator.

How is the middle mile different from vehicle routing?

In first- and last-mile work a parcel normally stays on one vehicle from pickup to drop-off, and the planning question is which vehicle serves which stops, in what order, usually within a single day. In the middle mile a parcel changes vehicles. The post likens it to a relay: freight is unloaded at intermediate centres, sorted, combined with other consignments and put on the next departure, sometimes over several days.

The post's worked example follows a parcel from a manufacturer in Groningen in the Netherlands to a customer in Versailles, outside Paris. It passes through regional centres in Utrecht, Antwerp and Paris, and waits in Antwerp for the next day's truck because the first one to Paris is already full. That wait is the crux of the problem: freight that misses a connection sits until the next scheduled departure, which can add a significant delay.

The authors single out three constraints that cannot be loosened without changing the nature of the problem, and conclude that, because of them, existing vehicle routing solvers cannot be applied to the middle mile as it stands.

  • Fixed timetables. Vehicles run to set schedules, so a plan must fit freight into departures that already exist.
  • Hub throughput. Each distribution centre can sort or cross-dock only a limited volume in any given hour.
  • Synchronisation. A shipment can leave on one vehicle only after another vehicle has brought it to the hub.

Modelling freight as flow on a space-time graph

The authors formulate middle-mile delivery as a multi-commodity flow problem defined on a space-time graph. Each node stands for one distribution centre at one time step. An arc either carries freight on a vehicle between centres or keeps it at the same centre until a later step, which represents storage and sorting.

The authors' 2024 conference presentation states the task compactly. The inputs are hubs, lines with timetabled rotations, vehicles with limited capacity and a set of shipments, each with an origin, a destination and a quantity, revealed over a planning horizon. The goal is to bring every shipment to its destination at the lowest cost. Time only moves forward, so the graph is acyclic, with waiting arcs at each hub and travelling arcs for each line rotation.

The slides describe the library as based on earlier work by Eberhard, Cuvelier, Valko and De Backer (2023), which recast middle-mile parcel routing as a goal-conditioned Markov decision process and tackled it by combining graph neural networks with model-free reinforcement learning.

How does MilleMiglia generate a realistic network?

MilleMiglia builds an instance in two procedures: first the logistics network, then the shipments that have to cross it. Every step samples from statistical distributions, so the output resembles a real network in shape and scale without copying any operator's records. According to the post, the distributions blend information that industry players publish with data that was disclosed privately.

Analysis: because each shipment is drawn from a path that exists in the generated timetable, every shipment appears to have at least one scheduled route by construction, capacity aside. That suits benchmarking, but it also means a default instance is unlikely to test how a solver copes with freight that has no viable connection at all.

  • Hub graph. The user sets the number of hubs and the graph density. The 2024 slides and the repository README describe a Barabási-Albert random graph (modified, in the slides' wording), which yields a few highly connected hubs; the post describes placing centres with gravity models or spatial clustering to reflect population and industrial density.
  • Lines and rotations. Lines are paths through the hub graph, with arcs sampled in proportion to node degree. Rotations then attach departure and arrival times to each line, drawn from a uniform distribution. The post notes that schedules link either two major centres or a major centre and its smaller neighbours, rather than arbitrary pairs.
  • Vehicles. Each line rotation receives one vehicle. Its capacity is drawn from a uniform distribution, and its cost rises in proportion to that capacity.
  • Shipments. The origin, departure time, destination and arrival time of each shipment come from sampling a path through the space-time network, and shipment sizes follow a Lomax distribution.

One file format, from toy problems to continental networks

Each generated instance is written to a single Protocol Buffers file, which keeps the data compact and readable by solvers written in many programming languages. The generator itself is an optimised C++ library, and the HAL abstract says it can produce instances with large numbers of hubs and shipments.

The authors contrast this with vehicle routing, where separate benchmark families cover capacities (CVRP), time windows (VRPTW) or pickup and delivery with time windows (PDPTW). MilleMiglia uses one format designed to hold the defining middle-mile constraints together: fixed schedules, hub throughput limits and transfer prerequisites.

The schema in the repository defines an instance as a named logistics network plus a list of shipments. The network holds hubs with map positions, lines with ordered stops, rotations with arrival and departure windows and a fixed price, vehicles with capacities and costs, a distance matrix and a time step for discretisation. A shipment's expected arrival time is marked as a soft constraint, and its revenue is optional.

The post also points to machine learning, since the generator can produce very large datasets for training models. The HAL abstract lists that use alongside benchmarking and studying how network configuration and shipment characteristics affect overall efficiency. The authors intend instances to cover a range of sizes.

  • Small. Comparable to academic toy problems and suited to testing exact algorithms.
  • Industrial. Large, continent-wide problems that call for heuristics or metaheuristics to reach good solutions.
  • In between. Medium sizes and levels of difficulty for graded testing between the two extremes.

What do the sources show, and what remains open?

The sources present a tool, not a set of results. Neither the post nor the HAL deposit reports solver benchmarks, generation times or a quantitative comparison with real networks. The HAL record holds the 22-slide presentation given at the 33rd European Conference on Operational Research in Copenhagen in 2024, together with an abstract.

The slides do place a generated hub graph beside a real one and plot the distribution of shipment sizes, which supports the realism claim visually rather than statistically. The authors list two items of future work: adding more complex features to the generator, and gathering feedback from the research community to refine it.

Analysis: three points deserve attention before treating MilleMiglia instances as representative. First, the sources describe the generator at different stages, with the 2024 slides and the README centred on a Barabási-Albert hub graph and the post describing gravity and clustering models, so the current code is the authority. Second, in the schema file we reviewed, a hub record holds only a name and a position, so anyone who needs hub throughput limits should confirm how the current release encodes them. Third, realism is built in by design and has not yet been measured against held-out operator data.

Why middle-mile logistics benchmarks matter for applied optimisation

Shared benchmarks are what make optimisation methods comparable. Vehicle routing has CVRPLIB, and the authors present MilleMiglia as a first step toward a similar standard suite for middle-mile optimisation. If researchers adopt it, claims about a new solver or a learned policy could be checked on common instances instead of private data that nobody else can inspect.

Analysis: for organisations that plan freight networks, a generator like this has practical uses beyond academic papers. Synthetic logistics data of this kind can be shared with external solver developers without revealing hub locations or volumes, scaled up to stress-test a method before it meets production data, and varied systematically to see how network density or shipment mix changes cost. The usual caveat for synthetic data applies: results on generated instances indicate how a method behaves, and they need confirming on an organisation's own network before any operational decision.

Code, licence and next steps

MilleMiglia is on GitHub under the Apache 2.0 licence, and release v0.0.1 includes a sample instance in text Protocol Buffers format. The README explains how to build the generator with CMake and run it from the command line, setting values such as the number of hubs, shipments and lines, the maximum line length and the maximum rotations per line; its example uses 100 hubs and 20 shipments.

MilleMiglia is the product of continuing joint work between Google and academic partners at the University of Brescia and ENPC Paris. The authors say they are developing a solver and API designed specifically for middle-mile operations, and they hope to launch a challenge that draws academic and industrial solver developers to the problem.

Questions and answers

What is MilleMiglia?

MilleMiglia is an open-source instance generator, written in C++, for middle-mile logistics research. It creates synthetic networks of distribution hubs, timetabled vehicle lines and rotations, capacitated vehicles and shipments, and stores each instance in a single Protocol Buffers file. Its purpose is to give researchers realistic benchmark problems without exposing any logistics operator's confidential network or demand data. Researchers at Google and academic partners introduced it, and the code is available on GitHub.

What is middle-mile logistics?

Middle-mile logistics is the bulk movement of goods between distribution centres. It sits between the first mile, where goods are collected from producers, and the last mile, where they reach customers, and it often spans regions or continents and several days. Unlike first- and last-mile delivery, a shipment may transfer between several vehicles at intermediate hubs, so plans must respect fixed timetables, hub sorting capacity and the timing of connections.

Why can't standard vehicle routing solvers handle the middle mile?

Vehicle routing solvers optimise tours: which vehicle visits which stops, and in what order. The middle mile adds transfers. Freight waits at hubs, changes vehicles and has to arrive in time to catch scheduled departures, often across several days. The MilleMiglia authors treat this as a multi-commodity flow problem defined over a space-time graph and conclude that, because of these dependencies, existing routing solvers cannot be applied to it directly.

Does MilleMiglia use real company data?

The generated instances are synthetic and, according to the post, reveal no private information. The post explains that MilleMiglia samples hub placement, demand and vehicle schedules from statistical distributions that blend publicly available industry information with data that was disclosed privately. The output is meant to resemble a real network in structure and scale without exposing any particular operator. The sources do not yet report a quantitative test of how closely instances match real networks.

Can MilleMiglia data be used to train machine learning models?

Yes, that is one of its stated goals. The post notes that the generator can create very large datasets for training machine learning algorithms, and the HAL abstract lists model training alongside benchmarking optimisation methods. Earlier work by some of the same researchers treated middle-mile routing as a goal-conditioned reinforcement learning problem, an approach that depends on many varied training instances.

References

  1. Lotfi, A., Petris, M., Cuvelier, T., De Backer, B., & Archetti, C. (2024). A novel instance generator for simulating middle-mile logistics networks [Conference presentation]. 33rd European Conference on Operational Research (EURO 2024), Copenhagen, Denmark. HAL: hal-04755189. https://hal.science/hal-04755189v1 (external site)
  2. Google OR-Tools. (2026). MilleMiglia: Instance generator for middle mile logistics optimization (Version 0.0.1) [Computer software]. GitHub. https://github.com/or-tools/millemiglia (external site)
  3. Eberhard, O., Cuvelier, T., Valko, M., & De Backer, B. (2023). Middle-mile logistics through the lens of goal-conditioned reinforcement learning. NeurIPS 2023 Workshop on Goal-Conditioned Reinforcement Learning. arXiv:2605.02461. https://arxiv.org/abs/2605.02461 (external site)
  4. Albert, R., & Barabási, A.-L. (2000). Topology of evolving networks: Local events and universality. Physical Review Letters, 85(24), 5234-5237. https://doi.org/10.1103/PhysRevLett.85.5234 (external site)

Original article

Lotfi, A., & Cuvelier, T. (2026, 18 September). MilleMiglia: A realistic instance generator for middle-mile logistics. Google Research Blog. https://research.google/blog/millemiglia-a-realistic-instance-generator-for-middle-mile-logistics/ (external site)

This is Bibha's independent summary of published research. Bibha is not affiliated with the authors or Google.

All news and research