Engineering

Stacking

Sep 30, 2026

This post features an unusual kind of rider in Cabify, a mostly ride-hailing company. In this case, we’ll talk about parcels (instead of people) riding vehicles. So let’s introduce Parcelo (Cabify Logistics’s pet) and see what makes it special:

IT Parcelo

In general, the main difference lies in how social parcels are: you can put several of them in the same car and they won’t ever complain about not knowing each other (don’t try to get strangers to ride the same car). Parcels are also typically more patient, and can wait depending on their so-called “delivery window”: e.g., hamburgers need to be delivered as soon as possible, while your Amazon order can wait a bit more.

This diversity of requirements takes us to distinguishing the main ways of delivering parcels: the different operations.

Logistics operations

One kind of operation we do in Cabify is cross-dock: a driver visits multiple pickup points to collect parcels in what’s known as a “first mile”. These parcels go to a warehouse where they get moved from the first-mile vehicle to a new vehicle that will do the “last mile”: a trip to their final destinations. The key trait of cross-docking is that parcels move directly from one vehicle to another, without intermediate storage. Visually, picking up and delivering through some first miles and last miles looks like this:

Cross-dock operation

In another kind of operation, parcels get delivered without a central warehouse, and it’s the closest to the ride-hailing problem, where you request a Cabify that takes you to your destination. It is the point-to-point delivery that takes place with food: once the pizza is ready, a driver gets it from the restaurant and takes it to the buyer as fast as possible to keep the pizza warm. It’s a similar case to what some online stores offer: you order a phone, a driver picks it up and you get it at home within 2 hours (same operation as with the pizza, just lighter time constraints).

Point-to-point operation

Note that there’s not a one-to-one mapping between operation type and delivery window. In general, the most urgent deliveries are served through point-to-point (like ride-hailing) instead of cross-dock operations, which by design add delay: parcels need to get sorted at the cross-dock, so it is typically used for same-day or next-day shipments.

So what’s stacking? We mentioned several parcels are collected in a first mile and delivered in a last mile. That involves complex routing algorithms to identify which parcels should go into which vehicles to save distance and driver time, both for the first mile and the last mile. Similarly, we could strive to optimize point-to-point operations, and try to deliver two or more pizzas if there’s a good route that can preserve the tight delivery window of food. This is the objective of “stacking” the pizzas (which literally means putting one on top of another!): applying those same routing techniques to make more than one pickup and more than one delivery in point-to-point operations, where no central warehouse is involved. Visually, the point to point operation with stacking would look as follows:

Point-to-point with stacking

Algorithms

We just casually mentioned the word “algorithms”, which sounds like we’re getting to the fun part. Since hundreds of companies are dedicated to logistics, routing is a well-studied problem with existing solutions.

The problem’s complexity is equally famous, and unfortunately brute-forcing is simply not an option since the number of combinations can explode with not too many parcels. Approaches where we check all possible solutions and filter by constraints will just not work (and no, there’s no foreseeable future where computational power will change this). Instead, algorithms rely on exploiting the limitations of each specific problem, or use heuristics to reach suboptimal solutions.

What are these “limitations of each specific problem”? There are actually a ton of cases. In general, this is called the Vehicle Routing Problem (VRP): delivering multiple parcels in multiple vehicles to a set of destinations from a single pickup point, but there are many variants:

  • Capacitated VRP: vehicles have a limited capacity, so you cannot carry an unlimited number of parcels per vehicle.
  • Distance-constrained VRP: there’s a maximum distance a vehicle can make due to fuel/charge limitations.
  • VRP with Time Windows: each destination needs to be visited with some time constraints (e.g., when requesting supermarket groceries to deliver between 2pm and 4pm).
  • Heterogeneous Fleet VRP: the available fleet has heterogeneous features, like different capacities, speed, or mileage.
  • Pickup and Delivery VRP: there’s not just a single pickup point, so vehicles can do two pickups in one stop, do one delivery, do two more pickups at another stop, then do three deliveries at the final destination.

The list can grow in boring ways, but the key lies in understanding how certain algorithms solve subsets of these problems. One typical one is the Clarke-Wright algorithm (also called the “savings” algorithm), which focuses on the most traditional VRP with a single pickup point and can reduce the complexity compared to other algorithms. There are variations for time window restrictions, or adapting it to first miles (by reversing routes compared to the original problem where there’s a single pickup point and multiple destinations). But it just won’t make it if more than one warehouse is required (i.e., no support for multipickup or multidelivery) or additional constraints are needed. Other algorithms include insertion (which iteratively builds a list of deliveries by inserting parcels), genetic algorithms (where solutions get iterated via evolution and natural selection), branch-and-bound (close to brute force, just with clever pruning), or linear programming (where constraints are represented by linear relationships).

In our case, Cabify uses an insertion-based algorithm that tackles the Pickup and Delivery VRP and is applied even when we limit ourselves to a single central warehouse for pickups: this reduces maintenance (one algorithm to rule them all, not one per case), while also facilitating the addition of new constraints. Order of complexity is not such a problem for real-time stacking, since pickup origins typically keep a low stock count (< 15 parcels), and the time constraints plus the limited space of vehicles make it unlikely to require processing thousands of parcels on each run. This contrasts with first-mile or last-mile routing, where you can easily have thousands of parcels to route, but real-time performance is not a requirement. We keep benchmarking other approaches, including open-source frameworks like Verso’s VROOM and Google’s OR-Tools.

About time and space

What exactly are these time and space constraints of parcels and vehicles? Quite obviously, a vehicle cannot take an unlimited number of parcels, so there’s a capacity constraint. Usually parcels are defined by weight and volume (length, width, and depth). Similarly, a vehicle’s boot (trunk) would follow the same model, determining which parcels could fit in. In practice, we don’t usually have such details: no one is measuring your burger menu, and stores doing 2-hour-delivery are also busy enough fulfilling orders and don’t want to measure parcels, so we typically have to deal with not having volumetric data and assume an approximate number of parcels vehicles can take.

In contrast, time constraints can be very strict. Not meeting a delivery window might mean not making money, regardless of how many minutes late. But there are many different times involved that we need to take into account:

  • Time to assign: the time required to find a driver, which depends on the drivers available at any given moment.
  • Time to pickup: the time it takes a driver to go to the first stop (which is always a pickup stop). This could be seen as the driving time to stop zero, but it’s special because we don’t know exactly where the assigned driver will be located. It’s determined by the search radius used when looking for drivers.
  • Pickup window: the valid time window when a parcel can be picked up. The pickup location is typically a warehouse that could have a closing time, which determines the pickup window. Similarly, parcels take time to fulfill (and food needs to get cooked), so a driver shouldn’t arrive too soon either.
  • Driving time: the travel time taken between stops, which should account for the expected traffic (peak hours can multiply by 2 or 3 the usual driving times in most cities).
  • Delivery window: the time constraints at the parcel’s destination. As mentioned, this is the most common constraint. For groceries, typically customers expect a window with start time and end time (e.g., receive my goods between 2pm and 4pm), while other cases don’t have a lower limit (so parcels can get delivered as soon as possible, without the possibility of arriving “too soon”).
  • Maximum start time: this is the resulting calculation of all times mentioned above. If a parcel takes 5 minutes to find a driver, 5 minutes to pickup, 20 minutes of driving time, and has to be delivered before 2pm, then it means we should start looking for a driver at 1:30pm the latest. Note this concept is applicable to individual parcels but also to “grouped parcels” (so the example before might need to start earlier if there are more future parcels to deliver with their own time restrictions).

All this is heavily connected to routing. In the typical travelling salesman problem, the different stops are visited according to the shortest path possible. But this is not necessarily the case when we have constraints for each parcel, which could cause some reordering to ensure we arrive just-in-time at each stop. Consider the following example:

Routing constraints differences

You can see that although route A is shorter, we might need to use route B if parcel C has a tight delivery window and we need to reach it earlier to meet the constraint.

Another interesting case is shown in the next example. Let’s assume there are no pickup or delivery window constraints. Which route is better?

Routes without considering return-to-origin

Route A looks shorter, while route B is odd because it circles back to an intermediate stop that could have been visited earlier. So why the question? The answer is: it depends. Route A finishes at a different location than route B, and that’s the key to deciding. If drivers finish deliveries in high-demand areas, that’s good for keeping them busy. So estimated demand is a deciding factor, but not the only one. For example, one of our clients uses dedicated fleet to cover their deliveries, since their warehouse is located in a low-demand area. These drivers are hired to only deliver the client’s parcels and are expected to return to the pickup point for more deliveries. This means we need to consider the path of the driver moving back to the pickup point (or to a high-demand area). Once we account for the back-to-origin segment (the dotted line), route B is the better choice:

Routes considering return-to-origin

Ready, set, go

What do we do once we have a set of routes for all the parcels that need to be delivered? Those routes are the output of our algorithm, but that does not mean we should dispatch a driver for each one right away. They represent a snapshot of the current optimal routing, but this picture could improve once we receive additional requests. Suppose we have a route with 2 parcels that, due to their time constraints, can wait for a bit: what’s the likelihood of receiving a request for delivering more parcels that fit in this same route? We could have a driver taking those 2 parcels along with, let’s say, 3 more parcels. So it sounds like we’re trying to see the future, but since parcels can wait (if their time windows are met), we can actually see the future just by waiting for it.

This takes us to the delivery triggers of routes, which are mostly two: the time trigger and the capacity trigger. The time trigger is straightforward: if a route involves parcels and we need to start the trip now because otherwise we’ll miss the delivery window, we start looking for a driver. Going back to our earlier example, if we have two parcels in a route and they can wait, we’ll wait; otherwise we’ll dispatch now. Technically, we’re checking against the maximum start time defined above: if we’re reaching it, we should start the delivery or we’ll be late.

The other trigger is the capacity one, which is even simpler. If a route involves as many parcels as the vehicle can hold, it’s ready to go. In our two-parcel example, if they fill the whole boot (trunk), it doesn’t make sense to keep waiting for future parcels, as we could start the delivery right away.

All this is orchestrated through continuous monitoring of pending parcels in what we call rounds: every few minutes we run the routing algorithm and check the route candidates against the triggers. If any trigger fires, the route candidate becomes a real delivery, and we dispatch it to a driver.

The Real World™️

How’s this performing right now? We’re currently stacking parcels for restaurants doing food delivery and a store that does 3-hour deliveries of consumer goods. Food delivery has tight delivery windows (nobody wants cold pizzas), and we’re delivering around 1.9 parcels per delivery (meaning we’re almost doubling the standard performance of “one meal per trip”).

Regarding the consumer goods store, it has a warehouse on the outskirts of Lima, which combined with the wide delivery window, offers a big opportunity for stacking. Typical travel time to the center of the city takes more than 30 minutes, and it’s a trip common to all parcels, so unless delivery addresses are very far from each other, it usually pays off to take many parcels at once. However, driving times to the center can reach 50 minutes during peak hours, and with almost an hour needed to fulfill parcels, we only have around another hour for deliveries during those peak windows. The actual destination locations matter a lot when building the route. On average, we’re stacking around 5 parcels per delivery, with more than 95% on-time-delivery success rate (the longer the route, the higher the probability that a delivery issue causes late deliveries for the last parcels).

Would you like to know more? Real-time operations such as stacking can be stressful at times, since stuff needs to work with no fuss, and we’re always demanding when on the buyer side (“I’m hungry, where’s my order?”). If you’re open to challenges and are eager to learn and work hard to tackle them, you would be a great fit for Cabify. There could be an open position waiting for you, so check out now: we’d be more than happy to get to know you and hopefully have you on board.

IT Parcelo

José Ignacio Fernández

Staff Software Engineer

Cerrar

Choose which cookies
you allow us to use

Cookies are small text files stored in your browser. They help us provide a better experience for you.

For example, they help us understand how you navigate our site and interact with it. But disabling essential cookies might affect how it works.

In each section below, we explain what each type of cookie does so you can decide what stays and what goes. Click through to learn more and adjust your preferences.

When you click “Save preferences”, your cookie selection will be stored. If you don’t choose anything, clicking this button will count as rejecting all cookies except the essential ones. Click here for more info.

Save preferences