Sep 30, 2026
This post features an unusual kind of rider in Cabify, a mostly ride-hailing company. In this case, we’ll talk about parcels (instead of people) riding vehicles. So let’s introduce Parcelo (Cabify Logistics’s pet) and see what makes it special:
In general, the main difference lies in how social parcels are: you can put several of them in the same car and they won’t ever complain about not knowing each other (don’t try to get strangers to ride the same car). Parcels are also typically more patient, and can wait depending on their so-called “delivery window”: e.g., hamburgers need to be delivered as soon as possible, while your Amazon order can wait a bit more.
This diversity of requirements takes us to distinguishing the main ways of delivering parcels: the different operations.
One kind of operation we do in Cabify is cross-dock: a driver visits multiple pickup points to collect parcels in what’s known as a “first mile”. These parcels go to a warehouse where they get moved from the first-mile vehicle to a new vehicle that will do the “last mile”: a trip to their final destinations. The key trait of cross-docking is that parcels move directly from one vehicle to another, without intermediate storage. Visually, picking up and delivering through some first miles and last miles looks like this:
In another kind of operation, parcels get delivered without a central warehouse, and it’s the closest to the ride-hailing problem, where you request a Cabify that takes you to your destination. It is the point-to-point delivery that takes place with food: once the pizza is ready, a driver gets it from the restaurant and takes it to the buyer as fast as possible to keep the pizza warm. It’s a similar case to what some online stores offer: you order a phone, a driver picks it up and you get it at home within 2 hours (same operation as with the pizza, just lighter time constraints).
Note that there’s not a one-to-one mapping between operation type and delivery window. In general, the most urgent deliveries are served through point-to-point (like ride-hailing) instead of cross-dock operations, which by design add delay: parcels need to get sorted at the cross-dock, so it is typically used for same-day or next-day shipments.
So what’s stacking? We mentioned several parcels are collected in a first mile and delivered in a last mile. That involves complex routing algorithms to identify which parcels should go into which vehicles to save distance and driver time, both for the first mile and the last mile. Similarly, we could strive to optimize point-to-point operations, and try to deliver two or more pizzas if there’s a good route that can preserve the tight delivery window of food. This is the objective of “stacking” the pizzas (which literally means putting one on top of another!): applying those same routing techniques to make more than one pickup and more than one delivery in point-to-point operations, where no central warehouse is involved. Visually, the point to point operation with stacking would look as follows:
We just casually mentioned the word “algorithms”, which sounds like we’re getting to the fun part. Since hundreds of companies are dedicated to logistics, routing is a well-studied problem with existing solutions.
The problem’s complexity is equally famous, and unfortunately brute-forcing is simply not an option since the number of combinations can explode with not too many parcels. Approaches where we check all possible solutions and filter by constraints will just not work (and no, there’s no foreseeable future where computational power will change this). Instead, algorithms rely on exploiting the limitations of each specific problem, or use heuristics to reach suboptimal solutions.
What are these “limitations of each specific problem”? There are actually a ton of cases. In general, this is called the Vehicle Routing Problem (VRP): delivering multiple parcels in multiple vehicles to a set of destinations from a single pickup point, but there are many variants:
The list can grow in boring ways, but the key lies in understanding how certain algorithms solve subsets of these problems. One typical one is the Clarke-Wright algorithm (also called the “savings” algorithm), which focuses on the most traditional VRP with a single pickup point and can reduce the complexity compared to other algorithms. There are variations for time window restrictions, or adapting it to first miles (by reversing routes compared to the original problem where there’s a single pickup point and multiple destinations). But it just won’t make it if more than one warehouse is required (i.e., no support for multipickup or multidelivery) or additional constraints are needed. Other algorithms include insertion (which iteratively builds a list of deliveries by inserting parcels), genetic algorithms (where solutions get iterated via evolution and natural selection), branch-and-bound (close to brute force, just with clever pruning), or linear programming (where constraints are represented by linear relationships).
In our case, Cabify uses an insertion-based algorithm that tackles the Pickup and Delivery VRP and is applied even when we limit ourselves to a single central warehouse for pickups: this reduces maintenance (one algorithm to rule them all, not one per case), while also facilitating the addition of new constraints. Order of complexity is not such a problem for real-time stacking, since pickup origins typically keep a low stock count (< 15 parcels), and the time constraints plus the limited space of vehicles make it unlikely to require processing thousands of parcels on each run. This contrasts with first-mile or last-mile routing, where you can easily have thousands of parcels to route, but real-time performance is not a requirement. We keep benchmarking other approaches, including open-source frameworks like Verso’s VROOM and Google’s OR-Tools.
What exactly are these time and space constraints of parcels and vehicles? Quite obviously, a vehicle cannot take an unlimited number of parcels, so there’s a capacity constraint. Usually parcels are defined by weight and volume (length, width, and depth). Similarly, a vehicle’s boot (trunk) would follow the same model, determining which parcels could fit in. In practice, we don’t usually have such details: no one is measuring your burger menu, and stores doing 2-hour-delivery are also busy enough fulfilling orders and don’t want to measure parcels, so we typically have to deal with not having volumetric data and assume an approximate number of parcels vehicles can take.
In contrast, time constraints can be very strict. Not meeting a delivery window might mean not making money, regardless of how many minutes late. But there are many different times involved that we need to take into account:
All this is heavily connected to routing. In the typical travelling salesman problem, the different stops are visited according to the shortest path possible. But this is not necessarily the case when we have constraints for each parcel, which could cause some reordering to ensure we arrive just-in-time at each stop. Consider the following example:
You can see that although route A is shorter, we might need to use route B if parcel C has a tight delivery window and we need to reach it earlier to meet the constraint.
Another interesting case is shown in the next example. Let’s assume there are no pickup or delivery window constraints. Which route is better?
Route A looks shorter, while route B is odd because it circles back to an intermediate stop that could have been visited earlier. So why the question? The answer is: it depends. Route A finishes at a different location than route B, and that’s the key to deciding. If drivers finish deliveries in high-demand areas, that’s good for keeping them busy. So estimated demand is a deciding factor, but not the only one. For example, one of our clients uses dedicated fleet to cover their deliveries, since their warehouse is located in a low-demand area. These drivers are hired to only deliver the client’s parcels and are expected to return to the pickup point for more deliveries. This means we need to consider the path of the driver moving back to the pickup point (or to a high-demand area). Once we account for the back-to-origin segment (the dotted line), route B is the better choice:
What do we do once we have a set of routes for all the parcels that need to be delivered? Those routes are the output of our algorithm, but that does not mean we should dispatch a driver for each one right away. They represent a snapshot of the current optimal routing, but this picture could improve once we receive additional requests. Suppose we have a route with 2 parcels that, due to their time constraints, can wait for a bit: what’s the likelihood of receiving a request for delivering more parcels that fit in this same route? We could have a driver taking those 2 parcels along with, let’s say, 3 more parcels. So it sounds like we’re trying to see the future, but since parcels can wait (if their time windows are met), we can actually see the future just by waiting for it.
This takes us to the delivery triggers of routes, which are mostly two: the time trigger and the capacity trigger. The time trigger is straightforward: if a route involves parcels and we need to start the trip now because otherwise we’ll miss the delivery window, we start looking for a driver. Going back to our earlier example, if we have two parcels in a route and they can wait, we’ll wait; otherwise we’ll dispatch now. Technically, we’re checking against the maximum start time defined above: if we’re reaching it, we should start the delivery or we’ll be late.
The other trigger is the capacity one, which is even simpler. If a route involves as many parcels as the vehicle can hold, it’s ready to go. In our two-parcel example, if they fill the whole boot (trunk), it doesn’t make sense to keep waiting for future parcels, as we could start the delivery right away.
All this is orchestrated through continuous monitoring of pending parcels in what we call rounds: every few minutes we run the routing algorithm and check the route candidates against the triggers. If any trigger fires, the route candidate becomes a real delivery, and we dispatch it to a driver.
How’s this performing right now? We’re currently stacking parcels for restaurants doing food delivery and a store that does 3-hour deliveries of consumer goods. Food delivery has tight delivery windows (nobody wants cold pizzas), and we’re delivering around 1.9 parcels per delivery (meaning we’re almost doubling the standard performance of “one meal per trip”).
Regarding the consumer goods store, it has a warehouse on the outskirts of Lima, which combined with the wide delivery window, offers a big opportunity for stacking. Typical travel time to the center of the city takes more than 30 minutes, and it’s a trip common to all parcels, so unless delivery addresses are very far from each other, it usually pays off to take many parcels at once. However, driving times to the center can reach 50 minutes during peak hours, and with almost an hour needed to fulfill parcels, we only have around another hour for deliveries during those peak windows. The actual destination locations matter a lot when building the route. On average, we’re stacking around 5 parcels per delivery, with more than 95% on-time-delivery success rate (the longer the route, the higher the probability that a delivery issue causes late deliveries for the last parcels).
Would you like to know more? Real-time operations such as stacking can be stressful at times, since stuff needs to work with no fuss, and we’re always demanding when on the buyer side (“I’m hungry, where’s my order?”). If you’re open to challenges and are eager to learn and work hard to tackle them, you would be a great fit for Cabify. There could be an open position waiting for you, so check out now: we’d be more than happy to get to know you and hopefully have you on board.
Staff Software Engineer