AI That Pays Back: Three Tests for Transport and Logistics Projects
Logistics leaders are not short of conviction about AI. In BCG’s August 2026 survey of executives at 30 leading global logistics companies, 97% ranked AI as a strategic priority, 70% had an AI strategy and 67% had a dedicated budget. Yet only 13% said AI was delivering measurable financial impact [1].
Other surveys point the same way. Descartes found that fewer than one in five transportation organizations use AI at scale: 19% of shippers and 15% of logistics service providers [2]. Trimble’s Transportation Pulse Report names inconsistent data as the biggest obstacle [3].
The interest is real. The payback, so far, mostly is not.
The algorithm is the small part
BCG’s explanation is useful because it is not about technology. Their rule of thumb is that an AI program that pays back is about:
- 10% algorithms
- 20% technology and data
- 70% people, process and organizational change
They also found that 80% of survey respondents named a single standalone tool as their most impactful AI, rather than a connected workflow [1].
Put simply, most of the value is decided outside the model: whether the data is good enough, whether the output connects to how work is done, and whether people use it every day.
That suggests a practical way to choose projects before funding them.
The three tests below are Optiyol’s proposed framework, drawn from our work with transport teams. They are not findings of the studies cited.

Test 1: Is it a decision you make often?
Every time a decision is made, an improvement has a chance to pay back.
A decision taken once a year, such as an annual carrier tender, has one chance a year. A weekly forecast review has 52. Daily route plans for five depots over 250 operating days are about 1,250 planning cycles a year.
Frequency does not make a decision more valuable on its own. A carrier tender can be worth a great deal.
But a frequent decision pays back in many small, steady steps, and each one is another chance to learn whether the project is working.
Test 2: Is there a fair baseline?
A project pays back only if its effect can be separated from everything else that changed.
The cleanest way is to keep a baseline: what would the operation have done without the new tool, on the same day, with the same inputs?
For routing, this is unusually practical.
Plan the same day twice under the same constraints, once the usual way and once the new way, with the same orders, vehicles and rules.
The gap between the two plans is an estimated saving. Then check it against actual results as the routes are driven, day after day.
For many other AI projects, the baseline is harder to define, and the result becomes a matter of opinion.
Test 3: Does its output reach the road?
A recommendation on a dashboard waits for someone to read it, agree with it and act on it. Each of those steps is a place where value can leak away.
When the output is the work itself, that leak disappears: a route sent to the driver app is the job list for the day.
Execution also closes the loop.
When drivers complete stops in an app, the plan learns what actually happened: which stops took longer, which addresses were wrong, when the day really finished.

Data quality follows daily use
Data quality and integration are the leading barriers to scaling AI in the Descartes survey:
- Data quality: cited by 45% of logistics service providers and 33% of shippers.
- Integration complexity: cited by 38% of logistics service providers and 39% of shippers [2].
It is tempting to treat this as a prerequisite, a cleanup project to finish before AI can start.
In practice, daily use exposes errors, and a feedback loop turns them into better data.
A wrong address in a monthly report waits for someone to notice. A wrong address on a route is flagged by the driver the same morning.
A finish time that never matches the plan shows up in planned-versus-actual comparisons within a week.
When those errors are fed back and corrected, tools that sit inside daily work become one of the most effective data quality programs an operation has.

Scoring three typical projects
Project | Frequency | Baseline | Execution |
|---|---|---|---|
Annual carrier tender | No | Yes | Partly |
Standalone forecast dashboard | Partly | Partly | No |
Daily routing connected to execution | Yes | Yes | Yes |
The carrier tender has a clear baseline in the current rates, but execution is only partly in its hands: awarded carriers do not always accept every load at the agreed rate.
The scoring is illustrative; results depend on how each project is set up.
A forecast connected to replenishment or staffing decisions would score higher than a standalone dashboard.
It does not say the first two projects are not worth doing. It says they will pay back in fewer, larger steps, and that their value depends more on what happens after the analysis is delivered.
Where we would start
Our recommendation: Start with frequent decisions, a fair baseline and a clear path to execution. Check estimated savings against actual results, and feed errors back into the data.
Start with one decision that passes all three tests, measure it against a baseline for a few weeks, and only then widen the scope.
A small project with a clear daily result builds more confidence than a large one whose value has to be argued.
Keep the people side in view from the beginning.
If 70% of the effort is people and process, the planners, dispatchers and drivers who will use the output belong in the project from the first week, not at go-live.
How Optiyol fits
Optiyol combines route optimization, a driver app and real-time tracking [4].
Plans go to drivers as their job list, and actual stop completions and finish times come back for comparison with the plan.
That makes daily route planning straightforward to measure against a baseline, and lets the data improve with use.
Frequently asked questions
Is route optimization AI?
It is decision optimization, a field that predates the current wave of AI. The tests in this article apply to any data-driven decision tool, whatever it is called.
Should we fix our data before starting?
Fix what blocks the first use case, then let daily use surface the rest. Waiting for perfect data usually means waiting indefinitely.
What about generative AI?
The same tests apply. A generative AI tool that changes a frequent decision, has a measurable baseline and sits inside daily work is more likely to pay back than one used occasionally.
Sources
[1] BCG (24 August 2026). Fink, M., Distler, J., Garcia Escudero, R., Weidmann, M., Mikulla, D., Giorgetta, F., Sanders, U. and Reynolds, J. Why AI Isn’t Delivering ROI in Logistics.
https://www.bcg.com/publications/2026/why-ai-isnt-delivering-roi-logistics
[2] Descartes Systems Group (16 September 2026). Descartes 10th Annual Study Finds Transportation Technology Investment Has Increased.
[3] Trimble (2026). Transportation Pulse Report 2026.
[4] Optiyol. Platform Overview.
