Maritime logistics
The arrival-time model that had no arrivals to learn from
A terminal operator wanted to predict vessel arrival times to plan berths, tugs and labour. The data the model would learn from turned out not to exist in usable form, so the money went to the contract that produces the data instead.
Evidence stopped a project that every stakeholder wanted.
Problem
Berth planning is done 24 hours ahead on estimated arrival times supplied by carriers. When a vessel is late, the gangs, tugs and cranes booked for it stand idle; when it is early, it waits at anchor. The operations director asked for a machine-learning model to predict true arrival times from AIS tracking.
Context
A mid-sized container terminal with a planning team of six. AIS feed already purchased. Carrier ETAs arrive by email and EDI, re-keyed by planners.
Current architecture
A terminal operating system of record, an AIS subscription feed used only for a live map, and a spreadsheet the planners actually work from. No historical store of predicted-versus-actual arrival.
Constraints
- Data: the only arrival history is the TOS berth log, which records when a vessel was worked, not when it arrived.
- Integration: carriers are contractually free to update ETA at their discretion.
- Skills: no data engineer; the BI analyst is at capacity.
Evidence
Each statement placed on the ladder before it was used.
ETA is missing from roughly 36% of AIS messages, and of the vessels that do publish an ETA, about 27% do not arrive within 24 hours of it. Over 75% of vessels fail to keep destination and ETA updated in the voyage data.
AIS voyage fields (destination, ETA and draught) are hand-entered by crew and carry documented human-error rates.
Berth planning errors cost the terminal a meaningful sum per year in idle gang time. The figure was estimated by the operations director, never measured.
A model trained on AIS position histories could beat carrier ETAs by enough to change a berth plan.
Any prediction must be available 24 hours before arrival, because that is when the labour order is placed. A better prediction at 6 hours changes nothing.
Questions that changed the answer
- What is a berth-plan error actually worth, measured rather than estimated?
- How accurate would a prediction have to be, at the 24-hour mark, to change the plan?
- Do we have any record of predicted-versus-actual arrival to train on, or only of work performed?
- Would the carriers share their internal ETAs under a data-sharing clause?
Options
Including the one nobody wanted to discuss.
Build an ETA model on AIS histories
Ingest AIS, reconstruct voyages, train a model to predict arrival at the 24-hour mark.
What it costs: Depends on a label (true arrival time) that is not recorded anywhere, and on inputs that are missing or stale for most vessels.
Buy a port-call prediction feed
Subscribe to a vendor product that publishes predicted arrival times.
What it costs: Transfers the same data-quality problem to a vendor, with no way to audit it, plus a recurring fee.
Fix the data first
Record predicted-versus-actual arrival in the TOS from today, and add an ETA-update clause to the next carrier contract round.
What it costs: Produces nothing for several months and is unglamorous. It is also the only option that creates the asset the other two need.
Do nothing
Keep planning on carrier ETAs.
What it costs: The cost stays, and stays unmeasured.
Economics
Four horizons, not one estimate.
- Build
- Model build was the smallest line. On the team’s own estimate (an estimate, not a measurement), voyage reconstruction and a historical store came to several times that, and were the part nobody had costed.
- Run
- AIS feed already paid for; retraining and monitoring would be a new permanent task for a team with no data engineer.
- Change
- Every carrier onboarding changes the input distribution. A model here is a subscription to continual rework.
- Exit
- Low for the build; medium for the vendor feed, whose predictions would have been embedded in the planning routine.
Decision
Do not build the model. Instrument the process that would produce its training data, and revisit in two planning cycles.
Why
The project's central assumption, that the terminal had arrival history to learn from, was false. What it had was a record of work performed. Without a label, there is no supervised model, however good the AIS feed is.
Why not the alternatives
- Build an ETA model on AIS histories: No arrival-time label exists, and the dominant input is missing or stale for the majority of vessels.
- Buy a port-call prediction feed: Buying does not fix the input quality; it only moves the unexplained error behind a contract.
- Do nothing: The cost of poor berth planning is real even though it is unmeasured. Leaving it unmeasured was itself the problem.
Trade-offs accepted
- Several months with no visible deliverable, in exchange for an asset that makes the decision answerable later.
- A contract negotiation the commercial team did not ask for.
Reversibility
Recording predicted-versus-actual arrival is a small, additive change. If the economics never justify a model, the record is still useful for carrier performance conversations.
Non-functional requirements
- Prediction must land 24 hours before arrival to be actionable
- Planners must see why a prediction changed, or they will not trust it
Decision gate
On the information available there was not sufficient justification to proceed with the model. The gate names precisely what would move it to GO: six months of predicted-versus-actual arrivals, and a measured cost of berth-plan error.
Implementation
Two days of work in the TOS to stamp arrival events, a nightly export to a single table, and a one-page definition of what counts as an arrival. No platform, no pipeline, no model.
What the decision was expected to achieve
Within two planning cycles the terminal can answer three questions it could not answer before: what a berth-plan error costs, how good carrier ETAs actually are, and whether a model could beat them. The decision was deferred, not avoided.
No outcome is claimed. This is an illustrative example, so there is nothing measured to report, and a real engagement would state what happened and how it was verified.
Lessons
- The most expensive part of a machine-learning project is often the label, and it is usually discovered last.
- A model cannot be more reliable than the field a deck officer types in by hand.
- “Don't build it yet” is a finding, not a failure, provided it names what would change the answer.
- A deferral is only a decision if it carries a date. This one did: two planning cycles, with the three questions written down.

