Ship Happens: AI can give you a forecast, but it still can't make the call

Ship Happens: AI can give you a forecast, but it still can't make the call

A WWII forecast, an AI model, and why I still think you need a design team

I recently watched the film Pressure and it sent me down a rabbit hole that started in meteorology, moved into thoughts about determinism and predictability, and ended up back in product and AI, which is probably inevitable given I’m a former philosophy student with an ADHD brain.

The film is a dramatisation of how in 1944, two teams of meteorologists looked at the same instruments and reached opposite conclusions about the weather over the English Channel. The American team, led by Irving Krick and working from historical analog matching, comparing the current weather map against decades of past ones and betting the atmosphere would behave the way it had before, read clear skies for the original invasion date. The British team, led by Group Captain James Stagg and drawing on the Bergen School's dynamical analysis of fronts and pressure systems, correctly predicted a catastrophic storm instead; however they went on to predict a narrow, uncertain lull that would provide a window of opportunity on the following day. Eisenhower had to choose which one to believe, with an invasion force already loaded onto ships and no way to undo the decision either way. He went with Stagg and the rest is history.

Whilst the historical analog method of making predictions based on previous patterns was shown to be faulty in the film, I instantly made the connection with the fact that this is fundamentally how LLMs work - making predictions based on vast amounts of historical data. This got me to thinking about the fact that even with AI we haven’t managed to crack predicting weather with complete accuracy. I also wondered what would happen if you replayed that exact decision today, with AI in the room; which of course, someone has already looked into: this month, researchers at the University of Reading (Ed Hawkins and Andrew Charlton-Perez) fed the actual 1944 conditions into one of today's best AI weather models, ECMWF's AIFS. However the results are more interesting than a clean "AI wins" or "AI fails" headline. AIFS captured the North Sea storm, but underestimated conditions specifically at Normandy, and had the storm moving faster than it actually did. For June 6th itself, the day Stagg bet everything on, it produced something like a 30% chance of strong winds and significant cloud cover at the beaches: a real, honest signal of residual risk, not a confident all-clear. There's a good irony buried in the methodology too: the researchers had fewer usable digitised historical observations available to them than Stagg had on hand in 1944, because a lot of the original paper records were never digitised. More data doesn't help if it isn't actually available to the model, which is its own small lesson about the difference between data existing somewhere and data being usable. The researchers’ own conclusion was that even now, with eighty-two more years of atmospheric science and a genuinely state-of-the-art model, human expertise remains essential for interpreting what a forecast is telling you and deciding what to do about it.

Determinism is not predictability

This is where my philosophy degree turns out not to have been entirely useless, as what it led me to thinking about isn’t: "AI still struggles with weather", but that two different claims are getting run together, and separating them changes what you should expect AI to eventually be able to do.

The first claim is determinism: the idea that the atmosphere's next state follows necessarily from its current state plus the laws of physics, nothing random, no dice being rolled anywhere in the system. Weather is about as clean an example of a fully deterministic physical system as exists in nature.

The second claim is predictability: whether anyone, or anything, human or machine, can actually work out in advance what that next state will be. It's intuitive to assume these are the same claim, that a fully determined system must, with enough data, become a fully predictable one; but this isn’t the case. The atmosphere is chaotic, which in the technical sense means tiny, unmeasurable differences in today's conditions get amplified rather than averaged out as they propagate forward, until a difference too small to ever measure produces a completely different outcome within days. This isn't a statement about the quality of our instruments or our models but a mathematical property of the system itself.

Once you see the split between determinism and predictability clearly, it changes what you should expect better AI to deliver. A stronger model, more data, more compute, can push the predictability horizon out somewhat. AIFS is a genuinely enormous improvement on what Stagg and Krick had access to. However what none of that can do is make the underlying system less chaotic. The gap between "here is the best available forecast" and "here is what we should actually do" isn't a symptom of AI still being early. It's what's left over once you've built the best possible forecasting tool, and it doesn't shrink as the tool improves. It just relocates, to whoever has to act on the number.

I wrote a year ago, in a post about learning to build with tools like Cursor, that AI can dramatically speed up how we build, but only thoughtful, skilled people can decide why and what we build. A lot has changed in a year, model capability most obviously, but I haven’t revised my view and I think the D-Day story is great example of why.

The last year has given us plenty of reminders of what happens without judgment in the loop: Ford had to re-hire 350 engineers it had just let go, at a higher rate, when it realised it’s automated system was missing quality issues; and a collective lawsuit demonstrated Workday's AI candidate-screening tools, used across a platform that processes over a billion job applications, allegedly filtered out candidates by age, race and disability without human review, in a case a federal court has allowed to proceed as a class action potentially covering millions of applicants over 40.

Why I think cutting design teams is a mistake

Sadly, I think a lot of organisations are currently making the opposite bet, at exactly the moment they should be more careful, not less. I recently watched a company decide it no longer needed a design team, on the logic that AI can now generate the screens and the copy.

💡

Producing an interface is not the true value of design. Understanding what a real, often anxious or non-expert user actually needs, whether an interface preserves their agency or quietly removes them from the loop, whether it communicates uncertainty honestly instead of hiding it behind confident-looking copy, that's judgment, not output.

It's exactly the Stagg role. And it matters more, not less, on AI products specifically, where the thing you're asking someone to trust is itself a probabilistic system that will sometimes be confidently wrong. I think this is a genuine, expensive mistake. Producing a screen was never the scarce or valuable part of what a good designer does, any more than producing a probability was the scarce or valuable part of what Stagg's instruments did. What a good designer actually does sits closer to what Stagg did with those instruments: work out what the output actually means for the specific, often anxious, often non-expert person on the other end of it, and design for that person rather than for the demo. Does the interface preserve someone's ability to make their own choice, or does it quietly make the choice for them and hope they don't notice? Does it tell someone honestly when the system underneath is uncertain, or does it dress that uncertainty up as confidence because that reads better in a screenshot? Those are judgment calls, not outputs, and they get harder, not easier, once the product itself is an AI system that will sometimes be confidently wrong.

The same trap, closer to home

It isn't only organisations making this mistake. I think a lot of individuals, including plenty of good product people, are quietly making a smaller version of it too. Even lawyers have started citing hallucinated law cases in live court cases. This is being termed "cognitive surrender"; people are deferring their thinking to AI entirely, rather than maintaining any critical check on it. AI can now produce a fluent, confident-sounding strategy document or deck in minutes. It's easy to mistake a well-written output for a decision that's actually been thought through.

💡

Writing the document was never the hard part either; weighing genuine trade-offs that don’t resolve cleanly, sitting with real uncertainty, being willing to be accountable if you're wrong, still is, and no amount of AI-generated polish does that thinking for you any more than a weather model could do Stagg's job for him. The scarce skill was never the forecast, the screen, or the deck, it was always the judgment sitting on top of it.

I definitely don't think the answer is to use these tools less. I use them constantly, for exactly the reasons everyone else does. The answer is closer to something I've learned building my own AI product: use the tool to get to a draft faster, then treat that draft as a first read of the weather, not as the decision itself. If a plan or a deck or a recommendation feels finished the moment the model stops generating, that's usually a sign the thinking hasn't happened yet, not that it has.

What doesn't change

Model capability is going to keep improving, on weather and on everything else. That's not really in question, and it shouldn't worry anyone who's paying attention to what's actually happening rather than what's being marketed.

What I keep coming back to is that none of it closes the gap Stagg and Eisenhower were standing in back in 1944. It just relocates it. Someone still has to decide how much to trust the number, the screen, or the draft, against what's genuinely at stake if they're wrong. The organisations and individuals who do well from here won't be the ones who generate the most. They'll be the ones who get more disciplined, not less, about the judgment sitting on top of everything AI now makes so easy to produce. So for goodness’ sake, keep hold of your design team!!