If I were properly rich, I wouldn’t want a yacht. I’d want a research department: three people in a room down the corridor whose entire job is that when I say “that looks odd”, someone goes and finds out.

(I am obviously reminded here of the bit when, as far as we could make out, YouGov UK were basing their polling questions on Jonn Elledge’s random musings)

Those of you who follow rail matters in the UK might have noticed an apparent increase in rail incidents in the last couple of years. I have had the RAIB in my RSS feed (ask a passing granddad) for the last 20 years and this was absolutely my subjective opinion from the frequency and severity of their reports.

Sadly, I don’t have a research department. I do have a Claude subscription. I mostly use it to automate basic tasks – I wouldn’t have been able to maintain PyCon AU 2026‘s social presence without it, for example – but it hasn’t been something I’ve used for research projects before.

But the app has been nagging me about its new capabilities recently, so I thought I’d give it a go. I went into cowork mode (I decided against trying to explain cow orking jokes to the computer), and asked for a preliminary view on the topic. I gave this prompt to Opus 5:

Railways in Great Britain are very safe, but it feels like there have been a higher frequency of train incidents caused by railway-internal issues (as opposed to crossing misuse by pedestrians or vehicles, or workplace safety issues with track workers) incidents in the last couple of years than the previous two decades.

Can you scan the RAIB website – it should be straightforward to categorise news pieces into initial releases/interim reports/final reports/safety digests to avoid duplication of incidents – and come up with a preliminary view on whether this is correct or just recency bias?

https://www.gov.uk/transport/rail-accidents-and-serious-incidents is a starting point, you should be able to find archives from there. No need to go back prior to the formation of RAIB in 2005.

What came back first was analysis rather than a document. I asked one substantive follow-up, then asked it to package the lot as a PDF with proper attribution and an explicit account of its own limitations. The result is thirteen pages with an appendix, a sources list, and a section explaining why I should not believe it too hard (the last section is there because I asked for it, but the model itself did keep highlighting its own limitations).

The answer, briefly, is that I was wrong. Not even wrong in an interesting way: once you divide by the number of trains actually running, the last couple of years look like the pre-pandemic railway to two significant figures.

Download the report here – it is worth a read if you are interested in This Sort Of Thing. I would be interested in people more experienced in this than me tearing it a new one, while noting that this is the standard you would expect from the equivalent human-penned report.

Two years ago I said this stuff would do a great “junior copywriter who’s very naive but actually competent” job. That was correct, and honestly it’s still mostly true as far as writing goes. But it’s also now a second-year associate consultant (AC2). It successfully pulled all 601 RAIB publications since 2005, coded every one of them to the year of the accident and a cause family, stripped out the duplicate interim reports, and then built a denominator that nobody actually publishes by chasing passenger train-kilometres through six different vintages of ORR PDF. None of that is clever. It is, however, the entire content of AC2 work.

Core point from the report:

At the 2017–20 rate, the traffic run in 2024–25 should have produced 22.9 higher-risk train accidents. It produced 23.

This is the perfect level of analysis for an “ah shit, nothing to see here” answer to the question I actually asked, with a much more interesting loose thread hanging off a wholly different question (p10):

The industry’s causal account [for SPADs] rests on cause forms completed for well under two-thirds of events.

I would be interested to see industry takes on this.

Let’s be honest about the quality here: it remains a preliminary report by an AC2. Its purpose is to survive the first meeting and not make the engagement manager look stupid. And OK, ultimately the reason that bar is clearable is perhaps more flattering to the computer than it is to the consulting industry.

(although, just as with an engagement manager report, I can tell it’s not total garbage because I know what a wrong-side failure is. There is very very strong Dunning-Kruger potential here, as with every aspect of LLM usage.)

But this isn’t one where I’ve coached it and led the work beyond the commissioning prompt. It answered the question with one substantive follow-up from me – this wasn’t a chatbot back-and-forth that I spent hours on. My only response to its first pass was this, based on the concerns it had flagged itself:

Let’s next consider:

  1. rates per train-km
  2. reasons for SPAD increase (it looks like the recent MML disaster was SPAD-related)

For these, it specifically flagged 1) “rates per train-km” as the issue I should look into, but didn’t feel it was within my brief. 2) MML disaster / SPAD increase was my read based on my existing knowledge.

It’s fun to have a proper research assistant on tap. I’ve spent more of my own time writing this blog than I did on the analysis, which is, with absolute certainty, the first time that’s been true of a data post here. I have just as much confidence in Claude as I do in the Excel work that I’ve done for the data analysis that I’ve posted here in the past (cite: Reinhart/Rogoff, 2010).

Fundamentally at this point if you’re basing your view on AI’s capabilities on photos from dodgy takeaway menus and your less adept colleagues’ annoying Copilot emails, then while it’s understandable to think the whole thing is like that, you are wrong.

At the same time:

  • The Singularity is pretend nonsense
  • Superhuman AI doesn’t even make conceptual sense – “cleverer than a human” isn’t a dial you can keep turning up, it’s a couple of dozen unrelated capabilities that all go somewhere different
  • I’m sceptical that we’ll even get something in my lifetime that’s capable of emulating a Big 4 consulting firm engagement manager (even in the pejorative sense)
  • I am rather cringing at the AI header images I made here 2-4 years ago, as everyone and their dog emulates the style, and definitely won’t be doing that again

But what we do have is another tool, like everything from the steam engine to the typewriter, that’ll make the world richer – and in the end Baumol will apply, which is to say that the wages of everyone who didn’t get any more productive go up anyway, and the winnings spread out across society.

Damn, I should write a blog about Baumol.

Image: Rail Accident Investigation Branch, from Report 16/2020. Contains public sector information licensed under the Open Government Licence v3.0.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.