People have found many uses for AI. Some use it as a coding tool, others for architecting and many just use it as a day to day companion for solving a variety of problems. More recently, AI has been useful as an archaeologist and a tour guide, to go from having little knowledge of a codebase to a more detailed understanding.
I once went to Naples and there I saw the ruins of Pompeii beautifully preserved. This whole civilisation was kept largely intact and explorable by the hard work of those wielding the shovels and pickaxes. The whole experience was brought together with the help of a multilingual expert who showed us around and answered questions on the ancient lives of the Vesuvians. Maybe you can see where I am going with the analogy here: while not every repository is covered with volcanic debris and fallout, time and unfamiliarity do tend to make their effects known and bury codebases with complexity and barriers to entry.
Codebases have always had a layer of rubble/ash/dust/tech debt/illegibility (mark as appropriate). As a new engineer coming to work on them or as even an existing one trying to remember what was going on for maintenance and upgrades, the depth of this layer has always been something to overcome before any productivity. Sometimes, context was locked away in the minds of the early engineers; other times, documentation simply is not informative. In defence of the engineers however, active codebases are living things and there is an aspect of learning from the code itself rather than having a manual for it. Nevertheless, for anyone outside of the context or without the history and experience of ‘living’ the code, old codebases can be difficult to dust off.
The influx of AI has brought with it many a Prometheus moment – and perhaps another one can be in the code archaeology world. To test this hypothesis, we decided to let an enterprise level Claude break out its shovels on a couple of repositories and see what it could deduce.
The Software Archaeology Mission
The mission given to me was to approach a new codebase with very little context, and to use AI to help fill in the blanks and reconstruct the context, requirements and business decisions that led to the resulting code. The results are very promising.
Given the breadth of this challenge, my initial thoughts were that Claude could struggle to define what was required. However, it quickly understood the brief and began to understand the repositories and also the commit histories too to build up a picture of how they came into being and what was prioritised. The inference it was able to pull out from this process was very useful in building up a picture of why the software was built and what the requirements were.
Of course, this inference is not the whole story – it must be validated where possible against the reality of the software. However in many cases this may not be possible. The original developers may have left, and even if they are still involved in the project, there is inevitably some context loss and historical irretrievability of artefacts, conversations and ideas. Much like with the Vesuvians themselves, it can be difficult to talk to history directly and therefore we must infer and validate with evidence where possible from what remains.
There is an aspect of being a good detective with archaeology. Archaeologists also look at piecing together evidence to present a compelling case. However good detectives sometimes make leaps based on their ‘gut’ feeling, which AI lacks. Knowing the code and the recorded commits is one thing; knowing the people who committed the code is another.
Machine & Human Learning
The exercise led me to wonder what is actually important to have as ‘history’. Given that AI is now increasingly responsible for writing the code, perhaps the most important thing to record are the prompts or the conversations the ‘director’ (i.e. the technical person responsible for implementing and creating the software) has had with the AI. When the actual code is no longer written by hand, the direction itself is what becomes important in figuring out the what, how and why behind the code.
There is also another world entirely which deserves its own research, which is where the AI directs itself and writes on its own. It is likely that archaeology in this area would be more of an interview rather than an investigation. What is in no doubt however is the fact that there has been an irreversible shift towards this as the human in the loop aspect continues to diminish.
Interestingly, one of the more fundamental aspects of AI has always been the ‘learning’ part of machine learning. That learning involves not only consuming and processing lots of raw data, but also recognising the patterns and building a narrative. Therefore it could be said that a certain level of archaeology is built into AI and it is one of its main skills. The difference in the future however is that while previously, most or all of the inputs were created by humans, increasingly AI will learn from things it has created itself. In some ways this mirrors the archaeology of humans; we can learn from dinosaur fossils and ferns as much as we can from our ancestors in Pompeii. What will be interesting to understand is how AI learns from itself and how it reinforces itself.
The Practical Matters of Software Archaeology
I gave Claude some very simple, more open-ended prompts (these have been sanitised for confidentiality) in order to find out what each repository did, what its purpose was and to infer the requirements that led to the code. Note that this is AI inference of human decision making, so it was not as accurate.
“You are a software archaeologist. Imagine we have just found this code with no real context or idea about its purpose. Look into the entire codebase and describe to me in simple terms what the purpose of the code is, its architecture and its main features and the wider context that you can develop as a narrative surrounding it. Infer the requirements and the decisions made from the code and its history. Summarise this in the form of two documents, one for the general reader and one for a Tech Lead who will take over this project.”
The purpose of keeping things relatively open ended and informative was to have a more narrative driven approach to the codebase, in a similar way that project handovers may have occurred between engineers. The results were useful and led to a number of follow up questions which might have been asked to uncover deeper aspects of the codebase and the implementation of certain features, especially once the lay of the land had been established.
“Based on the exploration, tell me how I can be productive in this codebase and modernise where necessary.”
“Outline fixes required for ongoing usage of this codebase for purely maintenance purposes; include dependency upgrades and knock on effects and describe the requirements for retaining existing functionality as is.”
Deeper Exploration and Using the Results
Armed with this information, conversations with the client become a lot richer, more contextual and outline what is necessary and what is more optional in terms of either maintenance or follow up work. This can help with prioritisation and smarter allocation of time and resources.
These prompts also led to some finding of ‘gotchas’ and plans to remedy these which may have been overlooked or unknown during the time of development. These may be hallucinations, not actual problems, known flaws or simply too minor to worry about however given the nature of the open-ended prompting it is expected that problems would be found.
To be metaphorical once more, the codebases used here represent Pompeii before the Vesuvius eruption but where the original builders of the city have moved on. AI has been used not strictly as a pure archaeologist but as something of an auditor and detective to go into the city, to look and to analyse and give new builders the broad picture. It has wisely pointed out the fact that there are areas where mini eruptions may occur and where something like a fully outdated dependency may lead to a wider and more cataclysmic eruption in the future. Where archaeology comes in is to look at what is there and infer the requirements and decision making in the city, to explain them in a useful way. That allows the new builders to make further decisions based on the foundations, preserve what needs preserving and upgrade where necessary.
AI’s role in Software Archaeology in Summary
It is clear that as code proliferates, either from humans or AI, good software archaeologists are going to be more and more necessary to either hand over knowledge or even to understand the purpose of the software itself.
The good news is that AI is very good at creating narratives and attempting to recreate context where it may have been lost or obscured. The risk is that this may not always be accurate and a healthy degree of skepticism should be employed by any human in the loop to ensure that false conclusions are not drawn. If there is nothing to validate against, this can be mitigated with this understanding: if code was written by a human, it is more likely to be understood by another human than an AI. This is simply because of the logical leaps that humans can make that AI is not quite capable of yet. The future however is going to be more machine-written code, and therefore it is important that context and wider decision making is recorded.
In short, documentation is still king, folks!