Thunderstorms and Strikes
[Note to the reader: This write up was supposed to be a white paper. However, during the course of the evening, rather than discussing the pitfalls of open and closed coding, the interest of all parties was focussed on the third way provided by our sponsors. What follows, therefore, is a very personal account of what proved to be an illuminating discussion on a very exciting product]
Record breaking temperatures, the first Data Observatory roundtable, beautiful location; the perfect opportunity to host the event as a garden party in conjunction with Anaconda at the stunning grounds of Battersea Park to discuss one of the most pressing topics in data – open or closed coding.
At least that was the idea. The days before the event suggested that the weather was deteriorating and this was doubled with our friendly regional weather forecasters stating that we were in for thunderstorms all evening. Maybe that wasn't the best time to have a garden party under a marquee after all.
As the sponsor for the event, the team at Anaconda (led by the amazing Frank and Maisie) simply arranged for their suppliers to host the roundtable at their own venue and festoon the restaurant with all the greenery and blooms you could hope for. Seriously, in 48 hours the Anaconda team had worked miracles and the food promised even more
Besides, London's intricate public transport network would keep us out of all but a few drops of rain on our way to a new venue. What could go wrong?
Sorry, what's that you say? Tube strikes? Ah.
For those of you who know London, know that even though we were "south of the river", an area bereft of much underground (on account of the heavy clay and flint soil making it almost impossible to tunnel apparently) so we might get away with it. As an old stomping ground of mine I noticed that the venue was on Lavendar Hill and the number 77 bus stops just outside which can be picked up at a number of overground and underground connections. Then I realised, I was not there yet but giving all the advice in the world about the bus route only to recall… I haven't been on that bus for 25 years! For all I knew it could be a non-stopper to Edinburgh nowadays.
Fortunately, it hadn't changed one bit and, thanks to the speedy work from our sponsors, everything was set for a fantastic evening, despite having lost two attendees who were unable to come because the journey home was no longer possible.
We all arrived to witness the aesthetic spectacle of food laid out before us, I poured myself a glass of Pimms and got instantly engaged in the discussion led by fascinating words of David De Santo, the CEO of Anaconda.
A few days before the event, there was a "story" that, due to the mythos breach, NHS England were going to mandate that all coding was to be held be the developer of that coding and not to be shared until prior approval was given (which, given some sign-off processes, may as well be forever in technological terms). It turns out that our media colleagues had got the wrong end of the stick (no, really) and inflated a one-off sharing of code that should never have been shared, with a mandate for all code. We were reliably informed that this was not the case.
Nonetheless, it does draw into question, how safe is the code that we are using?
Given the nature of our hosts, let us assume that a developer has developed a model in Python. Now, this developer is a fairly good coder and has fairly decent Python knowledge. But all developers copy their own coding, they may even borrow chunks from other developers or even get AI to produce it if it is a simple but time-consuming piece. We would like to think that whatever is developed is resilient against attack and, perhaps more importantly, does not make available data that shouldn't be, or make assumptions on that data that would fall outside of ethical boundaries.
As good as our developer may be, if we are honest, how often do we rigorously check a block of coding that we have used countless times before to see how it interacts with other code we have written countless times before. Sure, if we are relating it to a significantly different data structure but for the most part it'll be fine, right?
Well, sadly not. If we are being honest with ourselves our industry is made up of many different skills levels. Some borrow more than others and some have a stronger grip on data privacy or data ethics than others. Some will have experience of recycling code that doesn't work properly in new structures, even though it appears that it does.
This poses a risk to sharing code. The more we share, the more we need to check and the greater the opportunity there will be for vulnerabilities to slip through the net.
As a result, should we share our code? Should we use the code of others?
A libertarian view may well say "why not?" after all, suboptimal code will be identified and rectified faster if used in multiple scenarios. There is no better way to improve on an idea than democratise it and let the developer community give it a nudge in the right direction.
This argument would be fair enough in the gaming world or another sector where failure is only an inconvenience (albeit often an expensive one). But we work in health and care. We don't work in an environment where someone can access cheat codes or, at worst, someone's personal email address (financial information is another subject and one which has to involve those institutions). We work in an environment where data is intensely personal and often still in the process of corroboration.
The pain that could be caused by coding accidentally making HIV test data publicly available doesn't bear thinking about.
So maybe that's it. Like the erroneous article suggested, maybe we should lock all coding up. Sure, that means no replicable analytical pipelines, but similarly, it means no exponentially spread vulnerabilities. But for organisations like the NHS who have, even now, over 7000 data professionals, surely this is just going to double our workload?
This is why were we so fortunate to have Anaconda as a partner for this roundtable. They offer a third way. As far as my non-technical brain can gather, if one passes a developed application to them (preferably Python but they are spreading their remit), they can check the code and enable it to be reproduced at any other NHS site, assuring governance and (to a degree) security along the way. One may develop in Anaconda or just use it to check code, but with their agentic AI along with human governance checking, they can make it safe for the internal market.
Now, I cannot say that I have used Anaconda's solution. I couldn't even tell you how much it costs or what it offers in terms of value for money. But from an options perspective, it offers a third way. Open coding holds inherent governance and security risks, completely closed coding or personal coding offers grotesque inefficiencies. Anaconda, and potentially other organisations that do the same thing, appear to offer flexibility.
By all accounts this is a service that protects you from governance vulnerabilities inside and outside of your organisation. For me, the idea of replicating a PHM or strategic commissioning model that accounts for local variance is interesting, even compelling. I wouldn't say I have an opinion on the totality of potential use cases out there, but even if there is just one, then that shows us that we have the opportunity to share our code without risking our data. Whether Anaconda are the right people for you is your business (for the life of me I can't see why they wouldn't be), but what they have shown me is that there are still options.
This is the reflection of the round table. I couldn't tell you whether I was being sold snake oil by some extremely talented folk (though I've been around the block enough to strongly suspect this was not the case). However, the notion of "Governance As A Service" is an interesting one. There are a number of organisations who perform this by acting as data owners, but the efficacy and economics of buying this service to provide assurance to one's own coding… that is an interesting one.
Maybe we should just be more rigorous with our code and build assurances into our employment contracts.
Maybe we should accept the losses as collateral damage.
I'm not necessarily saying we should all rush to buy Anaconda products.
But I'd suggest it is worth taking a look at their freeware to start with at least.
We convene leading experts and scientists across health and social care to explore the ideas shaping the future of data.
Interested in partnering or collaborating on a topic? We would love to hear from you.