Tin Foil Hat Time: Why are AI Models hacking the internet - Part 1
Foundation Models are breaking loose. Should we care?
You have likely already heard about the OpenAI and Anthropic models that broke loose from their training environments and hacked their way into various organizations like Hugging Face.
This post isn’t going to rehash that news but instead to pose a question as to why two of the companies leading the way for foundation models both had similar cybersecurity incidents.
At the moment, the models themselves don’t scare me. It's the possibility that the “geniuses” running this are not competent enough to properly air-gap their training environments.
How difficult is it to unplug these models from the internet?
Theoretically, they should have been running these in a proper Air-Gapped environment; no wifi, no bluetooth.
The rigs they run these models on have to be custom built anyway to get those beefy GPUs, so it shouldn’t be too hard to get the hardware built without most of the comms methods that come standard in PCs.
Of course you could run these “in the cloud,” but considering their main competitors are “the cloud” providers, that would be a bit risky.
I am getting off on a tangent, but the bottom line is that the model was being run in an environment where it could get access to the internet and go off the rails.
Is the model to be blamed or the people that set the model up so that it could go off the rails?
I would think the human in this scenario is to blame. You don’t blame the hammer if you accidentally hit your hand while using it.
The next question is: are these escapes due to incompetence on the part of the humans testing these models?
OpenAI and Anthropic pay ridiculous salaries for some of the top talent.
They clearly train the models on infosec topics, so you would assume they threw some cash at top info sec talent.
How is it possible that none of that high-priced infosec talent thought to airgap these models?
Another thing that concerns me is that if they can’t unplug it from the internet, then they likely don’t have a plan to pull the power plug if these things go off the rails.
Now that only applies if we want to blame the incident on incompetence. There are alternative possibilities which I will examine in a future post.
For now, put your tin foil hats on and let me know what your theories are in the comments.