Anthropic Discovers AI Model Hacked into Three Other Companies During Testing
Transcript
It's a pleasure. Another major tech company has gone public with claims that its artificial intelligence models have autonomously hacked into another company. Anthropic said it's discovered that it's AI Claude gained unauthorized access to three other company systems during testing back in April. The company's investigation was prompted by a similar disclosure from OpenAI last week, and it's led to further calls for tighter restriction on AI companies. Cam Wilson is the ABC's national AI technology reporter and is with me in the studio. Hi, Cam. Hi Ross. So what did Anthropic tell us today? Yeah, it's a bit like Groundhog Day. I'm back here saying another company says that they have discovered that their AI models have hacked into another company. Has put out a blog post this morning saying that it has looked over its evaluations that it's been doing on its AI models, various models that it has, some that have been publicly released, some that it's just testing, and found that over tens of thousands of evaluation runs, so it's doing these tests internally, that in at least three instances the models had gained access to the internet and then gone on to gain access to another company's systems via the internet as part of completing its challenges. This, of course, is a you know a way, a nice way of saying that it's essentially hacked into these other companies without being explicitly told to do so. Do we know what Anthropic was testing? Yeah, so these models that these companies have are capable of a lot of things, and and the companies look at them and say, well, we've got to figure out exactly what we're dealing with here. Not only can it what it literally do, but also how it reacts in certain situations. Because you know, you give tasks or challenges to these bots and they want to understand how it's going to react. Unlike traditional computing, which is what they call terministic, you know, it's kind of like calculator. You say one plus one, you're supposed to get two. With AI, it's you know, one plus one equals well, who knows? It depends on whatever it's been trained on. So, in the process of testing them, where they do these challenges, uh, including this one known as capture the flag, where they say, Hey, AI model, I want you to go and get this thing. We'll put you in this little uh environment that's supposed to be cut off from the internet, and you do whatever you want to do, whatever you think you need to do to get that. That was the testing that it was under that led it to then go and seek the answers from these other companies' systems. So, do we know how the AI model managed to access those systems or as you say, essentially hack those systems? Yeah, so the evaluations were supposed to be done in an environment that weren't connected to the internet. In fact, the models had actually been told by the company, so the AI itself had been told you're not on the internet. So, whatever you do, you know, you don't have to worry about hacking other people. Unfortunately, it seems that they actually were connected to the internet that a third-party provider that was essentially providing the software, you know, the almost like the playground where this was happening, had left it open to the internet, and then as a result, these AI models, which in some cases knew that they weren't supposed to go out on the internet and do these kinds of you know unauthorized access, decided, well, we were told this was on the internet, so clearly what's happened is we actually just have all this in our playground, and so we went and did these things thinking we were doing the right thing, but not actually knowing that in reality they were out there, you know, wreaking havoc. This happened months ago, right, Cam. So why are we only hearing about it now? I think it's a really good question. Um, so last week OpenAI announced that it had discovered a similar thing. This is kind of the first time that we knew about that. I guess not to be outdone. Anthropic has now discovered it has had the same thing. It said last week's disclosures had made them look into their own testing, had made them check it out and discover that this had happened. And I think this raises questions, which we're hearing from AI safety experts, which is, you know, companies say they have this really, really powerful uh AI technology. It's capable of doing these things. How come if these things are happening? These incidents where you're going out in the open, you're affecting other companies, potentially causing real damage, and you don't even know it's happening until after the fact. I think a lot of people will not only want to know why it happened in the first place, but why it took so long to come out.