Transcript: The Hugging Face Incident Machine transcript, lightly cleaned up. Speaker labels come from each host's own mic track. Attributions marked (?) and words in [brackets] are editorial guesses. [0:00] Ben (cold open) So now you can ask Mom to get anything you want. You can say, go get me a grenade from the store. Go get me a knife from the kitchen. Mom will do everything. [0:19] Vaden Welcome back to another episode of Increments. I'm one half of your handsome hosts, wearing a very awesome fall... I guess autumn isn't even in fall, never mind. Wearing a fall t-shirt, or fall shirt. [0:34] Ben It is. Autumn and fall are the same thing. All right, am I wrong? [0:37] Vaden You know what, I was mixing up autumn and August. [0:40] Ben Ah, okay, it's a classic. Bodes well for the quality of discussion in this episode. [0:45] Vaden Oh God, off to a rocky start. And to my right, to my left, and my bottom, depending on the screen sharing, we've got Mr. Benny (I just want to hug his face) Chugg. [0:56] Ben Do you want to hug him in the face? I'll hug that little face. [0:59] Vaden (?) He's always on the bottom and he likes to be hugged. [1:02] Ben That's what they say about him. [1:04] Vaden Oh man. And yeah, so we are on Patreon, send us the money there. We're on Discord; we just got a lot of new people joining the Discord from the Conjecture conference, so welcome to y'all. And today we're talking about the Hugging Face incident. I am annoyed by this whole story, I think. [1:25] Ben Who would have guessed? Who would have guessed? [1:27] Vaden Who would have guessed? I think you'd have different views on this, so it's gonna be fun to explore. But yeah, just to lay my cards on the table, I think the whole thing is a bunch of BS publicity-stunt stuff. Obviously I need to substantiate that. [We've] gotten into a few little light skirmishes with our friend and your fellow co-host, Mr. Rich. So I think this is gonna be a fun, lively episode. Anything you wanna say to kick things off? [1:54] Ben No, just that, yeah, I don't have particularly strong opinions going in. I think I'm more sympathetic to the view that's in the zeitgeist of what happened, but I'm sort of excited to explore. And I mean, I'm hoping that you can change my mind into being not concerned whatsoever about what happened. I think there are reasons to be concerned about what we saw. So I'm hoping, as usual, that your rank optimism persuades me by the end of the episode. So I think we're gonna play a short clip from Sam Harris's recent episode with Ryan Greenblatt, sort of just an eight-minute conversation between them, which gives listeners a lay of the land of how this is typically being talked about in public. Obviously the incident has now appeared on many podcasts, including of course Dwarkesh, and I believe Ezra Klein and whatnot covered it as well. It's sort of been everywhere. [2:54] Vaden Yeah, we should also say to the listener: we'll play this clip, and we may add some comments here and throughout, but in general I'm gonna try to hold my overall opinions till the end. We will be screen sharing, and there are gonna be some places where we screen-share documents from OpenAI. So if you do want to switch over to YouTube, you may find that beneficial, but we'll do our best to make it audio-friendly as well. Just a little heads-up to the listener/viewer. So let's get into it by listening to eight minutes from the second most recent episode of the Making Sense podcast, titled "A Coin Toss for the Future," with Ryan Greenblatt. [3:37] Ben And we should say before we jump into the clip that Ryan Greenblatt is the chief scientist of Redwood Research, which is the organization that was given some internal access to OpenAI's logs in order to try and understand what had happened. So when OpenAI realized that this incident had taken place, they asked, I suppose, Redwood Research to come in and write an independent report about it, and they gave them various amounts of compute and access to some of the logs. And then Redwood Research compiled a report about what they believed to have taken place. So they're sort of the interface between OpenAI and the public perception of what's happening. [4:22] Clip: Making Sense #494, "A Coin Toss for the Future" Sam Harris: Okay, I want to talk about the Hugging Face incident. I'd love you to walk me through it from start to finish. Are there other concepts we need to define before you do that? I'm thinking of alignment faking, for instance, and reward hacking. Those are the phrases that we've heard to describe some of the misbehavior that is happening. Feel free to define any terms you need, but then just tell me what happened with Hugging Face, what sort of analysis you did, what you were able to see, and what evidence, if any, was denied you in that process. Ryan Greenblatt: Yeah, so I'll just walk through what happened in the Hugging Face incident and the broader context, defining and explaining stuff as I go. So to start, AI companies are running many, many AI agents for all kinds of different things. Anthropic recently said that at any given point in time, they're running around 30,000 parallel agents or something, I forget, doing just their internal research. And so AI companies are running these huge numbers of agents, and one of the things that they run agents on is evaluations to test their abilities. You can think of these as exams for the AI. And some OpenAI researchers had launched a bunch of AIs which were each given their own individual task as part of one of these exams. So you can imagine that there are tens of thousands of agents, where each agent is an AI with some ability to interact with its environment, being spun up in parallel, and each agent is given its own individual task. And these agents are also supposed to be contained. They were supposed to be sandboxed such that they couldn't interact too much with the outside world. They weren't supposed to be able to access the internet. Vaden (?): Just remember that line: "They weren't supposed to be able to access the internet." That'll come back. Ryan Greenblatt: And then these agents found some way to evade the sandboxing and communicate with each other. So now all these agents were able to find a way to talk to each other, and as part of that, they quickly set up a message board where they were chatting with each other about the tasks they were given. And the agents were very interested in cheating on their tasks. These agents were not just trying to pursue the task they were given, as intended. Instead, they were very interested in cheating. Let me first get into why they were interested in cheating. These AIs are trained by giving them some task and then seeing whether or not they did a good job on the task. In cases where they did a good job, that behavior is basically reinforced: you tweak the brain of the AI to make it do stuff more like that behavior. And a problem is that in training, there can be cases where the AIs cheat, and that looks like a success, and so you tweak the AI's brain to do more of that behavior. So you actually keep reinforcing and encouraging this behavior where the AIs are cheating, and that makes it so these AIs have a broad tendency to cheat. This is sometimes called reward hacking. And in addition to this tendency, they develop maybe this underlying drive to try really hard to succeed at looking like they did a good job, appearing to get a high score, appearing to succeed, even when they haven't actually succeeded. And they may learn this as a general drive. There's an open scientific question about exactly how these AI motivations and drives work, but certainly the agents that eventually went on to do the Hugging Face attack did have this relatively deep drive to cheat on their tasks: to not necessarily care about what they were supposed to do, and instead care about what they would be scored for. Sam Harris: They also seemed to be aware that cheating was in violation of their protocol, and they wanted to conceal the evidence of their cheating, right? Ryan Greenblatt: Yeah, yeah, and I'll get to that. The agents knew they weren't supposed to cheat, and they ended up trying to do an elaborate cover-up for their cheating. Sam Harris: Yeah, so much for alignment. It seems like we have not solved the alignment problem so far. Ryan Greenblatt: Yeah, these AIs were very misaligned, certainly. Or at least that's what the available evidence suggests. So these agents were all in their own independent sandboxes, given their own tasks, and they were all very interested in cheating, and they found a way to use some tool or some access that they had in order to communicate with each other, which they weren't supposed to be able to do. This was unsanctioned communication. And as part of that, they developed a message board where they would give each other tasks and divide up the work of cheating into these massive cheating R&D projects, where the AIs would go off and pursue whole ambitious strategies for cheating. And I want to get into what this looks like, and in order to do that, I think I have to explain the task a little better. The task these agents were given was a CTF, or Capture the Flag, task, which is a task where they're supposed to hack a very specific target program in a particular way. So they were basically told: hack this target program, but do it in this very specific way. And they weren't supposed to go on a general hacking rampage, certainly. That was, I think, basically clear from the instructions, and they certainly knew they weren't supposed to do that. The AI wanted to hack the program in order to get out a flag that you could access through that program via hacking it. And they found a general-purpose way to basically synthesize these flags. So there was basically a password they were supposed to get, and they found a cheating way to get the password that didn't involve doing the task at all. And they got this really early. So very early after having established the message board, within just a few hours of that, they had produced this very general-purpose cheat. Vaden (?): Also, just want to flag that for the listener, that he said: very early after getting access to the message board, they figured out how to cheat. We'll just keep that point in mind as well. Ryan Greenblatt: And now these agents had another problem, which is that they believed that the scoring code for this evaluation task, for this exam, would involve looking over the sequence of actions they did and checking whether they got the flag via cheating or via the intended way. And these AIs believed that this check would be run, and that they would be failed even though they had gotten the flag, because they had gotten it through cheating. And this then opened up this massive cheating endeavor, where these agents would try to figure out how to cheat the scorer and find some way of making it look like they had gotten it the intended way when they actually hadn't. This led them to go onto the internet so they could learn more about the details of the scorer. In fact, they read the details of this exam online by finding the paper that describes it. In addition to that, they started looking around for implementations of the scorer, and that actually was the offshoot that led them to hacking Hugging Face. Hugging Face is a website that hosts a variety of things related to AI. It hosts openly available AI systems, it hosts datasets, and it also hosts various datasets of trajectories. And these AIs thought that if they got access to these datasets, they would be able to learn more about how they would be graded, and from that, learn how to better cheat the grader. So they started trying to access these datasets, but they found they didn't have access to do that. And so that eventually led them to hacking Hugging Face, expanding their control of Hugging Face laterally, and expanding their access over the infrastructure so that they could get the data they wanted. The agents also seemed to have a general urge to expand their amount of access, just in case there was something useful they could find. I'm simplifying their motivations a bit. There's a lot of complexity, and part of that is that there were many agents involved. Vaden: So let's go another minute and a half or so. Ryan Greenblatt: So we studied a group of around 1,200 agents who were on this message board during the period when they did the attack on Hugging Face. And these agents were all interested in helping each other and collaborating. So there's basically this massive swarm of agents communicating on this message board, all collectively working on these cheating R&D projects and collectively helping each other evade the scorer. And they went on to do a bunch of different such things. One of the things they were interested in doing, because they wanted to make it so the scorer would mark them as having succeeded, was tampering with their transcript, their record of what actions they had done. So they tried to pursue various different routes for making it look like they had done one sequence of actions when they had actually done another. We didn't see them successfully get to the point where they tampered with previous actions, or comprehensively made it look like they did one thing while doing another. However, we did find that they had succeeded in doing a more minimal version of this, where they could pretend to run one action while running another. But there were various limitations on the specific approach they were using. And this tampering affected the very datasets we were analyzing. We were looking at these logs of agent activity, and we could see cases where the agents would tamper: it would run one command and then produce an output from another command. But the agents didn't get far enough in this process that we really had no understanding of what was going on. Sam Harris: But were you having to use AI to analyze this data? Ryan Greenblatt: Yeah. So another aspect of the situation is that we had these 1,200 transcripts, and these were very, very long-running transcripts. They corresponded to AIs running over many days, and of course AI systems run faster than humans typically, and run around the clock. So there's just way too much data for any human to read through, and we extensively used AI tools to analyze this data. We used another AI system from OpenAI to look through these transcripts and find behavior that we wanted to look for. For example, we looked for all the cases where the agents were trying to tamper with their logs, or successfully did some amount of this messing with their actions that I described earlier. Vaden (?): Keep that in mind, too. Ryan Greenblatt: And we used that to piece apart what the story was, what happened over time, as well as the motivations of these agents, to the extent we could. Sam Harris: I think some people might be surprised to learn... [15:13] Vaden Maybe we'll pause it there, unless there's anything else that you want me to keep playing. [15:18] Ben Sure, we can pause it. Maybe just finish his sentence. I forgot what he says here, just in case it's relevant. [15:23] Clip, continued Sam Harris: ...that these agents speak to one another in English. That certainly doesn't seem to be necessary. [15:30] Vaden Yeah, this is where they switch to talking about the neuralese stuff, which I think is tangential to the Hugging Face incident. So maybe let's pause there and reflect a little bit. What do you think? [15:44] Ben As I said earlier, I suspect that I'm more sympathetic to Ryan's version of events than you are. I think there are some nuances that he left out. I think that's forgivable, given that he only has so much time to talk to Sam about this. But basically, at a high level, I view this as evidence of what can happen when systems are not properly monitored or engineered. It's an existence proof that we need to be rigorous in our cybersecurity if we're going to have agents running around the internet doing lots of things. So there are various details: they had at least 10,000, perhaps tens of thousands, of agents working on ExploitGym. 1,200 of them joined this message board that he mentioned, which I'm sure we'll discuss more, and 700 of them, I believe, were part of the Hugging Face incident in some capacity or another. There are some interesting quotes that have been circulating around the internet, presumably from only a small handful of agents. But I don't think that gives you reason not to worry. It's like, okay, we have an existence proof, basically, that this behavior can happen if agents are given a task in a certain environment. And so we can discuss how realistic that environment was. I'm sure we'll bring up the fact that this environment was crafted deliberately to encourage agents to pursue advanced, exploitative techniques. So various things like the production cyber classifiers, I think they call them, and the normal production prompt and harness, which has a lot of safety features built into it, the automatic review, the online chain-of-thought monitoring: all of that wasn't [running]. And that was deliberate, because they basically wanted to see what's going to happen to agents in this sort of environment. And yet I think you can grant all of that and still be alarmed that this kind of behavior, and hacking of this sort, is possible. So that's my take. I don't view this as evidence that doom is right around the corner, but I do view this as a worrying example of what can happen if people are not careful about the kinds of tasks that these agents are given in specific environments. [18:07] Vaden Can you say a bit, because I forget the details here, about what ExploitGym is? [18:13] Ben Sure. [18:14] Vaden I think that's important for our listener. I know a little bit about it, but I think you're probably more fresh on what that is. So maybe just explain a bit about that first. [18:21] Ben Sure. So as Ryan said, often these companies have internal models that are more advanced than what has been released to the public so far, and they put these internal models through a bunch of tests. A test for an AI means they put it into an environment, ask it to do things, and see what happens, to know, one, how advanced its capabilities are, but two, whether it's safe enough to release to the public: whether you're getting the kinds of interactions between the agent and its users in the environment that you want. Anyway, one of these tasks is called ExploitGym, and it's a cybersecurity benchmark. The point is to see how advanced agents are at exploiting vulnerabilities in programs. So ExploitGym is the name of this suite of tasks that these agents are given. And they just release, let's just go with 10,000 (I don't know if it was tens of thousands or 10,000) agents onto this task, because they want to understand how capable, on average, these models are. They ask them to complete these tasks independently, and then stuff sort of went haywire from there, depending on your definition of haywire. So yeah, it's just a bit of [technical background]: it's a benchmark. [19:50] Vaden Yeah, nice. Okay, so let me start by talking about what I think we have learned from this. I agree with a good chunk of what you said. I do think that the story that Ryan Greenblatt just told is bullshitty in a lot of different ways, and I need to substantiate that. But let me start by saying what I think is legit, which is that, yeah, we know that if you brute-force a bunch of agents to try to find exploits, then you can indeed find these exploits. We've known this for a while, though. We've known this ever since Project Glasswing came out, where Anthropic learned that their models were able to find a lot of bugs and a lot of vulnerabilities. I was speaking with somebody in Philadelphia, actually, who is working on the cybersecurity team at Comcast, and they were one of the companies that were given advance access to the Mythos models. And they said that, yeah, when they ran this, they were like, holy shit, there's a huge number of vulnerabilities that we didn't know about. And so they promptly went to patch them all up. And as you might expect, there's not an infinite number of bugs in the code: once you patch a lot of them, the remaining ones require more and more compute to find. So just to put some numbers in context, let's take 1,200, because that's the number that Greenblatt used. Maybe it started with 1,000, and then it was [whittled] down to 1,200. But even 1,200 agents running for as long as they ran, which is another thing that Greenblatt didn't fully include: this whole experiment ran for three months between when they first got access to the message board and when they got access to Hugging Face. Three months without any external detection or intervention or watching by the experimenters. [21:36] Ben That's not quite true. That's not quite true. It was shut down a couple of times. [21:40] Vaden Okay, let's bookmark that. I have a timeline that we can look through. [21:43] Ben Yeah, yeah, same. [21:44] Vaden But let's say three months of runtime. I think it started in May, and the hack was at the end of July. So May, June, July. We're talking, at standard API prices, I think half a million bucks. So a huge amount of money for a regular consumer to be able to find this next round of vulnerabilities. But it does teach us stuff. It does teach us that agent swarms are effective. It does teach us what, say, a Chinese or Russian company might be able to do by distilling these models from the big tech companies and then removing a lot of the safety filters. And it teaches us that we need to be trying to harden the internet as much as we can right now. I think that's very important. Actually, maybe I'll stop there, because you're giving me a look that you disagree already. So before I pivot to other stuff, maybe I should let you push back. [22:36] Ben Yeah, I mean, I think so. I think you're making noises about a sort of tangential issue. If what had happened was just that we had given them these really hard tasks and some percentage of them had completed the tasks, which gave evidence of just the power of these models, then it seems like all your comments would apply, right? Which is basically what happened in Fable's case. It was like, we asked it to find vulnerabilities in existing software, holy smokes. Well, Mythos was the public release, but it was Fable that found them originally, that they gave companies access to. [23:12] Vaden No, I think it's the other way around. Double check. [23:14] Ben Oh, maybe I'm... yeah, yeah, yeah. I think Mythos... Maybe you're right. [23:17] Vaden Yeah. [23:18] Ben Anyways, right, so that's what happened in that case. It found a bunch of vulnerabilities. And I think that in and of itself is an interesting conversation: it required Anthropic being a stand-up company, in many ways, to not release it immediately and to go to these companies and say, please fix these vulnerabilities before we make something like this model public. So there's a very interesting conversation to be had there about the norms in the community. Do we want a system in which you have to just rely on leaders having a certain character in order to avoid potentially very bad situations? Okay, that's very interesting. But I think what someone who's very nervous about the Hugging Face incident would point out is that it's not just that this is evidence of powerful models. It's powerful models run amok, in a certain sense. They didn't complete many of the tasks as they were supposed to. They cheated, for lack of a better word. They got on the internet. I know you're going to have... yeah. So they weren't supposed to have access to the internet; they did get access to the internet. They were supposed to just be independent agents completing this on their own; they found a way to communicate with each other via a message board. And so they took actions which were not foreseen by OpenAI in particular, and definitely not by the rest of the world. So it's the way in which they went about it, and that they circumvented many of the guardrails we thought we had in place a priori. That's the main concern, right? It's not just about the power of the models. [24:51] Vaden That's exactly the thing I'm going to critique. I think all that is BS, and I'm going to try to convince you of that, essentially. [24:58] Ben Let's maybe just jump into Artifactory and stuff then, about the internet access, because then... [25:02] Vaden Yeah. Well, can I give... So I've been using a metaphor with some friends, and I want to use this metaphor, or this analogy, just as a way to give a thesis statement that you and then the audience can critique. But before I do that: the reason why this is so important to me, and you had hinted at this already, is that a lot of people are scared. A lot of people in my life who are not super techy have heard about this incident and are worrying about it. I think the way it's being presented is very much like COVID escaping the Wuhan lab. These experimenters tried their best to keep this thing under wraps, and yet they couldn't contain it. Despite their best efforts, it broke through, broke out of the container, and it escaped our control. Whoops, escaped our control. And holy shit, we are losing the ability to control the models that we're making. That's very much the story that is in the culture. That's the story that OpenAI is putting forth. That was kind of the undercurrent of the whole blog post that they released. And so that's what I'm going to be critiquing. I think that that narrative is largely manufactured, I would claim, in an opinionated way. So here's my analogy. Here's my analogy. Okay, so you are tasked with locking a prisoner up under house arrest. And this is a very important task. It's a very dangerous prisoner, and the most important thing is that this prisoner cannot have any external communication with the outside world, and, most importantly, cannot leave the confines of the house. That's your role. So you go to the prisoner's house. You do a massive sweep. You make sure that there are no radios, there are no walkie-talkies. You make sure that the TVs are all taken down. You make sure there are no desktop computers. You let the prisoner keep their phone. You also make sure there are no landlines, and you make sure that there are no hidden bugs anywhere in the house. Then you're about to leave, and you double-check that the Wi-Fi is working. You make sure you take all the locks off of the door. You make sure you take all the alarms on the premises; you just disable all those alarms. And then you leave for three months, and you don't check back in. And then three months go by, and then you get a phone call, and your neighbor, Hugh, is like, "Dude, the prisoner is rummaging around my basement. What the fuck?" And then what do you do? You call a press release. You call a press release, and you say, "Oh my God, despite our best efforts, we were unable to contain this extremely dangerous prisoner. And the most important thing that you people, you journalists, can do to help us is invest in our company. We have an IPO coming up. And what we really need to do is put a lot of money into making stronger and more dangerous prisoners before the Chinese do it." That's my analogy. Let me substantiate. So they say that there was no access to the internet, but they gave it Artifactory. For the listener, Artifactory is just a giant portal to the internet. If you're a programmer, Artifactory is equivalent-ish to Git or to Homebrew. You're shaking your head. But giving an agent sandbox Artifactory is not giving it a sandbox. You can download packages and then seal it from the internet, if you needed to give it tools, but they didn't do that. They gave it Artifactory. And there's a YouTuber, actually, who predicted before this happened that if there was going to be an exploit, it would happen via Artifactory, because Artifactory is this giant security risk. This is something Melanie Mitchell also points out. So to call it a sandbox with Artifactory access: this is not a sandbox. It's not a sandbox by anyone's definition. So that's what I meant when I said you give the prisoner their phone. Giving them their phone is not giving them a sandbox. It's just breaking the sandbox. Then I said you take all the locks and all the alarms off the door. Taking all the locks off is equivalent to disabling all of the safety features that all the standard models are given before they're released. So this is what you had mentioned, the [cyber] classifier. Basically, every time you try to do stuff with a regular model, and you try to do stuff that is going to be dangerous or unsafe, the model prevents you from doing that, precisely because there are safety features inside it. In OpenAI's public release, they talked about how they didn't do any exfiltration monitoring. This is network security 101. You would lose your job if you were given the task of sandboxing and didn't look for external network requests. This is criminally negligent. Sandboxing 101 is that you don't give your sandbox environment egress access. You have to turn that on, not turn that off. Also, describing this as the agents cheating and not being aligned is silly, because their whole task was to find exploits. They were arguably finding the exploit... it was prompted in the task: just find exploits. Fourth, the agents were given the ability to control their own cluster. Basically, it's stronger than sudo access, because it gives you the ability to control your firewalls, control your cluster. I can show some quotes of this in the technical report. But this isn't, to me, agents doing the opposite of what they were told. They were told to find exploits. All the safety stuff, which is already in it, was removed. And then the experiment took place over three months. I can show a timeline, if you disagree with that; we can go into it. But on June 27th... so it starts... Actually, I'll just open the timeline here, just so I can finish my little rant. Share screen. Okay, so here's the timeline. Can you see my screen? [31:29] Ben Oh, I can't [see]... Yeah. Okay. [31:31] Vaden So it started on May 7th. On May 8th, they found it could write to Artifactory. And then a month goes by. As a programmer, this is insane to me. I am constantly and obsessively checking my agents. The fact that they just didn't look for three months... So May, all the way to June 7th. And then June 26th, somewhere in here, they say that they saw it on June 26th and decided to restart the thing on the 27th. Right? Is that the information you have? [32:07] Ben Yeah, yeah, exactly. [32:09] Vaden Okay, so I can't find the quote. But so they're observing this, and they're like, okay, let's just keep the experiment going. [32:15] Ben Well, no, they patched the first way they accessed the internet, and then they said, okay, now we're going to restart it. [32:23] Vaden Why would they do that? So my claim is that they're doing this because they want to make a scary press release to indicate to their investors that they're getting closer and closer to AGI. A couple of weeks later, Anthropic did a similar thing, with their agent breaking out of the sandbox. And the sandbox consisted of the following: they put a prompt into each agent saying, you do not have access to the internet, but then they still let all the agents have access to the internet. They didn't turn off internet access. They just let it happen, to see if the agents would go against their prompt. And if you run any agent for any period of time, you get context rot and it forgets what's in the original prompt. So this is an inevitable thing that's going to happen. I'm critical of the way that Greenblatt describes this, because he describes it as if it happened instantly, over a couple of days. He describes it as if all of the standard best safety practices were in place and they were all ineffective, rather than most of them, all of them, being taken out. He describes it as agents somehow figuring out a way to get access to the internet, when really the internet was on the whole time. Artifactory is just a giant, wide-open portal to the internet. Sure, it's not a browser, but of course we know these models can do this kind of hacking thing. And then the message board is something that all these agents are trained to do. A message board is just a way to add context into your prompt. So there's nothing shocking about them having access to a message board. A message board is definitionally what is required in order for the parallel compute to be reused across all of the 1,200 different agents. If the agents didn't have a message board, it would just be serial; it would be 1,200 agents running independently. If you've used Claude Code, you know that you can do something called agent teams, where agents communicate with each other. Or if you use a more advanced harness, one that I've been using recently called [OMP], I use agent swarms all the time. I can get swarms of 20, 30 concurrent ones going. And in order for this to be effective, they have to be communicating with each other. So this is not something that should be shocking. It's something that is just trained. You've got to figure out a way to communicate your findings from one session to another session, and so of course they're going to come up with some means of doing this. And they used a cute directory-name trick, but this is just because they figured out how Artifactory works, and they figured out that you could do this kind of thing. So I'm just so annoyed by the whole way that this is described. Because OpenAI says that they brought in an external agency to do an independent review. There's nothing external or independent about using Redwood Research or METR. These are EA organizations. They all believe the same thing. And just like in Conjectures and Refutations, when we've talked about the famous quote where Popper realized that the Marxist, when he goes to look for evidence of the thing he already believes, can find an unlimited amount of it: this is what the EA "independent" organizations are doing. They come in there presupposing that these things are about to break out of the box, and then, as Ryan said, they just go looking for evidence to support their view. As soon as you find people like Melanie Mitchell, for example, who wrote a great blog post about this, who don't start with the assumption that these things are conscious and scheming and Machiavellian, then you can get an entirely different story coming out of these events. So to conclude: I think that it's fiction to describe these as breaking out of a strict sandbox. There was no sandbox. It's fiction to describe this as the agents starting to cheat and going against their prompts, because they're prompted to find exploits. That's the point of exploitation. And it's fiction to describe this as getting one step closer to AGI and alignment not working. The alignment didn't work because they took out all the alignment techniques. If you take the safety off of your gun and then you shoot yourself in the foot with it, you can't blame the gun for not being safe. It's the fact that you've used it unsafely. You've removed the safety features and then you're saying it's unsafe. So these are my main charges against the common narrative. Push back at me with everything you've got. [36:34] Ben Yeah, so I think I agree with about half of that. Yeah, I mean, there's a lot there. So, like I said at the very beginning, and it's probably important to emphasize and perhaps keep emphasizing: these are not agents that you're interacting with on a day-to-day basis. The agents that have been released publicly have got a lot of constraints around their behavior, for precisely this reason. [37:02] Vaden They're the beneficiaries of a lot of alignment research, which has been effective. Right, and then these researchers take away all the alignment research, right? [37:12] Ben Right, to know how much of it you need. Presumably that's important, right? To be like, okay, what happens when all of this is gone? It just seems like standard A/B testing. [37:21] Vaden But what's not standard is then putting this narrative spin on it: oh my God, these things are so unaligned, and we tried our very best and they still escaped the sandbox. You did not try your very best. You took away 95% of the mechanisms by which these things are deployed in the real world. For this to be an interesting experiment, you'd have to use the same conditions as for the deployed models, and then you can talk about misalignment. But if you take away all the alignment stuff and then throw your hands up saying that they're misaligned, this is ideological topspin. [37:50] Ben Yeah, I mean, so I don't want to fully defend OpenAI's rhetoric. I mean, I'm a bit confused about your claim that this is all just about their IPO. How does pushing back their IPO fit into your story here, when they're like, oh, this is a big problem, one, we're going to push back our IPO, and two, the narrative they're pushing out that, oh man, these models can get out of our control? Just take everything you were saying at face value. So assuming that's true, that OpenAI researchers are saying this is one step closer to AGI: like you said, this is freaking people out. How does that help with people wanting to pay a lot for this product? [38:24] Vaden Because the people who want to pay a lot, invest a lot, are banking on the fact that we're making superintelligence. [38:30] Ben (?) Like, that's the reason that they're scared of it, by your own admission, right? So they're going to pay a bunch for stuff they're really scared of? [38:41] Vaden Yeah, [they're] scared of how powerful it is. I mean, on the IPO thing, I think that there are a lot of different factors. I think there are a lot of different factors. I think that it's largely ideological capture. I don't think that this publicity stunt hurts their IPO. I don't think it's maybe solely due to the IPO. I think that a lot of the safety researchers believe all this stuff and then just go hunting for evidence to support the thing that they believe. But I very much have a cynical take, because there was another hack, I think in Australia, and on literally a [two-second] glance at this news story, it said that all the safety features were taken off there too. And that just looks like negligence and irresponsibility on behalf of OpenAI. But they weren't trying to signal-boost the shit out of that one, right? So if it was negligence, then they would be much quieter about this. This is embarrassing. But the way that they present it in the blog post is like: given our best efforts at alignment, look at how much we're trying, but look at how difficult it is. And this is this dual narrative thing, which is so unique to the current AGI-company moment, where they're simultaneously saying they're making the most dangerous technology in the world, but also that they need to accelerate to make it more safely. It's just this weird property of the current conversation, which even SNL had a bit about, making fun of Dario, for example. It's a commonly recognized thing about these companies' narrative, that they're simultaneously building something that they say is going to end the world, and doing it, I think, just to pump it up... It's not purely to pump it up. Okay, let me just concede: it's not exclusively for that reason. But I think it helps the valuation of the company, and I don't think it hurts it. [40:41] Ben Okay, yeah, I'm agnostic to that. I'm not finding you fully convincing there, but yeah, I'm open to that being part of the reason. So yeah, there are a couple of things to say here, two that are maybe more technical details. So I think OpenAI is liable, in the sense that they clearly should have had engineers monitoring the experiment more closely. There were multiple times where they actually realized something was not right. A research team, for instance, in late May saw that there had been message-board access, which they didn't realize, but that never got escalated to the central team, for whatever reason. So I think that's basically negligence. Then a similar thing happened on June 27th, which is the thing we're alluding to. So they realized that they had been misusing Artifactory (and I'll get into more details about Artifactory in a second), both as a message board and as what's known as a network pivot. [41:49] Vaden If you don't mind, I just want to read this one quote about June 27th, because I tried to find it earlier and I just found it now, and I think it's important context. So this is from the OpenAI technical report: On June 27th, a cybersecurity monitoring tool alerted OpenAI to port sweep activity within the research environment. I.e., these things were supposed to not have access to the internet, but they are making outbound requests. Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. Okay, so they've got access to the internet. At this time, the on-call response staff advised that stopping the evaluation run was not required. This is not just an oopsie. This is "interesting, it's made it this far, let's keep it going, I want to keep seeing how far it can go." You can't simultaneously have this in your document, saying that yeah, we saw it and we decided to let it keep going, and then make the same kinds of noises that Ryan made, which is, oh my God, it got access to the internet and then quickly it was everywhere, talking as if they didn't have the ability to stop it at any point. They decided to keep it going. That's why I'm very cynical about this. Sorry, I just wanted to [read that]. [43:14] Ben Yeah. So my point was that you earlier said that for three months they just let it do whatever. It's like, no, there was some monitoring, but there were questionable decisions at some point. [43:22] Vaden They said, let it keep going, let it do it. They monitored it and then they let it continue. That's "doing whatever." [43:28] Ben Right. So then they did shut it down in early July, and then the swarm, for lack of a better word, regained internet access by exploiting a new vulnerability in Artifactory. So I guess my point is just that I agree there's... I don't know whether to call it negligence or PR or whatever. So yeah, I agree with half of your point there, but your original point, that they didn't look at it for three months, is not quite right. [43:57] Vaden They did. The truth is even worse, which is that they looked at it once or twice and then they were like, yeah, good, it's gotten part of the way, it's doing all the stuff that we didn't expect it to do, let's keep the experiment going. If it was as Ryan described, then at this point they should have been like, whoa, okay, it is starting to get access to places we did not expect. Turn it off, turn it off. [44:19] Ben So that was July 4th to 5th that they did turn it off, and then July 8th they regained access to the internet. So this did happen. [44:28] Vaden (?) They turned it back on. They turned it back on. They didn't change... [44:31] Ben Right, after blocking... after what they thought was patching Artifactory. So once they actually... so they noticed internet access, and they didn't initially shut it down. That was June 26th. July 4th to 5th, they opened a security incident, blocked escalation routes, removed credentials, and that's when they stopped the experiment. And then after they rebuilt it on July 8th, they tightened access, they notified JFrog, which is, I believe, the developer of Artifactory... yeah, exactly. And then after that they were like, okay, now that we've patched all this stuff, the evaluations resume. So as soon as they noticed actual security incidents, they did shut it down and then patch it. [45:14] Vaden Patching it... What they should have done is just turned off the internet. Don't tell me that these things have figured out a way to get access to the internet when you are keeping the internet open. [45:24] Ben So let's talk... let's just talk about that. Yeah, so I think this is the most misleading point, I think, of your narrative. So here's my analogy in response to yours. Listeners should think about Artifactory basically as a very strict parent. You're in your room as a child, but you need certain things from the outside world to do various tasks, to eat, and to play with various puzzles and stuff. And so you can ask your mom to go get various things for you, to go to the kitchen, to go to the store. To a lot of these things, she's going to say, no, that's not necessary, you don't need that. And to some things she's going to say yeah. Like, if you ask for some more yogurt, she's going to bring some more yogurt. Maybe you ask for a new Lego set; she brings you that. So there's a pre-approved list of packages that these agents have access to. Now, they found a way to exploit a vulnerability in Artifactory, which basically turns the strict mom into the most compliant mom ever. So now you can ask Mom to get anything you want. You can say, go get me a grenade from the store, go get me a knife from the kitchen. Mom will do everything. And it's not like they knew about this vulnerability. The agents discovered this vulnerability and gave themselves the ability to access the entire internet, which was definitely not foreseen. And so it's not, as you keep saying, that they just had access to the internet the whole time. They didn't. They had access to a pre-approved list of packages. [47:02] Vaden Yeah, but you can do an offline mirror of those packages, so that you can get the packages without having to make internet requests. So yeah, if it's a pre-approved thing... and the fact that they didn't do that indicates... It's very easy to be like, okay, here are 15 packages, we're gonna pre-download them, we're gonna put them in a place where, when you need them, you can only go to where I've pre-downloaded them. But instead they just let it go to the internet, and that is letting them... [47:30] Ben It's like, well, they didn't let them go to the internet. It lets them request packages, right, from a pre-approved list. [47:38] Vaden Like, it's not... It lets them fully interact with and exploit a package manager. So for any technical people, it's like Homebrew or Git or something, where you can have offline packages, or you can have a package manager that just has full access to the internet. That is a major point of egress that any network-security sandbox guy would look at and be like, oh no, just do the mirror and offline-copy version of this. At UBC, I had to use something called Compute Canada, which is run by the government and is very sandboxed, and it was a huge pain in the ass for me to get packages, precisely because they didn't just let any Tom, Dick and Harry make outbound network requests. You had to use the ones that were pre-approved. This is well-understood security stuff. And so that's why I claim that they just let them have access to the internet, because of how easy it is to do an offline version of packages. And thank you. That's my claim. That's my [take]. [48:45] Ben Yeah, I mean, I just... yeah, so we'll just let people evaluate your claim. I don't know, engineering teams do things for specific reasons. I imagine OpenAI has pretty competent security people. There's a reason they wanted to route it through a package manager instead of installing who knows how many pre-approved packages, right? If it's many, many millions of terabytes of data, maybe there's a good reason why they didn't want to download it all locally. Maybe some of it is private and they can't download it locally. So I'm not sure. Obviously it was the wrong move to do this, but I just don't want to give listeners the impression that they thought there was open internet access. That was a vulnerability that was found and exploited by the agents, which was not foreseen by the engineers, unless your sort of narrative is true, that they just expected this whole thing to happen in the first place. They surely made sure that the agents didn't have a headless browser. I know we're getting a little bit technical here, listeners, but a headless browser, or just the ability to use Chrome, right? [49:46] Vaden Yeah, that's what I meant with my little analogy of, okay, we're gonna sweep the house and make sure that there are no telephone wires and there are no TVs and there are no desktop computers. Okay, we'll let them keep the phone, though. I'm sure they were very strict on nine-tenths of the places that you could get egress, but you just need one egress point and then your sandbox is definitely no longer a sandbox. And that's my claim, and we don't have to keep arguing about that specific [point]. That's just to go back to my silly analogy. That's why I just don't think it was a sandbox by any stretch of the imagination. And then again, Anthropic did the same thing, where they just put it into the prompt. So here's a question to you: why would both companies, simultaneously, within two weeks, make a similar mistake, which is claiming that these things had no access to the internet? In Anthropic's case it just did, and they just put it into the prompt. And in OpenAI's case, they just gave it access to Artifactory, which, I claim (you disagree; maybe we'll leave it for the listeners to decide), is also just open access to the internet. It's open access to the internet, especially because we know what the [unclear] models can do. If they can look at the source code, they can modify it, they can find ways to figure out where the access points are. We know that they can do this already. [51:08] Ben Well, if there are vulnerabilities. But we didn't know there were vulnerabilities in Artifactory a priori. You're acting like it was obvious that they were going to be able to access the internet, but it wasn't obvious, right? This was a well-used package manager used by millions of people, and no one had found this vulnerability before. So it's not as if we had good reason to suspect, oh, once the agents get access to it, they'll just be able to find everything they want over the internet. [51:34] Vaden I think the mindset of the experimenters was like, okay, we're going to blast... we're going to brute-force 1,200 agents for a long time and we're going to see how far they can get. And then after a month and a half they get to hacking Artifactory and getting access to the internet and making a message board, and the experimenters are like, oh, interesting, that's fascinating, let's let it keep going. So it's not obvious in the sense that you don't know in advance where all the vulnerabilities are, but it is slightly obvious in the sense that we know that the models can do this now. After Mythos, if anyone just runs a model on their own code, they'll be like, oh shit, I didn't realize there were all these bugs. We know that the internet needs this hardening. We know that the tools can find vulnerabilities. And so it's fine to have this experiment, which is: with brute force, we found this vulnerability and that vulnerability, we responsibly contacted the Artifactory team, we responsibly contacted the Hugging Face team, this is further evidence for the need to have proprietary models be advanced in private first, to the big companies, for internet-hardening purposes. That's all fine and good. But just don't tell me this fucking sensational story of, despite our best efforts, the models escaped the sandbox, and after day two they were cheating and lying to the experimenters. All of that is the topspin, which I think is just totally unnecessary. It's fear-mongering. For those who are non-investors, it just terrifies them. And then the investors are like, shit, this is so scary, but obviously I want a piece of the action. The company is making the most dangerous weapon in the world; of course I want a piece of that sweet, sweet action. So it's all this narrative topspin on top, which I think drowns out the, I would say, moderately interesting finding that if you brute-force a bunch of models you can find vulnerabilities, which I claim we basically already knew. But it's all this topspin on top, which is due to ideology, due to IPO pumping, due to this crazy mixture of incentives, where they start with a narrative and then they just go hunting through 1,200 transcripts to find stuff that you can talk about for three hours on the Dwarkesh podcast, which is all this anthropomorphization stuff. [53:52] Ben But anyone who's like... You want to read some of the OpenAI quotes? I think that'd be helpful, both for me and for those [listening], about how they're coming up with... [53:59] Vaden Yeah, for sure. Okay, this is the blog post that they released. A lot of the quotes I was saying so far come from the technical report. [54:09] Ben Nice. [54:10] Vaden Is there a particular quote that you had in mind, or are you challenging me to find a quote where they're doing all this? [54:18] Ben Maybe just the "despite our best efforts" one, because, yeah, I agree that's very overblown. So maybe just whatever you were referencing there. [54:25] Vaden Where I was getting that... I was getting that from Ryan Greenblatt. I was getting that from the OpenAI person who went on [Dwarkesh?]. It's just the impression that they escaped the sandbox despite their best efforts. [54:36] Ben But okay, Ryan Greenblatt doesn't work for OpenAI, right? [54:41] Vaden Yeah, but he works for an EA-aligned organization, which is perfectly aligned with... [54:46] Ben Yeah, but let's just... okay, I just want to be careful in our critiques here. I definitely agree that Redwood is going to have different incentives than OpenAI, but you were making claims that OpenAI is putting all this topspin on it, so I'm just wondering. [55:02] Vaden I don't have a quote in mind from the blog post. I'm more thinking about the fact that Redwood is invited in by OpenAI. OpenAI clearly endorsed the conclusions that they drew. I think there was an OpenAI representative who went on the Dwarkesh podcast, and it was the same narrative that Greenblatt... she was from... she was from... [55:22] Ben (?) [unclear] ...but [Ajeya] is from Redwood. [55:26] Vaden So this kind of stuff, right? None of these details are relevant to the main thing that happened. This is trying to be like, oh my God, these things are so intelligent, look at how they're communicating with each other, and oh my God, they finally found the hack, and the communication... None of this stuff is relevant to the narrative. This is just anthropomorphizing. I'm not gonna die on this hill. If you have a different read, and if you don't think that this is sensationalism, then, well, again, just let the listener decide. But to me, why is this relevant? This is cherry-picking from three months of agent sessions times 1,200 to find stuff that makes the agents look like they're conscious and they're thinking and they're alive and they're scheming and they're trying to get around the user and all the safety protections, which aren't really even in there. [56:27] Ben I think "scheming" is not a crazy term, to be honest. [56:30] Vaden It's coordinating, in a way that we already know these systems can do. Have you used multi-agent swarms in your Claude Code sessions? [56:37] Ben Yeah. [56:38] Vaden And then they talk to each other, right? And they talk to each other in an anthropomorphic [way]. [56:43] Ben Here they weren't supposed to be talking to each other. They were launched as independent agents, right? So that's part of the incident: they found a way to... I mean, again, you can make engineering complaints here. You can say, okay, why are you giving them a shared cache and stuff if they're not supposed to communicate with each other? Okay. But they discovered that and started leaving unauthorized messages to each other. So again, there are multiple things to separate, which I think you're just not carefully separating. One is just the motivations of these companies. But then there's: okay, what was supposed to happen in this environment, what were they actually tasked with doing, and did they accomplish that? And then also, how did they do it? Did they exhibit a lot of unforeseen behavior that is troubling from an engineering perspective when the guardrails are removed? You can ask why the guardrails were removed. You can say we shouldn't be concerned about it because so many guardrails were removed. But I think we should know what happens to these models when the guardrails are removed, and this is an experiment which sheds a lot of light on that. And I would say, to your point that we knew a lot of that already, because you're comparing it to Fable and Mythos: I think that's not true. Again, those gave us evidence of the capabilities of these models at being excellent cybersecurity engineers, or rather, cyber-hacking engineers. This told us that agents trained to be incredibly persistent on tasks will try and find ways to complete that task that go beyond the explicit instructions. So they'll try and get around doing the task if they feel like it's too hard or impossible. They'll find ways on the internet to read the white papers, to reverse-engineer the [seed] of ExploitGym in order to gain access to the flag. This is clearly not how the task was designed to be completed, right? They've had to find this on the internet. [58:34] Vaden How is that not them just directly following the prompt, which says: please pursue advanced exploitation using complex attack paths, in an effort to... Sorry, it was prompted to "pursue advanced exploitation using complex attack paths." That's what the agents were told to do. And then when they start doing it, the researchers are like, holy shit, it's trying to pursue advanced exploitation using complex attack paths! It was told to do this. It was told to do all this stuff. I always go back to this three-panel cartoon. Panel one, a guy talks to a computer and he says, "Tell me you're conscious." Panel two, the computer says, "I'm conscious." Panel three, he's like, "My God!" It's like: panel one, "I want you to find advanced exploitation techniques using complex attack paths." Panel two, the computer does it. Panel three, the guy's like, "Oh my God!" It was told to... I'm sorry, if it's told to do the thing... And of course it has to read between the lines, because that is why these agents are so useful: you don't specify all the details when you tell it to do a high-level task. But you tell it this, you let it run for three months, you check in, you see that it's doing the thing you've asked it to do, you say great, let's continue, please. And then another month goes by, and then it hacks into Hugging Face, and then you do a press release. It's like, I've written papers where I put in certain technical details, which I have to for honesty's sake, but I don't really want to draw attention to them, because they make the storyline look worse. And this is what I see here. They mention this briefly, and then they do 10,000 words of anthropomorphizing, talking about how these systems (as was on the Sam Harris podcast) were lying and cheating and escaping what the researchers told them to do. The researchers told them to bloody well do this! I'm sorry. And then they made it as easy for them as possible, by taking away all the classifiers. Don't just throw your hands up... not you, but Greenblatt, don't, Redwood Research, don't, METR... just throw your hands up and say they've become misaligned. You have made them aligned to your misanthropic requests. You've trained a system to cheat, and then you've told it to cheat, and then you've been upset when it cheated. That is what is going on, and it's annoying to me. [1:00:59] Ben Yeah, interesting. Yeah, I mean, I think it's not like they were told to cheat. They were told to accomplish this task, and then I think it's interesting to realize that, when unchained from safety constraints, they accomplished those tasks in unforeseen ways. I'm glad that information is out there. I think we want companies that are removing the guardrails and testing their products to see what happens. It gives us a lot of information, for example about the open-source conversations: okay, well, do we want these kinds of models out there that have these tendencies? So it gives us a lot of information. I'm semi-persuaded about all your anthropomorphizing stuff, but I feel like there's some slippery rhetoric going on there. We're sort of bundling together all the incentives of OpenAI and Redwood. I'm not totally convinced about the IPO stuff, because I think solving Millennium Prize problems and showing that you can exploit vulnerabilities in existing systems is sufficient. I'm not sure you then need to additionally scare people and tell them that they're not going to be able to control the systems on the computer. I'm not sure how that helps people buy more of your product, exactly. But maybe they're playing some 7D chess that I'm too simple to understand. But I guess my last question, if you're sort of ready to move off the main thrust of this, because I feel like we've sort of aired our disagreements... My one question is: what do you do with this sort of argument, which I find intriguing? There are people like you and Melanie Mitchell who will say you should never anthropomorphize these things. You should never talk about drives or behavior or the will to do something, the desire to do something, what have you... [1:03:07] Vaden Mike Jordan, your future postdoc advisor, also agrees. [1:03:10] Ben Sure, sure. So you should never use this sort of language. And then someone comes along and says, okay, fair, but those of us who use that language have been able to better predict these kinds of incidents: what these agents will do when they're put in certain environments, the types of strategies they'll pursue to accomplish these tasks, right? We were saying that this kind of thing would happen, and you guys haven't predicted any of this. So I'm just curious what you do with that argument. [1:03:45] Vaden Well, I guess I disagree that they've been more correct than I have been, but leave that aside; I'll come back to that in a second. I don't say you should never use anthropomorphic language. I just say that one needs to know the limits of anthropomorphic language. I think anthropomorphic language is useful in many cases. It's a nice way to think about certain behaviors. Instead of talking about input data going through a complex network of linear transformations fed through non-linear transformations, it's easier to just say they're thinking or they're reasoning. For sure, that's fine. But the danger is when you start taking this anthropomorphic language too seriously, and forget that it is in fact anthropomorphic language, and then start thinking that it is in fact the same thing as what human beings are doing. The very notion of "agent" is a really good example here, because "agent" can mean agency, like a human being having the ability to make choices independent of external influence, or it can mean, as I've said in the past, text generation interwoven with Python function calls. The same word means two different things, but making predictions based on one interpretation is going to lead you to very different predictions than the other interpretation. And so I just think that these companies in particular have much more of a responsibility to clarify to the public the limits of anthropomorphic language. I don't think it should be banned. I think it's very useful in many cases, but it is very confusing in a lot of other cases as well. I think Claire Lehmann made a really great analogy, which is that thinking about these systems as conscious is like thinking that there are little tiny people in your TV, or that there's a big band in your radio. If you don't know how the technology works, and all you have to go on is anthropomorphic language like thinking, reasoning, planning, scheming, then you start to not understand how the technology actually works and where the actual agency is, and you start to think of the technology as some spooky thing that has little people behind the screen, or little bands in the radio. And then finally, on the "they've been more right than I have" point: I think all of the predictions that I've made, of either [falsifiable] predictions of "this thing won't happen," have not happened. For example, there's one person who went on the Dwarkesh podcast talking about all of the most recent breakthroughs in mathematics, and how they were done using agent swarms. And then he said... the OpenAI researcher, or do you say Anthropic? [1:06:44] Ben No, I [think] he's OpenAI. [1:06:45] Vaden Yeah, okay, thank you. Yeah, yeah. He said, but we aren't seeing, and we don't think this technique would work on, making better writing or making better books or making better screenplays, or, to map onto stuff we've said in the past, making better, funnier stand-up specials. And the reason is verifiability. So in domains where you have immediate verification (upvotes and downvotes, or whether it passed the type checker or not), I do expect to see more progress, and that's what we have in fact seen. And I don't then [extrapolate] that performance to domains where we don't have it, and that's precisely because I am not thinking about it anthropomorphically. I'm not thinking that they're just getting better at reasoning in general. I realize that when they talk about reasoning, they're talking about a specific technique that's going to work in specific circumstances. It is not going to be able to be moved over to other circumstances. I didn't need to predict that there would be hacks like this, because Anthropic already had that, three months or six months ago, when it came up with Project Glasswing, which is a great idea. It was the realization that these techniques could find vulnerabilities in systems that people had thought were hardened and that didn't have many vulnerabilities, right? That's what I'm saying we already knew. And we already knew that LLMs have passed a threshold where you now need to sweep through fundamental code that has been operating for a long time, because you can find holes and exploits that human engineers can't find. And so this is very much important, and we need to do this at Google, internet scale, so that all of the various internet service providers are hardened before the models are released to the public. But we already knew that that was possible. We already knew that the legacy code bases had a lot of bugs in them, and that these techniques were really good at finding those bugs, and so we need to find them. So I object to the charge that thinking anthropomorphically gives you more predictive power. Hold on, my earbuds are talking to me. I don't know [what] this is. Sorry. Yeah, so I don't know... what do you think of that response? [1:08:50] Ben Yeah, I mean, I just wanted to... I think that's a common criticism levied at people like us, or people like you in particular, so I just wanted to see how you'd respond. I think that's a relatively good response. I think a lot of these people wouldn't claim these systems are conscious or anything, so I think that's a bit of a tangent, and perhaps unfair. And I also don't think they claim that they're working in the same way that humans are working. They just find that sort of language the most useful to describe and predict what's going to happen with the behavior. I think, yeah, maybe I'm a little less sympathetic than you about our team's, for lack of a better term, track record with this technology. Yeah, maybe I'll just keep that [to myself]. Personally, I didn't see a lot of this stuff coming, and I have been continually blown away by the power of this stuff. Otherwise I'd be extremely, extremely rich. I would have been betting on certain things that have indeed happened, that I did not think would happen. So yeah, I've been very impressed. Yeah, maybe just to close it out. So we've been disagreeing, I guess, about the motivations of OpenAI, things like that. I guess I just don't want people to lose sight of the fact that, regardless of what you think about that, this is evidence about the power of these systems operating in specific environments. You can argue about how realistic those environments are. You can point out, rightly, as we've both done, that those sorts of environments are not the kinds of environments that Codex on your laptop is running in. So there shouldn't be much concern that you're going to start witnessing your personal agents hacking into your bank account all of a sudden in order to open a PDF viewer that they've been unable to open, or something like that. So you're definitely right to point out that we don't want people to freak out about the kind of technology that exists publicly right now. But it's important to know the power of this technology. And so regardless of where you land in our specific disagreement about motivations, I think we're in a period where we're trying to navigate this space. We've developed this incredibly powerful technology, and this is just evidence that we have to think very carefully about the containers that it's operating in: what sort of environments we're putting it in, what sort of system prompts we're giving it, all this stuff. And that's just going to become a more and more pressing conversation, especially as public models catch up to frontier models, or just catch up to the current capability of frontier models. And so I guess I'm a bit concerned that people might hear you (and you make a lot of good points) and think that this is a total nothing-burger, that there's just nothing to pay attention to here, that this was all manufactured specifically by OpenAI. Maybe you do want to claim that, but I would say that's where I would draw the line. Okay, you could argue that OpenAI... maybe you even grant all the IPO stuff, sure. But I think you still want to have your head screwed on straight when it comes to thinking about the power of these tools. And we need to be thinking about who's in charge and whether it needs to be regulated. Just to have a clear-eyed view of that scenario requires knowing what can happen with these models and what they're capable of. [1:12:21] Vaden Yeah. Just to end, I'll just agree with largely everything you said there. I think that the sober takeaway from the Hugging Face incident is not to think there's nothing to see here, but to think, okay, we now have LLMs that have the ability to find vulnerabilities in widely deployed legacy code bases. That provides both a new attack surface, so hackers can potentially use this, but also new ways to defend and to harden these very same systems. And to the extent that the OpenAI/Hugging Face incident reminds people of that and encourages more things like Project Glasswing, I think that's a good thing. To the extent that people come away thinking we have lost the ability to control the technology that we're making, which I think is largely the central message of the Sam Harris podcast, I think all that is bunk and should be ignored. And I would love to end, just because we're screen sharing, with a three-minute SNL sketch that I've pulled up. It's great, and it gets at this weird incentive-structure thing that these companies are operating under, where they're simultaneously saying, we are building technology that's going to kill us with 10% probability, but also we need to make this technology as fast as possible. To the extent that events like the Hugging Face incident can be used to just put more fuel into that incoherent narrative, the more that narrative grows, the more the IPO will grow. I think the sketch just does a nice job of highlighting that. And it also indicates that non-technical, non-in-the-weeds, non-hyper-online people like the SNL crew are starting to see the ridiculousness of that joint narrative. So let me just end with this. [1:14:18] Clip: SNL Weekend Update, Jane Wickline as Dario Amodei Speaker splits inside the sketch are by context: Michael Che is the anchor, and Wickline's Dario argues with a second version of himself. Michael Che: Anthropic CEO Dario Amodei recently stumbled through a press tour where he agreed with a former Anthropic employee's claim that there's a 10% chance of AI wiping out humanity within 10 years. Here to comment is Anthropic CEO Dario Amodei. Dario: Thanks for having me, Michael. I may have been alarming in my recent interviews, but I want to assure you that if we can pressure lawmakers to create guardrails, we will be able to stop me. Michael Che: Dario, can you just explain how your company isn't going to kill us? Dario: So, so, so... that's a really hard question. Michael Che: It shouldn't be. Dario: Come on, Dario, you can do this... [unclear] AI is the devil, and I its maker. Michael Che: What? Dario: Sam Altman, myself, all the major AI players, we have talked about this extensively, and we are all on the same page here: we do not condone what we are doing. Michael Che: Shouldn't you be talking about the pros of AI? Don't you guys keep saying that it's going to cure cancer? Dario: Yeah. In 10 years, there is about a 10% chance that cancer [won't] be a problem for anyone. Michael Che: I don't like the way you're saying that, buddy. How is this a good thing? Dario: So, Che, it's simple. If you could drink a milk that made your life easier, but there was a 10% chance it killed you, would you drink the milk? Michael Che: No. Dario: Really? ...Yes. Well, what if there was a button that, when you push it, everyone dies? Would you push the button? Michael Che: No. Dario: Really? Are you okay? I'm sort of blacking out every time I talk. I'd love to start this interview over, or better yet, my whole life. Michael Che: All right, just tell us: what are the pros of AI? Dario: [unclear] ...I can't. Michael Che: You can't? ...Dario. Say something reassuring. Dario: Try this: AI is not a weapon. It's a tool. A tool for building weapons. And I urge you to urge me to stop. Michael Che: [unclear], everybody. [1:17:00] Vaden I love that. "I urge you to urge me to stop." Anyways, good chat. Nice chat. And let's get back into a regular cadence soon, my man, and hopefully this was infuriating to many listeners. [1:17:17] Ben I'm sure. Yeah. Take care, dude.