---
title: "Transcript: Zpk5Pwx2Yrm"
category: "transcripts"
videoId: "ZpK5PWX2YRM"
sourceLabels: ["YouTube transcript", "Cached transcript markdown"]
wordCount: "3931"
---

# Transcript: Zpk5Pwx2Yrm

## Source Video
- [YouTube](https://www.youtube.com/watch?v=ZpK5PWX2YRM)

## Local Cache
- `raw/sources/youtube-transcripts/ZpK5PWX2YRM.txt`
- 3,931 words

## Transcript

[music] two talks at a engineer Europe. One guy is saying code is free and deleted his ID and the other one is saying read every effing line of code. So should AI engineers still recode their agents output in 2026? I named this the Zopo continue and you guys probably have argued about this in Slack. You probably talked about this in the hallway track. So let's talk about this here because code got cheap attention didn't. As you may know back in December 2025, something big changed. AI engineering has changed forever and it broke its own trend line. Actually, Swigs, the organizer of AI engineer, is collecting evidence to that single moment in time at the website called wtfappen2025.com. I recommend you go and check it out. It's really, really funny. Uh, this is just one example from METR, the machine evaluation center, and it shows that models for the first time started completing tasks that would take engineers over 16 hours to do. And in fact, we've gone way up the curve, way up the trend line after that. This is the backdrop to everything that AI engineering is experiencing because we don't write code anymore. Most of us at least. I want to see one. Can you guys give me a raise of hands if you still handcraft and write code? Most of your code. Anybody here most of your code is written by hand? Amazing. This is the token maxing track after all. I I think the one person here who still writes code is maybe a little shy of raising their hand. That's okay because we don't type code anymore. We're not handcfters. We supervise. I like to say we're babysit agents. And the greatest example for this obviously is Boris Churn. You guys know Boris, the creator of Cloud Code, uh Adam Tropic. 100% of his code is written and authored by Cloud Code at this point. And he didn't stop being engineer. He moved up the layer. He still ships 20 to 30 PRs, maybe more. And recently he talked about he deleted his ID. I found it really funny. Just just no reason to just hand type code anymore. In fact 80% of entropics code is now AI written. And this is this stat is at least a few months old. It's likely more right now. And he's not the only one. Some of you have seen this chart from GitHub. Some of you have maybe remembered this chart while GitHub was down for you. Uh the reason is GitHub is on track to to to get 14 billion commits this year. All of 2025, all of yesterday last year was 1 billion. They're 14xing the number of commits. They're seeing 14x the number of commits, which is insane. And most of this is AI assisted. And it's a lot of code. And so the engineering has changed forever. And I want to tell you about AI engineer worlds fair. I been to every single one and I'll tell you about this later and a engineer is a great place to get the side guys of where our career is going and how is it changing. Okay, this one obviously is 3x bigger than last year. This is just one of the rooms. There's like a bunch of rooms 7,000 people I think we clocked in 36 tracks. And if you want to know what happens in AI engineering, you kind of have to be here. So this would be a little bit of a meta talk. So, one of the guys at AI Engineer EU talked about code the sheep. The other one talked about uh we should read every line of code. Let's listen to them for just a second. Okay, this is Ryan Lopa from OpenAI. I don't think he's he made it here, but this is Ryan Lopa from OpenAI. The models at this point are good enough where they're code at high quality that solve real user problems in real code bases. Code is free. It's free to produce, free to refactor, and it is not a thing to get hung up on anymore. Humans no longer need to concern themselves with implementation. The important thing is not the code, but the prompt and the guardrails that got you there. You can just simply say, "Do not produce slop." Don't accept slop. You won't get slop in your codebase. But to do that requires taking short-term velocity hits in order to back up or double click into a task to figure out what it is the agents are struggling with. So this is round of bubble. Okay. He came up on stage at the engineer and he opened with like hey I'm a talking billionaire and I want you to be as well. In fact the talking billionaire lounge that's in front of the leadership track that you guys see that's because of him. He came up with this concept uh and he got the golden card and everything. Uh, on the other side, the same conference, the other side, Mario Zetchner, creator of Pi. Slow the down. Everything's broken. And then there's people that say, "Our products been 100% built by agents." Yes, we know it sucks now. Congratulations. [applause] Agents are actually combounding boooos, which is my word for errors, with zero learning and no bottlenecks. and uh delayed pain. The delayed pain is for you. Those are my most beloved people. I don't even read the code anymore. Congratulations. Something is broken and your users are screaming. So, who you going to call? Not yourself because you haven't read the code. Non-critical code, sure, wipe slop ahead. Critical code, read every line. So, two folks, same conference, day after day, talk about the one anxiety that we all feel. Should we all still be reading code in 2026? By the way, these two folks are the number six and number seven most watched YouTube videos from AI engineer from all time. So they're obviously representing something that we're feeling we're talking about and this is being the leadership track something that folks that report to you are talking about. Okay. Should they still be reading code and what's the what's the level of quality? So they name the same exactly from both ends. Uh at this point I probably should introduce myself. Uh hi I'm Alex Walov. I'm the host of Thursday podcast. It's a podcast and newsletter. We go live every week to talk about AI. For the past three and a half years, we've been tracking every change in the engineering, every release from every lab, every model, and I'm also an AI evangelist with Weights and Biases and Core Weave. Um, what also should I tell you about myself that I've been covering AI engineers specifically since the first one in 2023, and oh boy, has it changed. And so, you can treat this as a dispatch from the front line because all of these people now are my friends. And we constantly talk about this in the speakers room, in the hallway track. I couldn't stop thinking about that tension. I couldn't stop thinking about that kind of disparity between the two folks. Okay? And I put them both on the line. Zner from one end, Lop on the other end. I called it a continuum. And I basically started asking people, hey, where are you on this line? Where are you a Zner? Do you still read every line of code? Are you a Leopolo? Do you just yolo and don't even look at code and think agents are good enough, etc. And I I got the framing wrong, but I'll tell you about this in just a second. Okay, so before this, I want you to be get honest with yourself. And again, if you don't write code or let me say this, if you don't babysit your own agents, but you you have reports that babysit agents for you, uh think about them when you answer this, okay? And be honest on the ZL continuum, where are you? And let's take um let's take by vote of hands. Who here has committed code that they've never looked at before? Amazing. Love that. Uh who here still reads every line of code of at least critical code? I see one cowboy over there. I love that, man. I'm going to talk to you afterwards. Okay. I want to understand exactly why you do this. Um and so who's right? Let's talk about who's right. Let's talk about where we are right now. And we start with Ryan Lopolo. If you get to meet Ryan over here, he is very AGI pill. I think even within OpenAI, the AGI organization, Ryan is kind of like the more AGI pill person. Uh if you had a chance to go downstairs and grab the AGI pills that's prescripted, I think Ryan had all of them. He works at OpenI. Where he sits, code is literally free. So are tokens. Uh it's we renamed Ryan uh do you guys know the D- yolo in Codeex? It's kind of like the skip dangerous permissions in cloud code. So we were named the Ryan Loapollo. YOLO. He's okay with it by the way. I asked him. So if we check his kind of side, the the folks like him against the data, they're actually right. The optimists are right, at least about output. This is from Ferros AI. I think I'm not the only speaker at this conference who cites this essay. It's uh sorry, this survey, it's new from April 2026. I think it's one of the best kind of evidence of where we're going that we can now site. Okay. 22,000 engineers were surveyed about code. They call this the acceleration whiplash and they're talking about my favorite stat on here and you can read this yourself. 861% increase in code deletion per PR. So us together with with the Asians will love deleting code. Uh Andromeic also said that they are shipping eight times more code per quarter than in 2025. But is it all good code? Okay, let's play a game and if you know the answer, you let me have my moment on here on stage. Okay, but if you don't know the answer, let's guess whose status page is this. I think I I hear a few answers. I think most of us guessed it. This is Quad. In fact, as you can see on the on the right, it was down when they took the screenshot. It was really funny. Uh uh this Entropic is the company that probably uses the most AI generated code and their status page looks like a Christmas tree. Now I'm not here to dunk on Entropic. Sar just did an incredible job back on stage uh talking about Cloud and etc. Um this may be due to scale. This may be due to other factors. Uh I'm not here to dunk on them, but [snorts] it just goes to show that they're not the only ones like this. Obviously, GitHub famously also suffers from a little bit of growth. Um output does not mean stability. Okay, so maybe this is a good example of what? Same essay, 31% increase in PRs merged with no review at all, human origentic. Don't do this. I beg of you, don't it's it's we'll talk about how to fix this in a second. So when you ship this fast and this much something gives and usually it's quality. So maybe Mario's right. Yeah, maybe the bill does come due in production. Same study, 242% increase in incident per kind of scary. The second study is also scary. Bugs per developer is up six times than 2025. So even on traffic conceds this, I don't know if you guys read the RSI essay they posted, the recursive self-improvement where they talk about, hey, what's does the future hold? They outline two scenarios. One of them says maybe the acceleration will stop and we're going to get used to this. They all they say that's actually not likely to happen. We just added this uh eventuality for clarity. We don't think that's likely to happen. What we think is going to happen is engineers and companies 10xing to 100xing to a,000xing their output and productivity. And then they say this, we as we began to push more code around the organization, human code review has become a new bottleneck. They're citing Amdo's law that shows that if you have an explosion of productivity in one area, another area going gets blocked. And nobody removes the human in these organizations. In fact, careers in entropic and careers in OpenAI. They're still hiring humans. So, nobody's removing the human. And they're both saying that human code of view is still a concern. And here's my ma culpa. I promise you I'll tell you where I got it wrong. The framing my mopa is the continuum is real. The ZL continuum is real. But it's not about the people. It's about the tasks. The continuum is real. It's not about the people. It's about the task. Same engineer could be a Ryan Lopa on one piece of code and has to be Mario Zner and read every line of other pieces of code. Different tasks just need different proof. If we look at them closely, I obviously character characterize them. I've practiced this word multiple times and I still got it wrong. Characterized them. They're a character uh on both ends for the ZL continuum. But if you look at them closely, what they're saying closely, they're actually not that different. Ryan's mechanism is moving attention up the layer. He's saying humans are unreliable at catching repeated mistakes of the same time. Repeatedly catching the mistakes of the same time. So when they you do catch a mistake during the PR review, write the documentation, the llinter and the reviewer needs to remember this once so the system will catch this type of bugs. He's not saying don't inspect your code. He's saying inspect the system, not every line. Mario from the other end is saying route by task. If it's not critical, let it rip. He said it. And if it's critical, you read every line. How do you know what's critical? Well, his answer is easy. You read the f code. Uh my answer to add to this is also you ask your clankers. They're great at looking at a large repository and telling you and telling you, hey, this line is actually critical. You should look at this area. This primitives over here are critical. So you ask your clanker. So they agree more than I kind of gave him credit for. And so I think at the beginning of this, the wrong question is should I still be reading code in 2026? I think the better question right now for all of us is what proof does this specific change need? What proof does this specific change need? And so I took Mario on the left obviously. I took Ryan on the right and then I took a bunch of other great AI engineers friends some friends uh many of the speakers at this conference and kind of distilled their advice down to a routing table. And uh they told me I think Swig told me on Twitter there's going to be one slide that I will that people need to take a screenshot of. It's going to be this slide. You don't have to read it with me, but at the end you're welcome to take a a picture of this. This is the your Monday artifact. Routing the change where the proof needs it. Routing the change to the proof that it needs. You read every line of authentication, money movement, permissions and reversible data. You inspect the critical path yourself and then obviously you keep going. Uh decomposing I think is very important. The more code is getting written, the more it's hard. Your eyes are starting to glaze over a a very long pull request. So splitting into atomic reviewable PRs, you know who's good at it? Agents, they're great at decomposing code. Ask them to do it. You verify that doesn't go away. This has been with us in in engineering, software engineering and AI engineering. It doesn't go away. Traces, evals, shadow mode. Come talk to me after after this talk. I don't have enough time, but shadow mode is a really cool one that I learned while preparing this talk. And then I think the most important one is separating. Many people have the same agent that writes the code, also inspects the outputs and writes the test. Separating is very important. If you don't separate, it's kind of like if I came up with an exam and then I took an exam and I scored myself on the exam. It's not not really productive, right? And then last one is engineer. Rails, observability, roll back. This is what Ryan Luplo talks about. Build a system that builds the system because read spends your attention once engineer makes the system remember, right? And you might sitting might be sitting there and saying, "Hey, did you hear the news? Alex Fable is back. What about Fable? What about Mythos? Is this still relevant at this next scale of capability? Because when I coined the ZL Continium, it was only 82 days ago. Mythos has just been announced. We weren't sure like what's going on. Only the people in Entropic got access to it. And Derek uh Shipar that was on stage from Entropic. He said about Mythos and and Fable, we used to check if cloud is doing the work right. And with Fable 5, I instead check if cloud is doing the right work. Let it land for a second. I don't know if you read the statement. When I read the statement, I felt like little chills at the back of my neck about the next like level of capability. Okay, we used to check if cloud is doing the work right with Fable and we check if cloud is doing the right work. And our favorite senpai who recently joined Entropic and is getting unnecessary heat on Twitter uh said this Andre Kapathy has said it's never felt so tempting to stop looking at code at all but don't do this in production. Senpai is great for the sole reason. Do you guys know the sentence uh this meeting could have been an email? So this presentation could have been Andre Capasi's one sentence. Okay, he's naming the anxiety from both ends. It's never been so tempting to stop looking at code. Don't do this in production even with fable. And so if you guys noticed uh I have a little thingy here. This uh it's so white you can see my little uh laser pointer. Do you guys see the arrow the capability drift arrow? This thing when I wrote the continuum I realized that it it's only a temporary place in time. Capability increases move us towards looop. So we're going to talk about capability increases as well because the review layer moves. If yesterday we inspected the outputs and we read the code uh and today we inspect the task direction and kind of like directed to the right proof maybe tomorrow we're inspecting the loops capability drift changes where proof belongs. It doesn't remove the requirement of proof. Talking about loops is that the next primitive I think most of this conference I think the zeitge guys for this one is going to be is token factories and co factories are real and is loops is a real thing that I need to be doing at this point by raising of hands who here heard of loops keep your hands up please and take them down if you are not running loops right now and you have no idea what they are there's a good perception there's a good uh number of people here who heard about loops and they started with both these folks Peter Steinberger creator of open cloud Open claw and now is open AI and Boris Churnney and pretty much within the span of two days both of them started talking about loops that became kind of the zeitgeist and loops are moving us from prompting each turn to designing the system that writes the actual prompts. By the way, do you guys know what's common between these guys and what's different between me and these guys? Their tokens are free. So when they talk about loops and their tokens are free uh they're not telling you hey you should be doing doing this right now specifically but because they work at bigger labs you can treat them as kind of a lighthouse that pointing where we're all going kind of like Gretzky skateboard where the pock is going to be they're going to tell us what all of our enterprises are going to get get up on and if it's if it's loops then let me at least give you a tlddr okay loops are basically fancy chrome jobs that run on a schedule but what they do is they discover a task and kind start writing a prompt for this task from the plan. They run the plan. They execute and most importantly for my talk here, they verify themselves and if it doesn't work, they try again. So an agent that loops grades its own work against a goal with less human intervention. But if the builder grades itself, you didn't remove the review. You hit it. Okay, this this connects to my routing table. This comes from Addiosmani recently at Google. He's also at this conference, a great engineer. Uh he said if if I wrote if I wasn't reviewing the code myself or relied entirely on automated loops to fix my code, let's say a bug comes up in Jira and my loop picks it up and starts fixing this, my product quality would suffer. I'd likely end up in a downward spiral digging myself into a deeper hole. So again, loops don't remove judgment, but they do raise the stakes on where you put it. So what about the future, folks? Nobody knows. Entropic did not know that cloud code is going to explode in them and this is going to be a billion dollar product. Uh nobody knew that coding agents and harnesses are going to be the generalized agent and now everybody's pursuing them. Folks at openi with codeex Elon with with gra code uh Google with with anti-gravity model capability is jumping at an insane pace. And what I implore and to tell you here is that flexibility is required. You need to keep been nimble to keep up with with the trends. This is why you're a engineer. And by the way, I told some folks here about my podcast Thursday news. If you want to keep tracking where that line moves, feel free to scan this QR code, join our, you know, newsletter, etc. Um, and I leave you with this because it's my time. I leave you with this. Not every line in 2026 needs your eyes. Every system still needs your judgment. Thank you. [applause]
