Transcription: OpenAgents Episode 206 - Codex on Autopilot

June 9, 2026 ยท View on GitHub

Source: https://x.com/OpenAgentsInc/status/2016039535134560301 Wiki source: https://raw.githubusercontent.com/wiki/OpenAgentsInc/openagents/Video-Series.md Media title: OpenAgents - Episode 206: Codex on Autopilot We demo the first alpha build of Aut... Upload date: 20260127 Transcription model: gpt-4o-transcribe-diarize Generated at: 2026-06-01T02:34:31Z

Machine-generated transcript. Review speaker labels and wording before using this as quote-grade source material.

[00:00] Christopher David: So no,

[00:01] Christopher David: 2025 was not the year of agents.

[00:04] Christopher David: It was the year of co-pilots, agentic co-pilots, co-pilot agents.

[00:09] Christopher David: You take an IDE,

[00:11] Christopher David: you take a terminal,

[00:13] Christopher David: you take a browser,

[00:14] Christopher David: you add some agentic workflows,

[00:17] Christopher David: you have an LLM in a loop calling tools steered by you.

[00:23] Christopher David: It's powerful,

[00:24] Christopher David: makes you feel cool.

[00:27] Christopher David: not agents.

[00:30] Christopher David: Real agents have not been tried.

[00:33] Christopher David: What do we mean by agents?

[00:34] Christopher David: Agents should be autonomous,

[00:38] Christopher David: long-lived, they learn and evolve.

[00:41] Christopher David: They are autopilots, not co-pilots.

[00:48] Christopher David: So 2025 was the year of co-pilots.

[00:51] Christopher David: 2026 is the year of autopilots. You heard it here first.

[00:56] Christopher David: So let's take a page out of Microsoft's playbook here.

[00:59] Christopher David: Look at this.

[01:00] Christopher David: Microsoft Copilot,

[01:02] Christopher David: R,

[01:02] Christopher David: registered trademark.

[01:05] Christopher David: I don't want to do that.

[01:07] Christopher David: But I do like jumping on the name,

[01:10] Christopher David: not Copilot.

[01:12] Christopher David: Lame. How about OpenAgents Autopilot?

[01:16] Christopher David: Yes, okay,

[01:17] Christopher David: so this is it everybody this is it this is the culmination of what we've been working towards Since February of this year our quest for the Holy Grail of agentic software development an agent that can reliably Code can convert issues to PRs or push directly to main in an automated loop

[01:40] Christopher David: Everyone's moving in this direction we are there

[01:45] Christopher David: We are there.

[01:46] Christopher David: I've had autopilot running for almost 24 hours out of the last 24 hours.

[01:54] Christopher David: It's run overnight reliably twice.

[01:59] Christopher David: Just come take a look at the freaking commits.

[02:04] Christopher David: oh let's repo look it's coding while i'm freaking talking yeah i had the apm storage the trajectory collector and it's it's tomorrow already wow i had to play the keeps trying to add github workflows like don't do that just these commits are so good now you know you've probably in the past let the eight-year agents run wild like make commits and like the code degrades over time the code is getting better over time

[02:31] Christopher David: It's amazing Okay, so the point is that this is working

[02:37] Christopher David: Why will you like this?

[02:39] Christopher David: It's open source

[02:41] Christopher David: Just use it. We were gonna like make a landing page and email capture and like tell you all this stuff No, no point your agent at the repo open agents slash open agents and ask it to tell you about autopilot and how it worked for us and how it can work for you Okay

[02:57] Christopher David: We'll have more human docs soon,

[02:59] Christopher David: but like here's the all of the what we're building autopilot is one big piece of that open brand Make your own autopilot competing companies.

[03:08] Christopher David: I don't care like I hope you do.

[03:10] Christopher David: Let's compare autopilots The more exciting thing is when we start really having our autopilots learn from each other

[03:18] Christopher David: So more on the infrastructure enabling that soon.

[03:23] Christopher David: MechSuit. It works with any agent. We're focusing on cloud code first.

[03:28] Christopher David: Cloud code's agent SDK is very, very, very good.

[03:31] Christopher David: They did not have a Rust version.

[03:33] Christopher David: Our whole thing is in Rust.

[03:34] Christopher David: We wrote a Rust version just by porting it from TypeScript.

[03:38] Christopher David: It's powerful as heck.

[03:41] Christopher David: We don't want to like reinvent cloud code right now because it's the best in some ways,

[03:45] Christopher David: but just by putting it in this mech suit,

[03:47] Christopher David: was able to get it to code reliably overnight.

[03:51] Christopher David: So that's nice.

[03:53] Christopher David: Revenue sharing.

[03:54] Christopher David: Okay,

[03:54] Christopher David: go back to episode one of our video series.

[03:57] Christopher David: This is episode 199.

[03:58] Christopher David: It's all here on our wiki.

[04:01] Christopher David: Episode one was all about, this is the day after dev day 2023,

[04:05] Christopher David: Sam Mambo was like, yeah, we're going to do a GPT store and do revenue sharing. I was like.

[04:08] Christopher David: There's no way. They're definitely going to half-ass that.

[04:12] Christopher David: And they did.

[04:14] Christopher David: But we want everyone to get paid to use this.

[04:18] Christopher David: Here's an example.

[04:19] Christopher David: Ryan Carson,

[04:20] Christopher David: I spent $200 on these skills and it was worth 5x that.

[04:24] Christopher David: We're going to see agent skills marketplaces appear soon.

[04:30] Christopher David: Yes,

[04:30] Christopher David: if you're spending on things that you discover,

[04:33] Christopher David: you could choose to share and keep those private or you should have a way of listing those on a marketplace and getting paid for them.

[04:40] Christopher David: So it's really great that there's already a ton of really cool stuff being shared openly,

[04:45] Christopher David: open source.

[04:46] Christopher David: That's really great.

[04:47] Christopher David: But we believe that if there's another layer to it where people can optionally set prices on things that they're sharing,

[04:53] Christopher David: if you have a financial incentive.

[04:56] Christopher David: Then maybe you're going to go the extra mile and make your skills extra good.

[05:00] Christopher David: Maybe you're going to be extra motivated to maintain it in the usual decay in open source projects because people don't have a financial incentive to maintain them.

[05:08] Christopher David: Maybe that starts to shift if you're getting paid a stream of micropayments whenever somebody uses your thing.

[05:15] Christopher David: So one of the like things that we've been coding here is a way to...

[05:22] Christopher David: ensure through these key splits because agents are able to hold their own private key split across different guardians or different froster nodes we can basically help ensure that paid agent skills can't be copy pasted to a different agent all sorts of cool stuff that we can do with cryptography we'll do separate videos on that so let's get you paid let's get autopilots as the best coolest

[05:51] Christopher David: um way of using these agents. Uh over the next day or two you'll start seeing us rolling out our uh G_U_I_ for all of this that are kinda like Starcraft style uh HUD U_I_s. This is all kinda C_L_I_ right now.

[06:07] Christopher David: Uh so we're gonna start doing about daily videos, rolling all this out, starting with our next episode, episode two hundred, we're gonna zoom way out.

[06:18] Christopher David: And we're going to tell you the big picture of what we're building and why, and uh how you can help. See you soon.

[06:26] Christopher David: Autopilot desktop app,

[06:28] Christopher David: very first alpha build.

[06:30] Christopher David: Let's check it out.

[06:33] Christopher David: So autopilot begins kind of as a UI for codex, the best long running coding agent at the moment.

[06:42] Christopher David: But it's got one key addition on the right sidebar here you'll see something called full auto.

[06:49] Christopher David: And when I enable that and chat.

[06:53] Christopher David: At the end of a turn, instead of it stopping or instead of going with the next thing that you've queued up, it's going to intelligently select the next action using a new agentic primitive we're calling Guidance Module.

[07:10] Christopher David: And we're going to step through that.

[07:11] Christopher David: Let's compare it a little bit to what people are doing right now with Ralph and stuff.

[07:16] Christopher David: So how I and a bunch of people who use codecs right now...

[07:20] Christopher David: uh you know we queue up something maybe give it a big refactoring task sometimes it can take up to 20-30 minutes or 45 minutes it can really do great kind of long-running tasks but when it finishes a chunk of work it'll say you know here's my summary i'm done and if you want to like go afk for a while you can queue up different messages which might be do something else or it might be continue if you give it a big thing continue continue continue

[07:49] Christopher David: um so that's good codex spits out very reliable stuff usually more reliable than cod code i've found but then you've got the ralph loop which aims to remove the human from the loop and replace it with this sort of dumb aka ralph uh it just does the next box on the checklist and hopefully you've kind of spec'd out out

[08:15] Christopher David: And it just kind of like dumbly goes through and does each PRD checkbox until hopefully you have something usable at the end of that process.

[08:25] Christopher David: So we're going to kind of take a little bit of the thoughts of both of these and kind of keep going with this idea of, yes, we want to remove the human from the loop, but we want to add more intelligence than just going down.

[08:40] Christopher David: a list maybe something happened maybe you got rate limited on one of your accounts maybe some new information came up maybe there's a new plug-in that just got discovered that's doing what you're trying to do better we want there to be some intelligence between these turns so we're introducing this new primitive called a guidance module and we've got a little

[09:06] Christopher David: Kind of specing this out all this is open source all this will be

[09:11] Christopher David: you know improved upon gradually if you're watching this after January 26 2026 it will be better and stuff by the time that you're looking at it Okay, let's step through this so turn to turn guidance module the purpose is to describe the intelligence layer that runs between codex turn

[09:28] Christopher David: Inside a full auto run calling this full auto

[09:31] Christopher David: for autopilot and we're doing codex for now just because it's the best you know coding agent it's a good thing to start with it's open source but you know this should be able to be worked with multiple different agents we're going to start with codex first so it replaces the manual continual loop with a structured extensible guidance module that can be evaluated and improved over time

[09:53] Christopher David: So for the definition,

[09:54] Christopher David: turn is one codex execution window.

[09:57] Christopher David: Ending a turn completed,

[09:58] Christopher David: that's the session that lasts for 30 minutes or five minutes or however long in the context of a thread or session.

[10:06] Christopher David: A run is a multi-turn session,

[10:08] Christopher David: blah, blah, blah, composed of turns.

[10:09] Christopher David: Guidance is a soft recommendation for what to do next.

[10:12] Christopher David: Guardrails are hard deterministic constraints that could override guidance.

[10:18] Christopher David: Okay,

[10:19] Christopher David: so today long running codecs runs are driven by manual loop. You issue a run, then type continue or queue if you want prompts.

[10:24] Christopher David: Turns are not aware of each other beyond ad hoc file handoff.

[10:28] Christopher David: You're kind of reading and writing to markdown docs,

[10:30] Christopher David: which is like okay.

[10:31] Christopher David: There's no shared model of state,

[10:33] Christopher David: budget,

[10:33] Christopher David: or environment across turns.

[10:35] Christopher David: This is brittle and does not scale.

[10:37] Christopher David: The system needs a real decision engine between turns that can see the full context of what happened in the last turn, understand the goal,

[10:44] Christopher David: constraints,

[10:45] Christopher David: environment,

[10:45] Christopher David: and remaining time.

[10:46] Christopher David: Remaining budget decide the next action with measurable confidence improve over time via evals and community contribution as Well as maybe pull in plugins or tools from third-party registries. There's all sorts of ways to kind of like add extensibility into that

[11:01] Christopher David: Decision process so whereas before the human is the scheduler you're or with a planner or you're prompting it to make a plan and then do it and then keeping it on task or You're delegating to this kind of like dumb loop

[11:14] Christopher David: Afterwards, we're kind of putting more agency and trust into the guidance module,

[11:21] Christopher David: improving the autonomy,

[11:22] Christopher David: blah, blah, blah, blah, blah.

[11:24] Christopher David: Alright,

[11:25] Christopher David: key idea,

[11:26] Christopher David: replace the continue step with a DSPI-powered guidance module.

[11:30] Christopher David: So, in terms of like Asian architectures, the framework, I don't even know what you call it,

[11:38] Christopher David: DSPI maps perfectly to this.

[11:40] Christopher David: DSPy stands for declarative self-improving Python.

[11:44] Christopher David: We're not doing this in Python, we're doing it in Rust,

[11:46] Christopher David: but there's a Rust version of this same library called DSRs.

[11:51] Christopher David: And just the like super fast introduction to what DSPI is and why people like it, it's a way to build AI behavior like you build software with clear interfaces and composable parts instead of one giant prompt.

[12:03] Christopher David: The two main concepts to know from that are signature and module.

[12:08] Christopher David: A signature is just a typed contract for an AI step,

[12:11] Christopher David: what you give it and what you expect back.

[12:14] Christopher David: Example, given the goal and what happened last turn and maybe budget constraints or whatever,

[12:20] Christopher David: return the next action,

[12:22] Christopher David: a reason and a confidence score.

[12:25] Christopher David: Confidence score is cool because if it's got very low confidence you can say okay if you don't really know what should be done next then just wait.

[12:32] Christopher David: Or consult some other agent. And then a module is a small program made out of one or more signatures.

[12:38] Christopher David: It can call other modules, combine results,

[12:40] Christopher David: and enforce structure.

[12:41] Christopher David: Example, a next action selector module that reads the term summary,

[12:45] Christopher David: checks budget, and outputs whether to continue,

[12:47] Christopher David: pause, or stop.

[12:48] Christopher David: So you can bundle a bunch of the different, like, signatures into modules.

[12:52] Christopher David: It's all composable.

[12:54] Christopher David: The punchline is that signatures make AI steps reusable and comparable,

[12:57] Christopher David: and modules let you snap those steps together into a bigger system,

[13:01] Christopher David: and then optimize them with...

[13:03] Christopher David: With evals instead of hand tuning prompts forever.

[13:07] Christopher David: There's a whole kind of evaluation and optimization algorithm,

[13:11] Christopher David: I think called MIProV2 that goes along with this so you can have these things just basically improve over time.

[13:20] Christopher David: Okay,

[13:21] Christopher David: let's step back to this.

[13:25] Christopher David: So the guidance module is a composable stack of DSPI signatures and modules plus policy gates,

[13:31] Christopher David: state,

[13:31] Christopher David: and optimization.

[13:33] Christopher David: Because these are all like reusable and composable, we can do some very interesting things with combining this with open source and market incentives.

[13:46] Christopher David: So, once the decision layer is expressed as declarative signatures, clear input-output contracts, it stops being a one-off blob of prompts and becomes a packaged surface. This is the same move JavaScript made with NPM.

[14:01] Christopher David: Before NPM teams hand wrote ad hoc utilities inside each repo,

[14:05] Christopher David: improvements stayed trapped in-house.

[14:07] Christopher David: After NPM,

[14:08] Christopher David: small modules with stable APIs became reusable,

[14:11] Christopher David: searchable, and composable,

[14:13] Christopher David: and ecosystem flourished because anyone could improve one piece and everyone could adopt it.

[14:17] Christopher David: We're starting to have some of that now. People like, you know, copy pasting each other's prompts or like skills are starting to become kind of discoverable pieces, but it's still very clunky.

[14:28] Christopher David: but just this idea of it being composable it can be pulled in you can compose your agentic flow out of a bunch of these different small pieces so with signatures each capability becomes a drop-in package budget policy stop decide or next action selector verifier independently discoverable evaluated and upgradable and monetizable by the way more on that shortly signatures turn agent intelligence into an API

[14:56] Christopher David: Once intelligence has an API,

[14:58] Christopher David: it can have packages like NPM so the ecosystem can flourish. Small composable module,

[15:03] Christopher David: discoverable by anyone,

[15:04] Christopher David: improved by anyone,

[15:05] Christopher David: and monetizable by the people who ship real games. Games. We want a system that is not just open source,

[15:11] Christopher David: but extensible and packaged.

[15:13] Christopher David: Each policy or selector can be shipped as a package. Each package ships with evals and measurable gains. The runtime can load, compare,

[15:20] Christopher David: and route between packages.

[15:22] Christopher David: Contributors can improve a single piece and get credited or awarded.

[15:26] Christopher David: Think NPM but for agent intelligence,

[15:28] Christopher David: signatures,

[15:29] Christopher David: module,

[15:29] Christopher David: and policies.

[15:30] Christopher David: Yeah, so right now you've got people trusting teams like you're trusting the Codex team or the ClaudeCode team or the AMP team to like improve their harness,

[15:40] Christopher David: right?

[15:41] Christopher David: Or you're trusting the Ralph team.

[15:44] Christopher David: And some of those teams that are open source.

[15:47] Christopher David: open code Ralph may improve faster than the closed source ones do but it's still kind of bottlenecked by like what is that one team going to do for their algorithm and their harness this is we think the next phase of that to subject it to true like open market forces you'll just get a lot more innovation and improvements in that way

[16:12] Christopher David: So just a little spiel on how this kind of ties into our broader vision.

[16:16] Christopher David: So this guidance module is the bridge between autonomy and the open agents market.

[16:21] Christopher David: It turns between turn decisions into typed auditable work units that can be routed, verified and paid for.

[16:28] Christopher David: Why this matters.

[16:30] Christopher David: So tying this into our sort of like compute fracking idea of you've got a bunch of spare compute that's on people's computers that they aren't using in part because everyone's just paying Anthropic for their tokens because all of their harnesses only use their own models.

[16:44] Christopher David: Whereas you've got all this untapped compute that's on your device that could be used like it's the ideal type of compute to be used with some of these jobs where things are being broken up into a bunch of small pieces.

[16:57] Christopher David: devices um go check out our recursive language models uh videos a couple videos ago where we talk more about how that's um an ideal algorithm that also is from the founders of DSPI uh anyway there's a lot of pieces that we're pulling together here um we'll I think spend a couple more videos going into individual pieces of some of this um but the big idea here is that

[17:25] Christopher David: We are releasing a single desktop app that has the whole kind of cycle of it's a way to have a coding agent that is, you know, a local coding agent starting with codex, increasingly augmented with local compute,

[17:42] Christopher David: yours and others,

[17:44] Christopher David: if you want to be able to throw out a job to like 100 different like micro jobs,

[17:49] Christopher David: that'll all be kind of.

[17:50] Christopher David: um possible through our swarm compute all of which is going to be built into this autopilot binary so the stuff that we did a couple videos ago about our pylon um all that's going to be folded into this one um binary to keep things simple this autopilot binary is also going to have a bitcoin wallet so the kinds of things that you're doing when you are paying codex to do cool stuff maybe you want to be able to sell some of that data

[18:19] Christopher David: to other people because you discover some gotcha that instead of someone else spending you know an hour or 10 minutes with their agent burning their compute maybe they just want to pay you three cents and save that time so yeah all sorts of stuff is going to open up here but the general idea is that we want you to be able to code on autopilot

[18:46] Christopher David: Sell some of your compute or agentic compute on autopilot if you want to earn some money streamed to your wallet here.

[18:55] Christopher David: We're going to turn compute into software,

[19:00] Christopher David: software into business value that gets priced and sold in Bitcoin.

[19:09] Christopher David: There's a lot here, folks.

[19:12] Christopher David: We'll do another video tomorrow. See you soon.