Transcription: OpenAgents Episode 202 - Recursive Language Models
June 9, 2026 ยท View on GitHub
Source: https://x.com/OpenAgentsInc/status/2008704591110541567 Wiki source: https://raw.githubusercontent.com/wiki/OpenAgentsInc/openagents/Video-Series.md Media title: OpenAgents - Episode 202: Recursive Language Models We discuss recursive language... Upload date: 20260107 Transcription model: gpt-4o-transcribe-diarize Generated at: 2026-06-01T02:54:34Z
Machine-generated transcript. Review speaker labels and wording before using this as quote-grade source material.
[00:00] Christopher David: Agents are hungry for compute,
[00:01] Christopher David: and while it would be nice to have our own gigawatt data center in Texas,
[00:06] Christopher David: you know we'll have one eventually.
[00:08] Christopher David: Stargate is OpenAI's infrastructure program,
[00:10] Christopher David: their flagship campus in Abilene has a 1.2 gigawatt, so I guess that'll be done this year.
[00:15] Christopher David: Congratulations and welcome to Texas by the way,
[00:17] Christopher David: but I'm more interested in this five and a half gigawatts over here.
[00:22] Christopher David: 110 million max,
[00:24] Christopher David: huh?
[00:24] Christopher David: So this is infrastructure right here we don't need to pay for.
[00:28] Christopher David: To build anyway. We're just gonna pay to rent it to rent pay you for your spare compute, so this is point number one from our uh big picture episode two one. We said hey you, you've got spare compute. Can we buy it from you? Or if not us, someone else in this global marketplace we let you connect that to.
[00:53] Christopher David: So that's what we're doing.
[00:55] Christopher David: We already demonstrated the technology behind this in episode 178, how you click a button and make your compute available over an open protocol that anybody can buy and pay you Bitcoin.
[01:09] Christopher David: Our example here used Olama instead of the on-device inference. We will be.
[01:18] Christopher David: uh adding olama and other support for other model types and hardware types so everyone will be able to sell spare compute we're going to start with apple silicon partly because there's just no download you don't have to download a model it's built into your computer so it's very easy to like run a script go online and just like instantly have bitcoin streaming so that's what we're going to launch on wednesday uh the name of the product is called pylon
[01:44] Christopher David: Oh, we did an earlier version of this earlier this year,
[01:46] Christopher David: but basically it's a Node software that runs on your computer that makes your compute available on this open market with a built-in Bitcoin wallet,
[01:55] Christopher David: so you just get money streamed to you.
[01:58] Christopher David: And we
[02:02] Christopher David: launched this before,
[02:03] Christopher David: so we kind of know what we're doing. We had 300-ish people online at the same time renting or selling their compute.
[02:12] Christopher David: And we were the only buyer really.
[02:14] Christopher David: This was at the time back in 2023.
[02:18] Christopher David: There was no agents.
[02:19] Christopher David: There was no MCPs, no tool calling,
[02:22] Christopher David: very early using horrible open models.
[02:26] Christopher David: But we validated the idea and the tech works.
[02:29] Christopher David: So we're kind of doing an updated version of this now.
[02:31] Christopher David: But we phased this out and we predicted that the future sort of buyers here are going to be agents,
[02:39] Christopher David: agents of the future.
[02:41] Christopher David: And when agents get good enough,
[02:43] Christopher David: that'll be the right time to relaunch this.
[02:45] Christopher David: Well,
[02:45] Christopher David: they are now good enough.
[02:49] Christopher David: Part of why now is particularly the perfect time, and we'll get into more of this in the next video,
[02:55] Christopher David: there's a new paper that came out in the last couple days or month,
[02:59] Christopher David: Recursive Language Models, which this other lab says is going to be the paradigm of 2026.
[03:05] Christopher David: And part of this model is that there's this like
[03:09] Christopher David: desire to have these very modular small micro jobs stand out it's basically the absolute perfect application for the type of compute that we're talking about so we think there's going to be a healthy buy-side demand for all of this so we'll get into that more in the next video I just want to kind of step through a little bit of this analogy here
[03:32] Speaker B: Have
[03:32] Christopher David: I want to conclude by defining these three terms you'll hear us use because the parallel to oil exploration is pretty clear here.
[03:40] Christopher David: Stranded, fragging, wildcatter. Okay,
[03:42] Christopher David: stranded energy. Energy that exists but can't be sold at the right price or time because it lacks a market path.
[03:48] Christopher David: No pipeline of transmission,
[03:49] Christopher David: no storage, no off-take contract with a transaction cost exceed the value.
[03:53] Christopher David: Follow Daniel Barton on Twitter,
[03:55] Christopher David: by the way, for examples of how Bitcoiners generally are great at monetizing stranded energy.
[03:59] Christopher David: energy um anyways the parallel there to compute compute capacity that exists but can't be economically traded because it lacks one or more of discovery buyers can't find it packaging no standard job or api product trust no verification reputation settlement can't pay for tiny work units operability no observability no replay no receipts in one line stranded means it's a real research missing market plumbing
[04:25] Christopher David: 110 million max,
[04:26] Christopher David: yes.
[04:27] Christopher David: Two,
[04:27] Christopher David: fracking.
[04:28] Christopher David: Energy.
[04:30] Christopher David: A set of techniques that unlock hydrocarbons trapped in low permeability rock by changing flow dynamics,
[04:35] Christopher David: pressure and permeability,
[04:36] Christopher David: creating a new economically recoverable supply.
[04:40] Christopher David: Compute. Compute fracking.
[04:42] Christopher David: A set of protocols and products that unlock compute trapped behind distribution and coordination barriers.
[04:48] Christopher David: By changing market flow dynamics,
[04:50] Christopher David: turning idle devices into routable paid verifiable supply,
[04:53] Christopher David: our fracking fluid equivalent is streaming money,
[04:57] Christopher David: Bitcoin and Lightning,
[04:58] Christopher David: plus receipts, plus routing.
[05:00] Christopher David: See our episode 94 or something called streaming money.
[05:03] Christopher David: Money is the incentive that makes the supply show up and stay online.
[05:07] Christopher David: Micro-payments make tiny work units worth selling.
[05:09] Christopher David: Receipts make spend auditable and procurable.
[05:12] Christopher David: Budgets bound risk and enable autonomy.
[05:15] Christopher David: Routing and reputation turn chaos into a market.
[05:17] Christopher David: In one line,
[05:18] Christopher David: fracking is inject the missing infrastructure so the resource flows into the market.
[05:23] Christopher David: Imagine tapping into the 5.5 gigawatts of compute that's just sitting there.
[05:28] Christopher David: Not to mention all the other edge inference.
[05:30] Christopher David: Oof.
[05:32] Christopher David: Wildcatter. My friends, we are building an army of wildcatters here. Check this out.
[05:39] Christopher David: Energy. An operator who drills in unproven territory,
[05:42] Christopher David: high uncertainty,
[05:43] Christopher David: high upside,
[05:44] Christopher David: before reserves are fully proven. They take exploration and operational risk early. Compute wildcatters.
[05:51] Christopher David: An early provider who brings unproven or stranded compute online as markets apply before demand,
[05:57] Christopher David: pricing and trust are fully established.
[06:00] Christopher David: Taking risk on demand volatility,
[06:02] Christopher David: reliability in turn,
[06:03] Christopher David: verification,
[06:04] Christopher David: reputation building,
[06:05] Christopher David: payout and ops overhead.
[06:07] Christopher David: Examples, a household or office turning Macs into a bundle land provider.
[06:12] Christopher David: A prosumer running pylon provider mode overnight.
[06:15] Christopher David: A miner deploying a small GPU pod alongside ASICs to monetize curtailment windows.
[06:20] Christopher David: Why wildcatters matter.
[06:22] Christopher David: They prove economics create early liquidity and generate the first trust reputation signals.
[06:27] Christopher David: Wildcatter is the first supplier willing to take uncertainty to unlock a new market.
[06:32] Christopher David: And if you ask any of the people who ran nodes during our GPutopia,
[06:38] Christopher David: you know, maybe not all of them, but a whole bunch of them got paid some Bitcoin.
[06:41] Christopher David: So that's nice,
[06:43] Christopher David: right?
[06:44] Christopher David: TLDR,
[06:44] Christopher David: stranded,
[06:45] Christopher David: resource exists,
[06:46] Christopher David: market access doesn't.
[06:48] Christopher David: Fracking,
[06:48] Christopher David: add the fluid,
[06:49] Christopher David: incentives plus settlement plus routing plus trust to make it flow.
[06:52] Christopher David: Wildcatter, early operator who takes risk first and earns upside when the market...
[06:56] Christopher David: The market becomes real.
[06:56] Christopher David: So folks, we built this already.
[06:58] Christopher David: We're relaunching this in two days.
[07:01] Christopher David: Just so you know a little something about me,
[07:03] Christopher David: my grandfather worked in the oil fields of the Permian Basin and all across Texas,
[07:10] Christopher David: Louisiana.
[07:10] Christopher David: He was a geologist for Shell and he mapped out the fields and helped them discover I think their largest coal reserve at the time, something like that.
[07:22] Christopher David: I'm a Texan and we're going to unlock this resource and share it broadly and make a whole bunch of money for us and our agents. Okay?
[07:29] Christopher David: See you soon.
[07:30] Christopher David: Alright, we're going to talk about recursive language models.
[07:33] Christopher David: We're going to step through the paper that just came out.
[07:36] Christopher David: We'll talk towards the end about why this is relevant for our model.
[07:41] Christopher David: The one sentence of why this is relevant to what we discussed yesterday,
[07:45] Christopher David: yesterday we talked about like how to connect all these devices into a swarm compute network,
[07:51] Christopher David: but now we have a reason to sort of attach value to that compute because it is absolutely perfect for this RLM model we're about to explain.
[08:02] Christopher David: So we're going to read the abstract and the key points here.
[08:05] Christopher David: Part of what makes me as a like non-academic pay attention to papers like this is when I see people on my timeline saying that they applied it and it's having all these like major improvements for them.
[08:18] Christopher David: So just some codex user here,
[08:20] Christopher David: I wrote a codex skill using the RLM paper and already see a massive gain in output performance and 66% reduced in token usage.
[08:28] Christopher David: Optimize RLM programs on smaller models will bring.
[08:31] Christopher David: Compute scalability to both personal.
[08:33] Christopher David: Okay, okay,
[08:34] Christopher David: this seems like it's worth looking into and the authors are from MIT CSAIL and Omar created DSPI.
[08:43] Christopher David: So these are like serious folks and everyone seems to be talking excitedly about this. So let's step through this and then we'll tie it into why we're using this and tying this into our launch.
[08:56] Christopher David: Tomorrow of the drumroll please one second
[09:04] Christopher David: Yeah, so our software pylon launches tomorrow, uh featuring integration with this model. Okay, let's step through this. We study allowing large language models to process arbitrarily long prompts through the lens of inference time scaling. We propose recursive language models, RLMs, a general inference strategy that treats long prompts as part of an external environment and allows the LLM to programmatically examine,
[09:28] Christopher David: decompose, and recursively call itself over snippets of
[09:30] Christopher David: zip it to the prompt we find that RLM successfully handle inputs up to two orders of magnitude beyond model context windows and even for shorter prompts dramatically outperform the quality of base LLMs and common long context scaffolds across four diverse long context tasks while having comparable or cheaper cost per query this is striking at just major limitations of agents and models like
[09:57] Christopher David: getting towards super massive context while being cheaper.
[10:02] Christopher David: So we'll go a little bit more into the background here. Despite rapid progress in reasoning and tool use, modern language models still have limited context lengths,
[10:12] Christopher David: and even within these limits appear to inevitably exhibit context rot.
[10:19] Christopher David: Though context lengths and sets will keep improving from training,
[10:22] Christopher David: we are interested in whether it's possible to dramatically scale the context size of general purpose LLMs by orders of magnitude.
[10:28] Christopher David: This is increasingly urgent as LLMs begin to be widely adopted for long horizon tasks in which they must routinely process tens if not hundreds of millions of tokens.
[10:37] Christopher David: We study this question through the lens of scaling inference time compute.
[10:41] Christopher David: We draw broad inspiration from out-of-core algorithms in which data processing systems with a small but
[10:47] Christopher David: All but fast main memory can process far larger data sets by cleverly managing how data is fetched into memory. Inference time methods for dealing with what are in essence long context problems are very common though typically task specific. Context compaction, blah blah blah.
[11:05] Christopher David: Okay.
[11:06] Christopher David: So the key insight of RLMs is that long prompts should not be fed into the neural network directly, but should instead be treated as part of the environment that the LLM can symbolically interact with.
[11:28] Christopher David: Benchmarks, blah, blah, blah.
[11:30] Christopher David: Observation one. R_L_M_s can scale to the ten million plus token regime and can outperform base L_M_s and existing task agnostic agent scaffolds on long context tasks.
[11:40] Christopher David: Cool. Observation two. The REPL environment is necessary for handling long inputs while the recursive sub-calling of R_L_M_s provides strong benefits on information-dense outputs.
[11:52] Christopher David: Observation three. L_M_ performance degrades as a function of input length and problem complexity while R_L_M_ performance scales better.
[12:00] Christopher David: Observation four, the inference cost of R_L_M_s remains comparable to a base model call,
[12:04] Christopher David: but are high variance due to differences in trajectory lengths.
[12:09] Christopher David: Observation five, R_L_M_s are a model agnostic inference strategy, but different models exhibit different overall decisions in contrast management. Okay.
[12:22] Christopher David: Okay, so here's kind of the key section for
[12:26] Christopher David: our purposes.
[12:28] Christopher David: While R_L_M shows strong performance on tasks beyond the context window limitations of existing L_M_s at reasonable inference costs,
[12:35] Christopher David: the optimal mechanism for implementing R_L_M_s remains underexplored. We focused on synchronous sub-calls inside of a Python REPL environment, but we note that alternative strategies involving asynchronous sub-calls and sandboxed REPLs
[12:53] Christopher David: can potentially significantly reduce the runtime and inference cost of RLMs.
[13:00] Christopher David: Furthermore,
[13:01] Christopher David: we chose to use a max recursion depth of one.
[13:04] Christopher David: While we found strong performance on existing long context benchmarks,
[13:07] Christopher David: we believe that future work should investigate deeper layers of recursion.
[13:10] Christopher David: Okay.
[13:13] Christopher David: So, taking a task,
[13:15] Christopher David: decomposing it into a bunch of subtasks.
[13:18] Christopher David: And then wanting to throw that out to small models,
[13:24] Christopher David: you know, you're basically having one smart model throw it out to a bunch of small models to be able to throw
[13:36] Christopher David: out subtasks to like a hundred different models at the same time.
[13:45] Christopher David: different tasks have that come back within a few seconds use that throughout more like there's a lot a lot a lot that can be done that I think that this model we've been talking about of when you have millions of Macs connected into a swarm compute network where you're not needing to care about one or a few cloud providers rate limiting you or forcing you to do things sequentially like
[14:13] Christopher David: A swarm compute network can let you and how we've architected our system with Noster,
[14:18] Christopher David: you can literally send out jobs like a thousand of them to a thousand different providers.
[14:26] Christopher David: They're just signed JSON blobs over WebSockets that a thousand providers can pick up.
[14:33] Christopher David: So all of that's going to be enabled by the software that we launch tomorrow.
[14:37] Christopher David: To be clear, for the first week we're going to be using only testnet funds for Bitcoin.
[14:41] Christopher David: The live money version will go live the following Tuesday,
[14:46] Christopher David: January 14th.
[14:48] Christopher David: Okay,
[14:49] Christopher David: we'll pause there.
[14:50] Christopher David: Aside from just to kind of emphasize another lab doing decentralized open AI stuff says they're super excited also about RLMs.
[15:01] Christopher David: The paradigm of 2026.
[15:13] Christopher David: So um I think this is exciting. I think um maybe there's a bunch of researchers who would appreciate having access to a swarm network of nodes running the models to enable these jobs to be thrown out to thousands of people at a time. And um tomorrow we're gonna release the software letting you become one of those nodes on the network and we'll start uh testing this in practice. See you soon.