232.md
June 9, 2026 ยท View on GitHub
A full transcription of the video is provided below:
Car Gonzalez: And we're back. OAPN. Happy Monday to you, Christopher.
Christopher: Happy Monday, Car. Uh, so we're going to go deep here on some concepts that will hopefully shed some light on why I'm so insistent that OpenAgents is the next trillion dollar AI lab. We're going to talk about energy, we're going to talk about compute, we're going to talk about what all these big labs are missing in their sprint to IPO, and it's going to be fun. We're going to talk a lot about Bitcoin. So, we're going to start here. Uh, so yeah, big theme in our upcoming episodes, like this one.
Bitcoin miners are the best in the world at finding cheap electricity. Let's look at what this lady, partner at Category Ventures, says. Anthropic and OpenAI charge 50% less for batch API with a 24-hour SLA because it can use spare capacity and smooth out demand. Great essay by @mariohsouto on the overlooked orchestration layer that co-optimizes energy and compute (forecasting energy prices, grid conditions, battery state, rack temperature, and workload priority) and why we need to think "electrons in, orchestration, tokens out" versus just "electrons in, tokens out." Well, this is right up our alley and, uh, to anyone who's familiar with the Bitcoin mining world, this should all sound a bit familiar.
Let's pop this open. We're going to skim down this and read some of it. "The Clay Jar and the Data Center." Jensen Huang describes Nvidia's business in a single phrase: "electrons in, tokens out." The phrase is doing more work than it appears to. It implies that the production of intelligence is, at its core, the conversion of electrical energy into useful computation, and that the rate of that conversion is what matters: the tokens per joule, the joules per dollar, the dollars per token. I think this framing is correct. I also think it points to a problem that most of the current AI infrastructure buildout is structured to avoid rather than solve.
The dominant response to the AI energy problem, judged by where capex is flowing, is to treat energy primarily as a procurement problem. Sign longer PPAs. Restart nuclear. Buy your own gas turbines. Microsoft and Three Mile Island, Amazon and Talen, Meta and Clinton, and the surge of new gas turbine orders backing up data center campuses across Texas and the Mountain West are each rational responses to real constraints such as scarce firm capacity, long lead times, and the strategic need to lock in resources before competitors do. But they share an implicit assumption: that AI workloads are continuous, massive, and inflexible, and therefore demand continuous, massive, inflexible supply.
I think the procurement view is correct as far as it goes, and incomplete in a way that will matter a great deal in the next few years. The reason is that the demand side of an AI data center is not flat, and treating it as flat leaves most of the available economic value on the table. Orchestrating energy end-to-end, from how it is procured and stored through how it is allocated to workloads in real time, will be one of the most strategic layers of the inference era. This holds regardless of the generation mix. Renewables-plus-storage, nuclear-powered, and gas-fired campuses all face the same underlying problem: a demand side that varies through the day, and a supply side whose value depends on the moment. Orchestration doesn't go away when the generation mix is firm.
Why inference is not flat. There is an assumption embedded in the procurement view that deserves to be made explicit: that AI workloads are continuous and inflexible, and therefore the right job is to procure continuous and inflexible supply. The first half of this is becoming less true every year. A training run is a single coherent computation, fragile and tightly synchronized, that cannot be paused without significant cost. Inference is not. Inference is millions of independent requests with wildly varying latency requirements. Some are real-time chat, where every additional second of latency degrades the product, but some are batch document processing that can wait six hours without anyone noticing. Some are agent workflows where the user has already gone to lunch and will not check the result until a few hours later. Some are scheduled analyses that could run any time before the morning report. In aggregate, inference has substantial temporal flexibility, and the proportion of inference that is genuinely latency-bound is shrinking as agentic and long-running workloads grow.
I'm going to pop over briefly to mention this other article we talked about on a previous episode called "The Inference Shift," where it makes essentially the same point, partly that, well, agents are going to need a lot of compute, and they say, yeah, people kind of draw distinction between training and inference. But he says you actually have to break inference out into two different types of inference. There's answer inference, which is when you're sitting waiting for an answer and you want it to be super fast, and then doing a task, what he calls agentic inference. So Cerebras, who talks all about, "Oh, it's got to be fast," they're talking about answer inference. But the architecture for agentic inference will look a lot different, and also down the road, probably pretty soon, agentic inference is going to dwarf answer inference because there's going to be a lot of long-running machines doing agentic inference, and that's going to be the dominant pattern, and we're not really structured for that today.
Okay, back to here. The market is already pricing this. Both OpenAI and Anthropic offer a 50% discount on their batch APIs, which run with a roughly 24-hour turnaround. That is the going rate for temporal flexibility in inference today, and the fact that the two largest model providers converged on the same number is not an accident. It is the market revealing that a meaningful share of inference does not need to run immediately. Agentic workloads are accelerating the shift. Gartner projects that 40% of enterprise applications will integrate task-specific AI agents by the end of 2026, up from less than 5% a year earlier, and inference workloads, broadly, are projected to rise from roughly a third of total AI compute in 2023 to two-thirds by 2026. The category that is growing fastest is the category with the most temporal flexibility.
This flexibility is, today, almost entirely unexploited at the energy layer. The systems that schedule inference jobs don't see the price of electricity. The systems that manage electricity don't see the value of the next token. The two control loops run in separate buildings, optimized for separate objectives. A facility that could see both at once would, on most days, find itself with hours of cheaper electrons and hours of more valuable workloads, and the gap between them is pure margin that nobody is currently capturing.
So Car, a few weeks ago, we had, uh, the afterparty of the Texas Energy and Mining Summit right here in this building, and we had, uh, you interviewed a few of the folks there. Uh, what did you learn from them?
Car Gonzalez: Yeah, so I think one of the biggest things I learned is that everybody is, uh, it was good that we had the gathering here because we saw so many people finally get to see us and then they kind of put all the, put all the stuff together, and it's right there. It's like power, compute, inference, energy markets, hardware supply chains, economic coordination. And I put some graphs together here and it just talks about more about the global data and the electricity demand and how it's just going to keep increasing. Um, we also, I know you partnered with, uh, with HODL with, uh, sovereign compute and that's something that, uh, I haven't put that in here yet, but it's something that I think is going to lead to even more demand, as far as, like, it's not cloud computing, it's actually, like, a decentralized sovereign way of giving people access to their own private, um, cloud or AI.
Christopher: Yeah, it's like you've got, and they're a little unique because they're putting non-GPU like CPUs on top of the same underlying economics where they're finding, like, super cheap power, they're connecting it to the grid, they're taking advantage of all the ERCOT, um, curtailment and all that stuff, but like, they're powering CPUs and I guess increasingly soon some GPUs, but like, we're using their, for a bunch of stuff. Um...
Car Gonzalez: Yeah, and so, um, I'm going to put this together. I probably will add a little bit of what, uh, you showed today. But as you can see there, the ERCOT, because one of the coolest things about Texas is we have our own, our own grid that we control. And so we are already, we already are used to, or at least I should say, like, most of the Bitcoin miners here already have grown accustomed to, uh, leveraging Bitcoin mining and how to work around those constraints, and I think when you add AI and kind of this model here, uh, it really opens things up for, uh, you know, a frontier AI lab to take advantage of this.
Christopher: For sure. And like, take a look at this prediction. "Looking ahead. By 2028, I think at least one major AI lab will be publicly reporting that a meaningful share of its inference cost is being managed by a combination of storage orchestration and workload routing to regions and times of day where firmed power is cheapest. By 2030, I think the gap in dollars per token between the best and worst data centers will depend as much on energy orchestration as on the underlying silicon. And I think the unified control plane across procurement, storage, dispatch, and workload will become a recognized layer of the AI infrastructure stack. The granary-keeper." Hey Car, do you know any, um, soon to be major AI labs that are based in Texas?
Car Gonzalez: There is one, it's called OpenAgents.
Christopher: OpenAgents! Hang on. And they keep talking about Texas, Texas, Texas, and ERCOT in here. Uh, you know, wouldn't it be great if, uh, we had people, like, on our team who, like, knew about ERCOT and Texas and mining and all this kind of stuff? So let's just say, we're not going to say too much about, like, all what we're preparing here, but let's just say, like, envision that there's a unified model of Bitcoin miner profitability that incorporates AI compute, that includes tools for all of the kind of orchestration and measurement that you're learning about here, and imagine that that is tightly coupled with a product suite from a frontier AI lab that's able to take the best possible use out of that compute. So, so, take a look at that energy. So here's the metric that we're going to be optimizing for that nobody's heard yet. It's called "accepted outcomes per kilowatt hour." Okay? So go ask your, go ask your AI professors and your favorite AI labs, "Hey, what's your guys's accepted outcome per kilowatt hour?" They won't know what the heck you're talking about because we're kind of defining this metric. But this is it. You have a stream of electrons, what is the cost of turning that into an accepted agent task? Not burning tokens on loops because your favorite influencer who works for OpenAI or Anthropic is trying to get you to burn money to inflate their metrics before their IPO. No, no, no. What is the absolute most cost-efficient way of converting electron to accepted agent work? And the name of the company that's going to bring you the solutions for that, it's called OpenAgents. We'll leave it there for now. See you soon.