Transcription: OpenAgents Episode 186 - Actions Per Minute

June 9, 2026 ยท View on GitHub

Source: https://x.com/OpenAgentsInc/status/1948617421654245436 Wiki source: https://raw.githubusercontent.com/wiki/OpenAgentsInc/openagents/Video-Series.md Media title: OpenAgents - @TauriApps Episode 186: Actions Per Minute Flawed ground-truth data ... Upload date: 20250725 Transcription model: gpt-4o-transcribe-diarize Generated at: 2026-06-01T03:47:33Z

Machine-generated transcript. Review speaker labels and wording before using this as quote-grade source material.

[00:00] Christopher David: So how do we measure the effectiveness of our agents?

[00:03] Christopher David: Well, it came out yesterday that one of the big benchmarks for measuring the effectiveness of frontier agents had a bunch of bogus data in it.

[00:13] Christopher David: Some nonprofit did a nice analysis of it and smart people are like, ooh,

[00:17] Christopher David: this is actually pretty bad.

[00:19] Christopher David: Humanity's last exam had a bunch of bogus data. People were measuring and optimizing for the entirely wrong thing and agents are getting all messed up as a result of it. Now.

[00:28] Christopher David: Who made humanity's last exam?

[00:29] Christopher David: Hey, Grok.

[00:30] Christopher David: Oh, it was designed by some AI safety institute in collaboration with Scale AI.

[00:37] Christopher David: Oh,

[00:37] Christopher David: the dude who's helping Meta now become a closed source lab now.

[00:44] Christopher David: Throw enough data slop at the wall from your overseas data farms to get hired,

[00:49] Christopher David: but you know you're screwing with the rest of the ecosystem.

[00:51] Christopher David: Okay, why are these people designing benchmarks that people care about?

[00:55] Christopher David: Screw that.

[00:57] Christopher David: Who should we trust instead to make benchmarks?

[01:00] Christopher David: You know who I think would be really good is developers themselves,

[01:05] Christopher David: who may also be gamers.

[01:07] Christopher David: Gamers know what to measure.

[01:10] Christopher David: Min-maxing? Oh my goodness.

[01:13] Christopher David: We're playing StarCraft here,

[01:14] Christopher David: people.

[01:15] Christopher David: We are inspired by StarCraft. And you know what I want to measure is what is the amount of...

[01:23] Christopher David: effective actions per minute that my agents take.

[01:26] Christopher David: I want to maximize that.

[01:29] Christopher David: So let's take a page from StarCraft.

[01:35] Christopher David: In StarCraft, actions per minute measures how rapidly players interact with StarCraft through mouse clicks and keyboard commands serving as the primary metric of mechanical skill in competitive play.

[01:45] Christopher David: What's the equivalent of that for agents?

[01:49] Christopher David: So I think for now we're going to draw a little parallel between StarCraft APM and the APM of our agents and that APM would be what?

[02:06] Christopher David: Messages to and from the agent and tool calls.

[02:10] Christopher David: We'll start with that.

[02:11] Christopher David: We can sophisticate it over time,

[02:12] Christopher David: blah, blah, blah.

[02:15] Christopher David: Cloud code also has some built-in telemetry support.

[02:20] Christopher David: So we're gonna start sophisticating kind of the analytics here. And so um anyway, just took a very quick and dirty fast pass at this.

[02:33] Christopher David: Just give me the actions per minute

[02:36] Christopher David: of me.

[02:37] Christopher David: analyzing all of the cloud cloud code conversations because it's all on my desktop saved all of the two months of conversations that I've had just go and analyze it what's my APM over the last hour six hours day week month lifetime I have a lifetime APM of 2.298 any of you are welcome to download the towery app and run it yourself let's see

[03:05] Christopher David: So here's our baseline,

[03:08] Christopher David: and we're going to try to improve this.

[03:11] Christopher David: By the way,

[03:12] Christopher David: check this out.

[03:13] Christopher David: Let's open up a chat window.

[03:15] Christopher David: Ooh.

[03:18] Christopher David: Yeah.

[03:19] Christopher David: Oh, another chat window?

[03:20] Christopher David: Oh, yeah.

[03:20] Christopher David: Oh, yeah.

[03:23] Christopher David: OmarK,

[03:23] Christopher David: eat your heart out.

[03:25] Christopher David: All right. Any feedback on APM?

[03:31] Christopher David: measurement benchmark stuff let us know we are going to keep a dock of apm on our github repo that'll be at open agents inc slash open agents docks there'll be something in here that says apm.md and we'll keep that as sort of a running spec of how we're measuring a apm there could be different versions of that effective different sorts of vetting but uh yeah let's give it a try

[03:58] Christopher David: See you soon.