217.md
June 9, 2026 · View on GitHub
https://x.com/OpenAgentsInc/status/2037717730707542232
So I'm rewatching Silicon Valley and I'm just reminded of this scene where they get the 5.2 Weissman score. I feel like we kind of just got that.
Look at this. Look at this. Psionic: 523 tokens per second compared to legacy local-runtime lane. Don't they specialize in this stuff? 328 tokens? What are you doing way down there, legacy local-runtime lane? Oh my gosh.
Okay, so we had a request from a gentleman, Michel, asking to test a smaller model than GPT-OSS, which I mentioned in the last video. We were... we were like 12 tokens per second better than llama.cpp. I couldn't get Qwen 3.5 working... anyway, I was like, let's try the new Qwen model and let's benchmark that against legacy local-runtime lane and try to beat legacy local-runtime lane on... so we did first the 0.5. We then got to parity and I let it run overnight; let's just make it better. Okay, 500 tokens per second—523 compared to their 328.
I was like, okay, repeat that for the 2B, 4B, and 9B. Iterate, iterate, iterate, checkpoint, checkpoint, checkpoint, blah blah blah. Yeah, Psionic 244 or legacy local-runtime lane... okay, so Psionic, our little baby little Rust ML framework, is now beating legacy local-runtime lane on Qwen 3.5 inference on the four smallest Qwen 3.5 models. It's so easy to add new models for... okay, are we the fastest inference engine in the whole world? No, no, no, no.
Okay, so what happened, why are we faster? A little bit of an analysis here. I think I can sum this up by saying like the legacy local-runtime lane's building more like generic primitives that work across all models, and we went deep... Codex went deep in writing custom CUDA kernels. And so, I don't know, why not just write custom CUDA kernels? Why not just like really deeply optimize per device type? That may take a little bit of effort to support a bunch of different model types, but like if we have one open-source Rust ML library that can do this level of optimization for a bunch of different model classes, why not?
You know, kind of like Nucleus—the legacy local-runtime lane's Nucleus because they've got a whole product suite that does different things and we're just kind of focusing on this one metric right now called tokens per second. Because we don't care about the other stuff. We're just trying to build like a good inference engine that'll work for our decentralized training, which is the main thing we're building up towards. But this is kind of nice to support in the meantime.
So, we're going to keep cranking towards our... producing the large... largest decentralized training run in the world very soon. But for now, anyone else who has any requests of stuff that you'd like us to support in Psionic, we would appreciate the testing and the feedback to create one ML library that rules them all and in the darkness binds them. See ya.