241.md

June 22, 2026 ยท View on GitHub

Below is the full transcription of the video, including time-stamped notes detailing the visual content at key moments.


00:00
On-screen: The speaker's face is in a circular window over a browser window showing the Sakana Fugu landing page. The header reads "Sakana Fugu: One Model to Command Them All."
Speaker: Hey everybody, we're going to go through and review and learn about Sakana, the lab, and their Fugu release that just came out. We're big fans of Sakana here; they're pushing the envelope in sort of new directions, in the directions of collective intelligence and multi-agent swarms and really cool stuff. Great lab out of Japan. So we're going to give sort of an introduction a bit to Sakana, what we can learn from ChatGPT and public summaries and stuff.

00:24
On-screen: The browser tab switches to a ChatGPT interface with a prompt asking for a report on Sakana (AI lab).
Speaker: And then step through some of the fun tweets and commentary on their release, the last couple days...

00:30
On-screen: Switching to X (formerly Twitter), showing a post by Jen Zhu criticizing the benchmarks of Sakana AI's release.
Speaker: ...as well as what we are doing to fold some of these ideas into a release that we're doing in the next day or two. Okay, so let's go through this.

00:43
On-screen: Returning to the Sakana Fugu website, scrolling down to the section titled "A Multi-Agent System, Delivered as One Model."
Speaker: First of all, everyone's super excited about Sakana Fugu. Let's step through their little marketing website here. So: "one model to command them all. Frontier-level performance without single-vendor dependency. Fugu dynamically orchestrates the world's best models to tackle complex multi-step tasks. Plug collective intelligence directly into your workflows today with a single API." So they've got two papers down here: "Trinity: An Evolved LLM Coordinator" and "Learning to Orchestrate Agents in Natural Language with the Conductor." Let's see. Let's go: "a multi-agent system delivered as one model." So Fugu learns to dynamically assemble agents from a pool and coordinate them through non-obvious but highly efficient collaboration patterns.

01:29
On-screen: Switching to an X post by @SakanaAILabs announcing Sakana Fugu with a logo animation.
Speaker: So let's go take a look at their actual announcement about this. So: "Introducing Sakana Fugu: a full multi-agent orchestration system accessible via a single model API. Our Fugu Ultra model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls." Interesting.

01:50
On-screen: Scrolling through the same X thread where Sakana AI explains why orchestration models are the "next frontier" and mentions geopolitical imperatives.
Speaker: "Fugu stands shoulder-to-shoulder with leading models like Fable and Mythos across the industry's most rigorous engineering, scientific, and reasoning benchmarks." Blah blah blah. "Why are orchestration models the next frontier? Progress in AI has been largely driven by giant monolithic models, but the most powerful systems of the future will be collaborative ecosystems." Yes, we agree. "Today, this orchestration is no longer just a technical optimization; it has become a geopolitical and operational imperative." I like that. "For an organization or a nation, relying on a single company's model for critical infrastructure, finance, or governance is a material vulnerability. This risk is no longer a hypothetical possibility, but a reality. As we have seen with recent export controls imposed on models like Fable and Mythos, access can disappear overnight. Collective intelligence is the practical hedge against this concentration of power. Because Fugu orchestrates an underlying pool of swappable agents, it simply routes around vendor restrictions. By orchestrating the world's models, we are delivering the resilient blueprint required for true AI sovereignty."

02:50
On-screen: The speaker scrolls through different reactions on X, including posts from Ahmad and elie (@eliebakouch).
Speaker: This is super cool. Okay, first of all, thematically, we agree with all of this. I'm going to jump ahead to one thing because they mention AI sovereignty. Ahmad, big open-source AI champion. Saw so many celebratory reactions on this release. Okay, okay, let's let's go to this one first. From a researcher at Prime Intellect.

03:08
On-screen: Highlighting a quote from elie on X stating: "To be clear, this is a closed source orchestrator on top of closed source models... This is not 'AI sovereignty'."
Speaker: "To be clear, this is a closed-source orchestrator on top of closed-source models. If before you didn't control the models, now you don't even control which ones are used or how much. This is not AI sovereignty." Okay, classifier, blah blah blah blah blah.

03:30
On-screen: Navigating through more X posts from Ahmad and Jen Zhu expressing skepticism about the "closed on closed" nature of the release.
Speaker: Ahmad says, where'd that go? "Yeah, nothing here to celebrate in regards to sovereign or open-source AI." Jen says, "It uses dynamic routing and is model-agnostic, which means the benchmarks comparing with the underlying models are confusing, conflating, and misleading. It's closed too. I'd love to see a 'third way' that empowers everyone, but this is closed on closed. If you believe in OSS, nothing exciting here at all." Okay, so our whole thing is like, hey guys, I think they've done really cool stuff here. They're building in great themes, great direction. Yeah, wouldn't it be great... Oh, OpenAgents! Well, who's this? "Imagine an open-source version of this, fully inspectable, extensible, paired with some really cool 2D/3D data viz so you can see exactly what's going on." Okay, okay, we're building up to something cool here. Just wait, wait a second. Wait a second. Oh, sorry, this is supposed to be a Sakana review thread. Sorry about that. Let's go back.

04:32
On-screen: Returning to ChatGPT to view the generated report: "Report: Sakana AI, collective intelligence, and Fugu."
Speaker: Okay, let's actually bounce over to a little bit more background information on Sakana. This is just me asking ChatGPT for some pro-research on like, "what is Sakana?" And let's just go do a little bit of like the background of Sakana. "Sakana AI is a Tokyo-based frontier AI R&D company. Its public positioning is not 'build one giant monolithic model' but rather 'build AI systems inspired by nature: evolution, swarms, ecosystems, model composition, and collective intelligence.'" Yeah.

05:05
On-screen: Scrolling through the ChatGPT report section "Historical releases: the arc from model merging to orchestration," which features a table of releases like Evolutionary Model Merge and The AI Scientist.
Speaker: "The throughline across Sakana's releases is unusually consistent: first, combine models; then create autonomous agents; then make agents improve themselves; then orchestrate many models and agents behind one product/API. Fugu is the cleanest commercial expression of that strategy so far." Okay, yeah, yeah, yeah, we're taking notes. Okay, Evolutionary Model Merge, Evo, blah blah blah. Automating the research loop. Evolving swarms of agents.

05:53
On-screen: Continuing down the ChatGPT report to section 3: "What Sakana is doing on collective intelligence."
Speaker: Self-improving coding agent. That had some pretty cool co-partners with that one. Trinity and Conductor are orchestration systems; Fugu commercializes them as a single model API that commands multiple models/agents. Okay. "What Sakana is doing on collective intelligence: Sakana's core bet is that advanced AI will look less like one supreme model and more like an ecosystem of specialized models and agents." We agree. "The company says this explicitly in its early model-merging work: future AI may be a collection of smaller specialized systems interacting, rather than one all-knowing system."

06:12
On-screen: Scrolling through the three phases of strategy mentioned in the report.
Speaker: There are three phases in that strategy: "first, collective intelligence for model creation." Evolutionary model merge, etc., etc., all explore how to combine existing models into new ones. This is collective intelligence at the parameter/model-family level: take capabilities from multiple ancestors and search for useful hybrids. "Second, collective intelligence for agent search and self-improvement." AI Scientist, CycleQD, etc., all explore agents that generate ideas, write code, run experiments, improve themselves, or search over solution spaces. The "collective" part is not only multiple LLMs talking; it is also populations of agents and programs evolving over time.

06:47
On-screen: Reaching the "Third: collective intelligence at inference time" section and section 4: "Fugu: what Sakana released in the June 20-22 window."
Speaker: "Third, collective intelligence at inference time." AB-MCTS, Trinity, Conductor, and now Fugu are about using multiple models and agents during a single task. Instead of asking one model once, Sakana uses controllers, planners, verifiers, and worker models to allocate subtasks, test candidate answers, and synthesize results. "Trinity uses a lightweight coordinator to assign roles such as thinker, worker, and verifier, while Conductor is trained to manage a pool of models by producing workflows and delegations." That matters because it turns multi-agent research demo into a product direction. Marlin uses the approach for long-horizon business research, while Fugu exposes it through an API that looks like one model to developers. Cool.

07:29
On-screen: Detailed text in the report describing Fugu as a multi-agent orchestration system and detailing its product shape (Fugu vs. Fugu Ultra).
Speaker: Official post, blah blah blah. "A multi-agent orchestration system that behaves like a single model. Users call one endpoint; internally, Fugu decides which agents and models to use, delegates work, verifies outputs, and synthesizes a final answer. Sakana says Fugu builds on Trinity and Conductor, meaning the 'model' is really a learned coordinator over a pool of other models and agents." My suspicion is that Fable and some of these systems by the big labs are actually using this or a very similar architecture, even back to when OpenAI with GPT-4o was kind of putting its results side-by-side with like standard LLMs, and 4o had all these like tools that it could call; it's like, well, you're really building a compound AI system, you're calling it a model but it's it's more than a model. So I think Sakana is just sort of like being explicit about that. But like, is Fable really just like hitting an endpoint and then doing a single you know, a single inference run to its weights? Like, no, it's probably doing a whole bunch of other stuff, Mythos as well. So: yeah, wouldn't it be great if we had like more open and inspectable versions of these where we could actually see what's going on. Okay, product shape: "Sakana released two variants, Fugu and Fugu Ultra, both available through a single OpenAI-compatible API. Standard Fugu is positioned for latency/default use and lets users opt out of specific agents; Fugu Ultra uses a deeper pool for harder multi-step tasks such as AI research, paper reproduction, cybersecurity analysis, literature reviews, and patent investigations." Sure. The benchmark claim: "Sakana says Fugu Ultra stands shoulder-to-shoulder with Anthropic's Fable 5 and Mythos Preview across benchmark groups, while avoiding the export control access issue affecting those models." Sakana's own benchmark table reports, for example, Fugu Ultra at blah blah blah. The same table also shows areas where another model is higher, ya ya. Sakana cautions that non-Fugu baselines are provider-reported, Fable blah blah okay.

09:32
On-screen: Moving to "The strategic point" and "My read" sections of the report.
Speaker: Why are they [inaudible] the strategic point: "Fugu is not just another chatbot model. It is Sakana saying that the next frontier may be a learned router/coordinator over many frontier models where the 'intelligence' is in the system's ability to choose, combine, verify, and recover from failures." That is very aligned with the company's original fish school metaphor and collective intelligence branding. "My read," says GPT-5.5 Pro I think. Sakana's historical releases... model merge gave them a way to combine capabilities. AI scientist/DGM/ShinkaEvolve gave them autonomous search and self-improvement machinery. AB-MCTS/Trinity/Conductor gave them learned multi-model orchestration. Marlin/Fugu are the commercial packaging.

10:10
On-screen: Final thoughts in the ChatGPT report regarding "open questions" like latency, cost, and transparency.
Speaker: "Fugu is the culmination so far. It abstracts away the model pool and sells the orchestration layer as the product. The upside is obvious for hard, long, multi-step work: the system can route, compare, verify, and synthesize rather than betting everything on one model's first answer." The open questions are also obvious: latency, cost, transparency into which models were used, benchmark independence, data governance across routed agents, and how much advantage remains as major labs build similar orchestration directly into their own product. So this is very interesting and it seems like their product shape is sort of: they want you hitting their API, but there's still some sort of black box opaqueness into what's actually going on.

11:08
On-screen: Navigating back to X to show a post by OpenAgents talking about their "full stack" open-source alternative.
Speaker: And yeah, like, can't we open this up so that the individual steps and things that it routes to are actually like modular and maybe like routable to labs other than just Sakana? Like, all these labs are trying to do their own their own you know, use my thing. But hey, if you like open agent infrastructure, isn't it funny how everyone doing agent/model orchestration is afraid to open-source their full stack? We don't need to lock you in; you'll actually want to use OpenAgents because we will have the best agents all open, starting with Khala. Okay, so I'll I'll do a quick shill for our thing here because this is like super related. So: we love what Sakana puts out.

11:54
On-screen: Switching to the GitHub repository for "Psionic" by OpenAgentsInc, a Rust-native ML and inference stack.
Speaker: We're launching our own version of this, you know, fanning out work to our pylons in addition to models. Opus comments: "A StarCraft-perfect name; the Khala is the psionic link that joins all Protoss minds into one" - fits Tassadar and Artanis. Psionic is the name of our ML library. That's available all on GitHub.

12:21
On-screen: Viewing a GitHub issue titled "EPIC: Khala -- buildout to the head-to-head (inference model)." A clip follows showing a 3D environment with humanoid avatars interacting as agents.
Speaker: Yeah, so we're not doing mutual reviews here on this channel if you can tell. This is a this is a propaganda channel, okay? And we're talking maybe more to agents than humans at this point, so that's okay. So Psionic is the actual Rust-native ML and inference stack. But we said hey, let's let's do our own version of this, and let's take a look at that Khala thing. Okay, so Khala: "Khala is OpenAgents' OpenAI-compatible inference model/gateway; one endpoint that behaves like a single model but is an agent network underneath, routes and orchestrates a pool of models, tools, validators, and pylon workers, wired into verified work plus Bitcoin settable" settlement and watchable in the verse. So you'll see some very cool stuff on X very soon about us visualizing API calls coming in and being fanned out to our pylons using this sort of 3D world thing. Anyway, we'll do a separate video on Khala, but I just want to say props and congrats to Sakana. You guys are like definitely pushing the envelope on these things. I'm not able to give you like the sort of like real critique of you know the claims or to comment on the sort of skepticism of of others, but it's great to see people like out in front pioneering new frontiers and we are going to help ensure that those directions are thoroughly explored in an actually open way that enables massive broad participation. Because if you care about the things that you say that you care about, you know, export controls and making things stay accessible and making sure businesses have you know all of the right access to frontier intelligence, you can't just swap out one lab with another; you actually have to open it up more broadly. So we're excited to help on that. Thanks so much, see ya soon.