This page examines one specific frontier AI safety concern: architectural embedding as a possible source of frontier-level capability. The question is whether an existing AI core, embedded within a larger functional architecture, could achieve dramatically increased capabilities.
By "architectural embedding," here I don't mean architecture embedded in the AI – safety layers, guardrails, or other structural features built into the model itself. That's a real and important area of work, but it's not what I am talking about.
I also don't mean an agentic architecture built around an AI – a model wired up to tools, memory, email, calendars, or task workflows. That kind of setup extends an AI's reach: it can now touch more systems, complete more tasks, act in more places. But the intelligence doing the acting hasn't changed. It's the same core, just with a longer arm.
What I mean is a third thing: an architecture that changes the AI's functional role itself – not by giving it more to do, but by changing what kind of intelligence it is, so that the system as a whole reasons at a level the core model alone never would.
Not architecture in the model. Not tools around the model. An architecture that changes what the model's intelligence is.
Just like with other tools, sometimes a single detail can magnify power immensely. A mixture of saltpeter, sulfur, and charcoal results in a very hot, bright, and rapidly burning fire with a hissing sound and beautiful sparks, but with a single functional detail – when it is contained within an enclosure – it became the basis for powerful armaments and a mining tool that inspired the invention of the more powerful explosives used today. As another example, steel rods can be used for hitting, poking, and prying at things, but when shaped into circles, given spokes and an axle, they become wheels. Current AI capability likewise sees dramatic increases in capability if treated as a core embedded within an effective functional structure.
The BAL-looping framework as presented on this website is one candidate functional architecture relevant to this question. Through philosophical discussions and scientific exploration, the two protagonists propose that the human brain works on cybernetic principles involving a core consisting of an inner model of the environment and goal-seeking behavior. They then discuss how, in the human's case, this is leveraged by subjective experience, which in turn is generated by an exaptative loop through the expressive and receptive communication channels.
Readers interested in the architectural question can begin with dialogs 1 and 2, or with the BAL-looping scientific papers.
As AI systems become more advanced, it’s crucial to stay ahead of the curve. It won't be long before AI designers begin using these cybernetic functional architectures, or alternatives that can likewise boost an AI core dramatically. That would add great effectiveness, but coupled with risks – just like any powerful tool. For readers interested in the architectural proposal itself, the AI Safety Brief page provides a concise overview of the BAL-looping framework and its potential relevance to AI.
One way to start is to read the Prelude to the Seven Dialogues – a thoughtful introduction that gives an overview of the BAL-looping framework in everyday accessible language, over drinks at the bar.