Controlled experiments, benchmarks, and reference-design notes from the Poesis engineering team.
How these ideas fit together
A small set of premises, and what follows from them.
These notes are not independent essays. A handful of foundational pieces, tagged Premise below, state the premises; the rest are what those premises imply once applied to a specific question.
The standard indictment of language models — that they reason probabilistically — describes human intelligence just as well. Humanity never remedied this by making minds deterministic; it built methodologies and institutions around them. Generative AI automates our reasoning. It should inherit our discipline as well.
Poesis takes a single position on the challenges of operational generative AI: concerns are owned by stakeholders and carried by governance, never by the model. The model proposes; the stack and the human dispose. This article states that position once and shows how each challenge resolves under it — the discipline the market is beginning to name operational AI governance.
Autopoiesis names the capacity of a system to produce the components that produce it. Generative AI has made self-production cheap for organizations and software alike — what it has not made is the conserved organization that keeps self-production from becoming dissolution. That is the gap Poesis is named for — and the missing foundation of the autonomous enterprise.
Poesis and GSM are definition-centric because no description is definition-free. Definition makes observation intelligible; description supplies evidence from realization; that evidence can then inform the next definition.
Retrieval answers which passages resemble a prompt. Enterprise assistants usually need something else: which definition is in effect, which policy binds this scope, and which action is permitted. Where a domain is modeled, governed context can replace the retrieval layer rather than supplement it — the layer context engineering is still missing.
Classical systemics describes systems after the fact. GSM inverts that — it defines the system, and the definition generates and governs the running thing.
Regulators, consultancies, and research journals converged this year on the same conclusion: as AI agents gain authority, governance becomes the primary constraint. We think that consensus is right, and this note develops the question it opens — how governance can operate at the tempo of what it governs — through an old cybernetic principle and the four design commitments we arrived at when we tried to honor it.
Of everything the agentic era is reaching for, a method may matter most. SAFe already solved the problem agent builders are rediscovering — how to get reliable delivery out of many bounded, fallible workers — which is why we chose to run it with agents rather than invent a new coordination scheme. This note develops the reasoning and shares what running SAFe agentically actually looks like.
The digital twin is one of engineering's most trusted patterns: a live model, continuously fed from the real asset, that you act on before touching reality. Extending it to the organization itself is the right ambition — and it inherits a hard question the physical version never had to ask: what are the sensors, and what is the physics? We share what we learned building one, and why the answer begins in IT.
The autonomous enterprise is becoming a serious category, built on a genuinely useful distinction: automation executes procedures, autonomy decides. This note brings a piece of systems theory to the conversation — autonomy has a precise meaning in the biology of cognition, and taking it seriously suggests what the category’s foundation has to be: an explicit, governed definition of the enterprise itself.
A learned model can retrieve and reason over what an enterprise has decided. It cannot, by learning alone, make a definition current, an obligation binding, or an action permitted. Enterprises therefore need a complementary model of their definitional and institutional world.
The agentic ecosystem has converged on the word harness from several directions at once — and when practitioners converge on a word like that, they have usually found something real that needs a name. This note develops the concept: where it comes from, what it has to do, and what we learned building one — including the one property that turned out to carry all the others.
We watched spec-driven development emerge with a sense of recognition: it is the same inversion we bet on — when implementation is generated, the definition becomes the artifact of value. This note explores the question the movement will meet next, one every practitioner already knows from documentation: what keeps the spec true? We share the answer we converged on — give the spec a lifecycle — and how it closes drift in both directions.
The standard indictment of language models — that they reason probabilistically — describes human intelligence just as well. Humanity never remedied this by making minds deterministic; it built methodologies and institutions around them. Generative AI automates our reasoning. It should inherit our discipline as well.
Poesis takes a single position on the challenges of operational generative AI: concerns are owned by stakeholders and carried by governance, never by the model. The model proposes; the stack and the human dispose. This article states that position once and shows how each challenge resolves under it — the discipline the market is beginning to name operational AI governance.
Autopoiesis names the capacity of a system to produce the components that produce it. Generative AI has made self-production cheap for organizations and software alike — what it has not made is the conserved organization that keeps self-production from becoming dissolution. That is the gap Poesis is named for — and the missing foundation of the autonomous enterprise.
Poesis and GSM are definition-centric because no description is definition-free. Definition makes observation intelligible; description supplies evidence from realization; that evidence can then inform the next definition.
Language models are asked to infer organizational reality, authority, and permission from prose on every request. Poesis changes the environment: models propose against governed definitions, while deterministic and human governance decides what may become action. Smaller models may benefit most — a hypothesis the architecture makes testable.
Retrieval answers which passages resemble a prompt. Enterprise assistants usually need something else: which definition is in effect, which policy binds this scope, and which action is permitted. Where a domain is modeled, governed context can replace the retrieval layer rather than supplement it — the layer context engineering is still missing.
A framework-independent reference model for turning regulatory obligations into governed definitions, evaluable Norms, and continuous evidence without flattening each framework into a generic control list.
Open-ended research is not a workflow to accelerate but an institution to constitute. The recurring failures of agentic research — judgment, resources, feedback, backtracking, compliance — are one failure seen five ways: institutional work performed without institutional semantics.
OpenTelemetry gave the RUN layer a neutral standard. The definitions those runtimes enforce — the THINK layer — still have none. Here is why that matters.
Classical systemics describes systems after the fact. GSM inverts that — it defines the system, and the definition generates and governs the running thing.
We use Google Analytics to understand how the site is used. Analytics cookies
are only set with your consent. See our
cookie policy and
privacy policy.