The foxo Papers · Vol. I

The Complicated Agentic Matrix

An agentic operation is a system of two flows — energy and information. Regulate the flows, and the complex becomes merely complicated.
Author FoxSoft Published July 2026 Version 2.2 Series The foxo Papers License CC BY 4.0
Abstract

The dominant explanation for why multi-agent AI systems fail — that the models are not yet capable enough — is at best incomplete. In the largest empirical study to date, benchmarked open-source multi-agent frameworks failed on 41–87% of tasks, and on the human-labelled subset roughly four in five failures traced to specification and inter-agent coordination rather than to the base model.1 We argue that an agentic operation is best understood not as a collection of intelligences but as a regulated system of two governing flows: energy — the compute and capital it dissipates — and information — the data that coordinates it. Reliability, cost, and viability are largely properties of these flows, not of the agents. Drawing on Prigogine's dissipative structures, Shannon's information theory, Ashby's law of requisite variety and the Conant–Ashby good-regulator theorem, and Landauer's principle — which couples the two flows by pricing every bit of computation in energy — we give an objective, quantitative account of when such a system holds together and when it collapses. A raw swarm is complex in the precise sense of the Cynefin framework: emergent, non-linear, ungovernable by specification. It becomes governable only when a control plane meters its energy flow (holding return-on-energy above unity) and routes its information flow (keeping the coordination channel ahead of the variety it must absorb). Such a plane does not remove the complexity; it renders the operation complicated — measurable, bounded, and steerable by a single operator. This is the foundational thesis of the foxo grid.

How to read the claims in this paper
established peer-reviewed result.   model our definitional / derived model.   approximation an operational proxy, not an exact law.   hypothesis a foxo conjecture we hold as testable, not proven.

01Operations as a flow system

Strip away the mystique of the word "agent" and an operation is a physical process. It draws energy from its environment, moves information internally, produces value, and exports waste. A software company run by agents is not, at the level that determines whether it works, a cluster of minds; it is a flow system — and flow systems are governed by conservation laws and channel limits, not by the cleverness of any single component.

This paper takes that reframing literally. We model an agentic operation through two governing flows: energy — the compute, tokens, and capital it dissipates — and information — the data that keeps its parts coherent. Other things matter — permissions, latency, trust — but we claim they enter the fate of the operation through these two flows, and that node intelligence is a term inside them rather than the binding constraint. The binding constraints are thermodynamic and informational. We characterize each flow quantitatively, show that they are coupled by the physics of computation, and then show that a control plane is nothing more or less than a regulator of the two flows.

The first evidence is the failure data. The 2025 MAST taxonomy (Cemri et al., UC Berkeley and collaborators)1 analysed 1,642 execution traces across seven open-source multi-agent frameworks and found task-failure rates of 41% to 86.7%. On its human-labelled subset, ~79% of failures fell under specification / system-design and inter-agent misalignment rather than the base model.established One honest boundary must travel with every citation of this number: these are benchmark tasks run on research frameworks, not production incident rates, and the authors expressly forbid comparing the frameworks against one another. What the data supports is narrower but sufficient for our argument: failure in these systems is dominated by how the agents are organized, not by how capable each one is. Read through the flow lens, these are not cognitive failures; they are flow pathologies — energy misdirected or unmetered, information lost or uncontrolled. Making the node smarter cannot fix a flow.

Data · where multi-agent systems fail (MAST, 2025)
Failure rate across 7 frameworks (benchmark tasks)41.0 – 86.7%
System-design / specification failures (human subset)41.8%
Inter-agent misalignment (human subset)36.9%
Verification / termination failures (human subset)21.3%
Gain from a role-specification fix (ChatDev)+9.4%
Gain from adding a verification step (ProgramDev)+15.6%
Cemri et al., Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657 v3 (2025). Benchmark frameworks, not production; per-framework rates are not comparable per the authors. Interventions helped but "task completion rates still remain low."

02The failure is structural: complex, not complicated

Two words, routinely confused, mark the difference between a system you can govern and one you cannot. The Cynefin framework2 draws the line. A complicated system has many parts but is knowable: it decomposes, its cause-and-effect can be recovered by analysis, good practice can be codified (a jet engine). A complex system has many interacting, adaptive parts whose behaviour is non-linear and emergent; the whole exceeds the sum, small perturbations cascade, and cause and effect are visible only in retrospect (an economy, an ecosystem). Complex systems cannot be commanded — only sensed and steered.

A population of autonomous agents is unambiguously complex: each adapts to the others' outputs, the coupling is dense, and consensus and contagion dynamics emerge that no agent intends. The founding error of the "bag of agents" is to treat this complex object as complicated — to believe a sufficiently detailed specification can make the swarm correct. MAST's fourteen failure modes1 read, one by one, as flow pathologies: conflicting objectives and role unclarity are misdirected energy; communication breakdown and state desynchronization are information-flow failure; missing verification is unmeasured flow. The reliability problem is structural, and its structure is the two flows.

COMPLEX emergent · sense & respond COMPLICATED knowable · analyze & plan CHAOTIC CLEAR agent matrix governed matrix + flow regulation
Figure 1. The Cynefin domains. A raw agent matrix sits in the complex quadrant — ungovernable by specification. Regulating its two flows does not simplify it; it moves the operator's experience of it across the boundary into the complicated, where it is knowable and steerable.

03The first flow: energy

Every action an agent takes dissipates something real — compute cycles, tokens, and ultimately capital. Prigogine's account of thermodynamics is our lens here, not a claim that an agent literally obeys the heat equations: Ilya Prigogine's Nobel-winning work on dissipative structures established that a system can hold order far from equilibrium only by continuously dissipating energy drawn from its environment10 — order sustained through flux, "order through fluctuations."established A whirlpool, a flame, a living cell: each keeps its form only while energy passes through it, and dissolves the instant the flow stops.

An agentic operation is such a structure. Its order — coherent, goal-directed work across a shifting lattice of agents — is not a resting state but a pattern held up by a continuous flux of compute and capital. We collapse that flux into a single cost rate: let Φin be the fully-loaded operating cost per unit time — tokens, GPU-seconds, tool and API calls, and human-review time, each priced into a common currency — and Vout the rate at which the system exports value in that same currency. So defined, both sides share units and their ratio is dimensionless. The defining question of viability is not a threshold on either quantity but their ratio — a return-on-energy (more plainly, a return on spend):

(1)ρ ≡ Vout / Φin      thrives if ρ > 1,   breaks even at ρ = 1,   dies if ρ < 1  model

ρ is the objective, measurable metric of an autonomous operation: a system with ρ > 1 pays for its own order and grows; a system with ρ < 1 is consuming itself. It is, deliberately, ordinary unit economics — value over fully-loaded cost — given a physical reading. The threshold ρ = 1 is meaningful only when numerator and denominator are booked in the same currency and all costs are loaded; a value score divided by a dollar cost has no break-even at one. This also reframes the budget: the constraint is not a ceiling imposed from outside but a metabolic rate — a cap on the cost flux the structure may draw. The operator's problem is to find a policy that maximizes long-run value under that flux:

(2)π* = arg maxπ  𝔼[ Σt γt r(st, at) ]    s.t.   dE/dt ≤ Φmax,   at ∈ 𝒜safe

What Schrödinger called negentropy — a living system feeds on order and exports disorder15 — is here made operational: the matrix imports low-entropy energy and structured intent, and exports value plus waste heat. To keep ρ > 1 is to stay alive.

Measuring valueout: the dual metric

ρ presupposes that Vout can be measured — which is trivial once revenue flows and subtle before it. Yet a pre-revenue operation still produces value: it builds a pipeline of prospective customers, and a pipeline is a portfolio of options whose expected value is positive even when out-of-the-money. We therefore split the numerator into a realized and an expected part, and weight the expected part by a confidence that decays with neglect:

Vout(t) = rev(t) + d/dt Σleads c0(stage) · e−λ Δτ · ACV · e−r tclose  model

The pipeline form shown is one domain adapter — a sales operation's value function; a support, engineering, or research operation would substitute its own verified-outcome measure, and only the confidence-decay discipline is meant to be universal. Here Δτ is the time since a lead's last real engagement. Two disciplines keep this honest. First, both ratios are reported together — ρrealized (proven) beside ρexpected (believed, confidence-banded) — so the distance between belief and proof is never hidden. Second, the decay term makes the value model itself a dissipative structure: pipeline value is sustained only under continuous engagement and bleeds away under silence. A system therefore cannot manufacture ρ > 1 on a dead pipeline — starved of engagement, ρexpected decays back to the truth. Value, like order, must be paid for continuously.

04The second flow: information

The second flow is information, and it needs a substrate. Coordination is, at bottom, the reconciliation of state across agents — information that must move. Two quantitative facts govern it.

Integrity decays with unverified length. Consider the most naive composition: a chain of N agents, each individually correct with probability p, where the task succeeds only if every step does. End-to-end reliability is

(3)R(N) = pN  (independent steps, no verification)  model

At a generous p = 0.95, twenty steps succeed only 0.952036% of the time. This is deliberately the worst case — it assumes independent, unverified steps with no retries or recovery; correlated failures can be worse, and verification can be much better. Its point is not to predict a rate but to show the shape: unverified length degrades reliability exponentially, so the architectural job is to insert verification and restoration before N grows. Reintroduce a per-step verify-and-correct with detection d and recovery r and each factor becomes p + (1−p)·d·r — which is exactly what the control plane's verification step buys, and MAST measured a +15.6% task-success gain from adding one such step.1

Coordination cost is a channel problem. If every agent must reconcile state pairwise, the number of integration paths scales as

(4)Cpair = N(N−1)/2 = O(N2)   vs.   Cbus = N interfaces

The architectural answer is a shared communication fabric — a bus, or in its classical AI form a blackboard11: agents publish to and read from a common, structured workspace while a controller watches it and activates agents as their preconditions are met. What collapses from O(N²) to O(N) is the number of interfaces each agent maintains — it now integrates once, with the bus, instead of with every peer. The message traffic does not fall for free: if every agent still consumes everything every other agent posts, delivery volume stays quadratic and a naive shared workspace simply floods each agent's context.11 Holding traffic near-linear takes real routing — topics, subscriptions, aggregation, and back-pressure — layered on the bus. The bus makes cheap coordination possible; routing policy makes it actual.

And a channel has finite capacity. Let Cbus be its throughput (bits per unit time) and H(D) the entropy rate of the disturbances the operation must absorb. Ashby's law of requisite variety — given a rigorous information-theoretic form by Touchette and Lloyd, who show a controller acts as a communication channel and that the entropy it can remove from a system is bounded by the information it gathers about it316 — gives an operational bound on the information flow:

(5)Cbus ≥ H(D)   — else control is lost  approximation

We state (5) as an operating heuristic, not a measured law: H(D) for a live business environment cannot be read off directly, and adequate control also depends on latency and on having a response for each disturbance, not on bandwidth alone. Estimating H(D) from telemetry is an open problem we name in §08; in the architecture (Vol. III) the bound is tracked through proxies — utilization, event-latency tails, backlog age, and the rate of disturbances with no matching response policy. The direction is what matters: a controller under-varied in what it can see is as helpless as one under-varied in what it can do, and the bus is where both variety budgets are spent. The Conant–Ashby good-regulator theorem13 motivates the sensing half — effective regulation benefits from carrying an internal model of the system — though, as we discuss in §06, that theorem says less than its slogan and does not by itself force any particular division of labour between machine and human.

the matrix dissipative structure energy Φin compute · capital value Vout entropy · waste data bus Cbus ≥ H(D)
Figure 2. The two flows. Energyin: compute and capital) enters, is dissipated as work, and leaves as value Vout plus waste — viable only while ρ = Voutin > 1. Information circulates on the bus, which must satisfy Cbus ≥ H(D). The control plane regulates both.

05The coupling: information is physical

The two flows are not independent accounts — they are one, joined by the physics of computation. Landauer's principle12 states that erasing a single bit of information dissipates at least

(6)Ebit ≥ kB T ln 2   (≈ 2.9 × 10−21 J at 300 K)  established

The bound is real and now experimentally confirmed — Bérut et al. measured the heat released erasing a single bit and found it approaches the Landauer limit.12 Information is physical: every bit moved or erased has a thermodynamic price. The energy cost of coordination is therefore bounded below by the information it carries,

(7)Ecoord ≈ Ibus · ε   where ε ≥ kB T ln 2  approximation

Two honest caveats sharpen rather than weaken the point — and we make them because the naïve version of this argument is wrong. First, real datacenters run some nine orders of magnitude above the Landauer floor,12 so at operational scale the coupling constant ε is not kBT ln 2 — it is the market price of compute, the ₹ or $ per token. Eq. 7 is thus an engineering cost model, not a derivation from Landauer; the theorem supplies only the sign and the floor. Second, that is precisely why the coupling bites in operations: the constant joining information to energy is large. The direction is exact and unavoidable — every bit of coordination is paid for in compute, and every unit of compute buys a bounded quantity of information.

The consequence unifies the paper. You cannot fix a coordination failure by spending without limit, because the cost flux is capped at Φmax (Eq. 2). You cannot economize compute without economizing information, because coordination cost is Ibus · ε (Eq. 7). Efficient coordination — near-linear traffic rather than quadratic — is therefore identically cost efficiency, which is identically a higher ρ. Reliability, cost, and viability are three views of the same two-flow account.

Data · why the coupling constant matters
Modern computing vs. the Landauer floor~10⁹×
Enterprise generative-AI spend, 2023 → 2025$1.7B → $37B
Orgs attributing >5% of EBIT to AI6%
Inference-cost cut from routing/orchestration (FrugalGPT, RouteLLM)85–98%
Landauer gap: Bérut et al., Nature 484 (2012). Spend and EBIT: industry surveys, 2025. FrugalGPT (Chen et al., 2023); RouteLLM (Ong et al., 2024). Spend is scaling far faster than measured value, and unit economics turn on orchestration — not model price alone — which is exactly what ρ is built to expose.

06The control plane as a two-flow regulator

Everything above resolves into a single design object. A control plane governs an agentic matrix by regulating its two flows — nothing more exotic, and nothing less. Its four properties are four acts of flow regulation.

1 — Observability: sense both flows. The plane measures the energy account (Φin, Vout, ρ) and reads the information state off the bus. In the spirit of Conant–Ashby, this is the regulator building its model of the matrix13; the operator's pulse is that model, rendered legible.

2 — Objective alignment: direct the energy toward value. Every agent optimizes one shared, budget-bounded objective (Eq. 2) rather than a private script, so divergent local effort sums to a coherent gradient that maximizes ρ. The MAST failure of "conflicting objectives" is, in flow terms, energy pulling in different directions; a shared gradient removes it.

3 — The escalation boundary: bound the excursions. Actions that are irreversible or low-confidence — those with large, unrecoverable consequences in either flow — are gated to a human:

(8)act(a) ⇔ p( ok | a ) ≥ τ  ∧  reversible(a);   otherwisehuman  model

This bounds the blast radius of emergence. Any model the regulator builds is a compression of a complex system and will miss states it did not anticipate, so bounded autonomy is prudent on those grounds alone. But we do not rest the human boundary on a theorem: the deeper reason a human holds the irreversible actions is normative — accountability, legal exposure, and the values encoded in the objective are not the machine's to redefine — a point we develop in Vol. II. The escalation gate is where that principle is enforced.

4 — Structured coordination: route the information. Coordination follows a topology — orchestrators, leads, agents, sub-agents over a shared bus — that keeps interfaces linear in N and, with routing policy, holds traffic near-linear within capacity Cbus: honouring the requisite-variety bound (Eq. 5) and keeping the compute cost of coordination (Eq. 7) sub-quadratic. The lattice grows to the work and prunes when the work is done.

Together the plane commands requisite variety over both currents: it meters energy — holding ρ > 1 under a flux cap — and routes information — keeping Cbus ahead of H(D). That pair of measurable conditions is the core of what we mean by "governance" — accountability, permissions, and audit sit alongside it — and satisfying them is what carries the operation across the Cynefin line: the matrix stays complex; the operation becomes complicated.

07Isolation as flow containment

One property is a precondition rather than an optimization. When a control plane governs many matrices at once — many products, many tenants — each matrix's two flows must be sealed. If energy or information couples across tenants, failure propagates along the coupling and ρ becomes unaccountable: you can no longer say which operation paid for which value, nor contain one matrix's emergence from another's. Strict isolation — row-level at the data layer, and ideally host-level beneath it — is therefore flow containment: the boundary condition that keeps each matrix's complexity local. Governing many complex systems at once is possible only if each is sealed from the rest. Isolation is the license to operate at scale.

08Discussion

The reframe redirects the field's reliability roadmap. If most multi-agent failures are coordination rather than cognition1, then investment aimed only at stronger nodes is mis-aimed; the larger and more durable gains are in flow regulation — and flow regulation compounds across model generations rather than resetting with each one.hypothesis This is our central conjecture, and it is falsifiable: a well-regulated two-flow system running a modest model should beat an unregulated swarm running a frontier one, on cost per verified outcome, and should hold that edge as both improve. We state it as a claim to be tested, not a result we have proven.

It also gives the operation a single north-star metric — ρ — and a short list of open problems we do not pretend to have solved: estimating ρ and the disturbance entropy H(D) from live telemetry; learning the escalation threshold τ from outcomes rather than setting it by hand; and measuring how much predictive power the regulator's model actually loses in compression. The operator's job reduces to three levers: keep ρ > 1, keep the channel ahead of the variety, and tune τ. Steering, not doing.

What this paper is and is not. It is a systems framework and a set of measurable conditions, not a proof and not a benchmark. The physics and cybernetics are used as lenses that yield operational quantities (ρ, the requisite-variety bound, the escalation gate); where a claim is our conjecture rather than an established result, it is tagged hypothesis above. The empirical anchor — the MAST failure data — comes from research frameworks on benchmarks, not production; the strongest validation, a live operation's measured ρ over time, is work in progress and is not reported here. We would rather state that plainly than dress a framework as a finished science.

09Conclusion

An agentic operation is two flows — energy and information — coupled by the physics of computation. Its failures are flow pathologies, not deficits of intelligence, which is why a smarter agent does not save a swarm that fails four-to-eight times in ten. Govern the flows — meter the energy so return-on-energy stays above unity, route the information so the channel stays above requisite variety, and gate the irreversible to a human — and a complex, ungovernable matrix becomes a complicated, steerable operation. That two-flow regulator is the grid. Volume II asks how a system held above ρ = 1 does not merely run but lives and thrives; Volume III specifies its architecture.

§References

  1. [1] A. Cemri, et al. "Why Do Multi-Agent LLM Systems Fail?" arXiv:2503.13657 v3, 2025 (NeurIPS 2025). (MAST — the Multi-Agent System Failure Taxonomy; 1,642 traces, 7 open-source frameworks, benchmark tasks.)
  2. [2] D. J. Snowden and M. E. Boone. "A Leader's Framework for Decision Making." Harvard Business Review, Nov. 2007. (The Cynefin framework.)
  3. [3] W. R. Ashby. An Introduction to Cybernetics. Chapman & Hall, 1956. (The Law of Requisite Variety.)
  4. [4] S. Beer. Brain of the Firm. Allen Lane, 1972. (The Viable System Model.)
  5. [5] J. H. Holland. "Complex Adaptive Systems." Daedalus, vol. 121, no. 1, 1992.
  6. [6] Anthropic. "How we built our multi-agent research system." anthropic.com/engineering, 2025. (Multi-agent beat single-agent by 90.2% on an internal eval, but ~80% of the variance is token usage and the system costs ~15× the tokens — first-party, non-peer-reviewed.)
  7. [7] H. Maturana and F. Varela. Autopoiesis and Cognition: The Realization of the Living. Reidel, 1980.
  8. [8] S. Casper, et al. "The 2025 AI Agent Index." MIT, aiagentindex.mit.edu, 2025.
  9. [9] L. Chen, M. Zaharia, J. Zou. "FrugalGPT." arXiv:2305.05176, 2023; I. Ong, et al. "RouteLLM." arXiv:2406.18665, 2024. (Orchestration/routing cuts inference cost 85–98% without proportional quality loss.)
  10. [10] I. Prigogine and I. Stengers. Order Out of Chaos. Bantam, 1984. (Dissipative structures; Nobel Prize in Chemistry, 1977.)
  11. [11] L. D. Erman, et al. "The Hearsay-II Speech-Understanding System." ACM Computing Surveys, 1980. On the modern shared-workspace caveat (a single blackboard floods agent context): "Terrarium" and related blackboard/event-bus multi-agent work, 2025.
  12. [12] R. Landauer. "Irreversibility and Heat Generation in the Computing Process." IBM J. Res. Dev., 1961. Experimental confirmation: A. Bérut, et al. "Experimental verification of Landauer's principle." Nature 483, 187–189, 2012.
  13. [13] R. C. Conant and W. R. Ashby. "Every Good Regulator of a System Must Be a Model of That System." International Journal of Systems Science, 1970. (Motivates the internal-model requirement; its scope is narrower than the popular slogan.)
  14. [14] C. E. Shannon. "A Mathematical Theory of Communication." Bell System Technical Journal, 1948.
  15. [15] E. Schrödinger. What Is Life? Cambridge University Press, 1944. (Negentropy.)
  16. [16] H. Touchette and S. Lloyd. "Information-theoretic approach to the study of control systems." Physica A 331, 140–172, 2004. (Requisite variety made rigorous: the entropy a controller can remove is bounded by the information it gathers.)