EXP_002 / AGENT SYSTEMS
When agents work together.
What actually changes when AI stops being one assistant and becomes a team.
01 The question
What happens when AI stops being a single assistant and becomes a team?
Planners, builders, reviewers, specialized models, shared tools, parallel execution. On paper, an incredibly powerful idea.
In practice, I learned quickly that adding more agents does not automatically create a better system. Sometimes it just creates more context. A lot more context.
02 Context explosion is real
More agents, more context.
One of my earliest orchestration experiments used AION with multiple specialized agents. It worked — and it exposed one of the biggest problems in agentic systems: context grows frighteningly fast.
Agents generate plans. Other agents review them. The original agent reacts to the review. Tool outputs accumulate, files get rediscovered, decisions get repeated. Before long, useful reasoning is buried under millions of tokens of conversation and duplicated context.
Context is not free memory. It is a resource that needs to be engineered.
03 The harness can make or break the system
The model matters. So does everything around it.
A good harness controls what an agent sees, which tools it can use, how work is delegated, how results come back, and how much history needs to survive between steps.
This became especially clear with Pi. Instead of treating the conversation itself as the orchestration layer, I could keep the main agent focused and delegate bounded tasks to specialized subagents — with simpler control over context, providers, responsibilities and execution.
A great model inside a bad harness can be expensive and unreliable. A good harness can make smaller models dramatically more useful.
04 Determinism vs. agency
Guardrails, not instructions.
Make the workflow too deterministic and agents become expensive scripts. Give them complete freedom and they explore endlessly, duplicate work, reinterpret goals or spend enormous amounts of tokens on problems that were already solved.
The sweet spot sits in between: define the objective, constrain the important boundaries, expose the right tools, set clear completion criteria — and let the agent decide how to get there.
Orchestration shouldn’t dictate every step. It should make good decisions easier and bad decisions expensive.
05 Trust agents. Decide with data.
A claim is not evidence.
Trust the agent to execute, but use evidence to decide whether the execution worked: tests, CI pipelines, static analysis, benchmarks, review agents, token usage, runtime, diffs, logs.
The more autonomy I give an agent, the more important observability and objective validation become.
An agent saying “the task is complete” is useful information. A green pipeline is evidence.
“The task is complete.”
UNVERIFIED- TESTS
- CI PIPELINE
- STATIC ANALYSIS
- BENCHMARKS
- REVIEW AGENTS
- TOKEN USAGE
- RUNTIME
- DIFFS
- LOGS
06 Failed experiments were useful too
Frameworks become systems too.
Not every orchestration attempt survived. My experiments with dsh-workflow never became the production workflow I originally imagined.
That failure was valuable. It showed how quickly orchestration frameworks introduce complexity of their own — abstractions, plugins, configuration, coordination rules and hidden assumptions that eventually become another system to maintain.
The best orchestration layer isn’t the one with the most features. It’s the one that stays out of the agent’s way.
The question changed
How can I make several agents work together?
What is the minimum orchestration required for several agents to work well together?
Agentic systems are not about maximizing agents, prompts, tools or model calls. They are about keeping specialization, autonomy, context and verification under control — because when agents work together, intelligence isn’t the only thing that scales.
Complexity does too.