Ensemble Exploration
Spawn several agents on the same problem with different instructions, compare their output, and pick a winner.
Sometimes you do not know which approach is best — LRU vs LFU vs ARC for a cache, three different refactor strategies, two competing library choices. Ensemble Exploration spawns several fresh agents on the same problem, each given a different instruction for how to attack it, and lets you compare the results side by side.
This is triggered through MadoAgent — you ask MadoAgent to explore approaches, and it decides when a fork-and-compare shape fits better than a single task. Each fork spawns as a brand-new tile with its own independent context, so divergent approaches do not contaminate each other.
What you can do with it
- Run up to 5 forks on the same problem — each fork gets the shared prompt plus its own variant instruction.
- Mix agent types per fork — Claude, Codex, or Amp, so you can pit different models against the same task.
- Compare forks side by side — the Exploration panel's Compare view shows each fork's last response in a horizontal lane.
- Score forks with the auto-critic — one click asks the critic to rate each fork 0–10 on correctness, completeness, clarity, and actionability, and name one recommended fork.
- Pick a winner manually — the critic only advises; selecting a winner is always your click.
- Keep runs across relaunches — recent runs are restored on restart so the panel survives an app relaunch.
How to use it
- Open MadoAgent with Cmd+Shift+M (⌘⇧M).
- Ask for an exploration, for example "try three different caching strategies for this bug and tell me which wins".
- MadoAgent spawns one tile per fork, each with the shared prompt plus its own variant instruction. The right drawer auto-switches to the Exploration tab so you see the run the moment it starts.
- Watch each fork's status dot in the Exploration panel —
running,done,cancelled, ortimeout. - When forks have output, click Compare to expand a side-by-side lane of each fork's last response.
- Optionally click Score forks to run the critic. The critic rates each fork and highlights a recommended one.
- Click Pick winner on the fork you want. The run closes as complete and any other running fork in that run is marked done.
Common use cases
- "Try LRU, LFU, and ARC for this cache and tell me which to use." — three forks, one winner, picked after reading the Compare view or the critic's scores.
- "Refactor this module two ways — extract a class vs split into modules — and compare." — two forks with different refactor strategies.
- "Pit Claude and Codex on the same task." — two forks, one per agent type, to see which model handles the task better.
Tips and best practices
- Cap forks at what you can actually compare. Three to five is the sweet spot; more than five and the Compare view becomes hard to scan.
- Use the critic as a tiebreaker, not a decider. Read the Compare view yourself before clicking Pick winner — the critic can be wrong.
- Forks with no output yet are skipped by the critic, so wait until forks have produced something before scoring.
- Dismissing a run cancels any still-running forks rather than leaving them marked
runningforever.
Related features
- MadoAgent — the panel where you trigger ensemble runs.
- Right Drawer — the Exploration tab where runs and forks are shown.
- Orchestrator Mode — for a single headless worker instead of several forks.
- Settings — the AI provider the critic uses.
Limitations
- Forks are capped at 5 per run. Requesting more rejects the whole call rather than silently truncating it.
- The critic only advises — it never sets a winner. MadoHub does not auto-pick a winner.
- The critic uses your configured AI provider's cheaper/faster model tier, so its judgment is only as good as that model.
- Ensemble runs persist for the last 7 days (capped at 20 runs) — older runs age out.