MadoHub Docs

Ensemble Exploration

Spawn several agents on the same problem with different instructions, compare their output, and pick a winner.

Sometimes you do not know which approach is best — LRU vs LFU vs ARC for a cache, three different refactor strategies, two competing library choices. Ensemble Exploration spawns several fresh agents on the same problem, each given a different instruction for how to attack it, and lets you compare the results side by side.

This is triggered through MadoAgent — you ask MadoAgent to explore approaches, and it decides when a fork-and-compare shape fits better than a single task. Each fork spawns as a brand-new tile with its own independent context, so divergent approaches do not contaminate each other.

What you can do with it

  • Run up to 5 forks on the same problem — each fork gets the shared prompt plus its own variant instruction.
  • Mix agent types per fork — Claude, Codex, or Amp, so you can pit different models against the same task.
  • Compare forks side by side — the Exploration panel's Compare view shows each fork's last response in a horizontal lane.
  • Score forks with the auto-critic — one click asks the critic to rate each fork 0–10 on correctness, completeness, clarity, and actionability, and name one recommended fork.
  • Pick a winner manually — the critic only advises; selecting a winner is always your click.
  • Keep runs across relaunches — recent runs are restored on restart so the panel survives an app relaunch.

How to use it

  1. Open MadoAgent with Cmd+Shift+M (⌘⇧M).
  2. Ask for an exploration, for example "try three different caching strategies for this bug and tell me which wins".
  3. MadoAgent spawns one tile per fork, each with the shared prompt plus its own variant instruction. The right drawer auto-switches to the Exploration tab so you see the run the moment it starts.
  4. Watch each fork's status dot in the Exploration panel — running, done, cancelled, or timeout.
  5. When forks have output, click Compare to expand a side-by-side lane of each fork's last response.
  6. Optionally click Score forks to run the critic. The critic rates each fork and highlights a recommended one.
  7. Click Pick winner on the fork you want. The run closes as complete and any other running fork in that run is marked done.

Common use cases

  • "Try LRU, LFU, and ARC for this cache and tell me which to use." — three forks, one winner, picked after reading the Compare view or the critic's scores.
  • "Refactor this module two ways — extract a class vs split into modules — and compare." — two forks with different refactor strategies.
  • "Pit Claude and Codex on the same task." — two forks, one per agent type, to see which model handles the task better.

Tips and best practices

  • Cap forks at what you can actually compare. Three to five is the sweet spot; more than five and the Compare view becomes hard to scan.
  • Use the critic as a tiebreaker, not a decider. Read the Compare view yourself before clicking Pick winner — the critic can be wrong.
  • Forks with no output yet are skipped by the critic, so wait until forks have produced something before scoring.
  • Dismissing a run cancels any still-running forks rather than leaving them marked running forever.
  • MadoAgent — the panel where you trigger ensemble runs.
  • Right Drawer — the Exploration tab where runs and forks are shown.
  • Orchestrator Mode — for a single headless worker instead of several forks.
  • Settings — the AI provider the critic uses.

Limitations

  • Forks are capped at 5 per run. Requesting more rejects the whole call rather than silently truncating it.
  • The critic only advises — it never sets a winner. MadoHub does not auto-pick a winner.
  • The critic uses your configured AI provider's cheaper/faster model tier, so its judgment is only as good as that model.
  • Ensemble runs persist for the last 7 days (capped at 20 runs) — older runs age out.

On this page