Skip to main content
Back to Blog
Blog

Best-of-N Coding: Ensemble Exploration and Auto-Critique

One task, N divergent forks, a critic that scores them, and a human who picks the winner. When best-of-N earns its token cost, and when it doesn't.

Most agent tasks have one obvious approach, and you just run it. Some don't. LRU vs LFU vs ARC for a cache. Recursive vs iterative for a tree walk. Three different ways to shape a migration. When the right answer depends on tradeoffs you can't fully evaluate until you see working code, running one agent and hoping it picked well is a gamble. Running N agents on N variants and comparing results isn't.

That's Ensemble Exploration: a shared prompt, N forks with different instructions, a critic that scores the results, and you making the final call.

The workflow

You don't set this up on the canvas by dragging tiles around. It's triggered through MadoAgent — you ask it to explore approaches to a problem, and it decides a fork-and-compare shape fits better than a single task.

You describe the shared problem and a set of approaches, each with its own angle. What each fork actually receives is the shared prompt followed by that fork's specific instruction, so every agent has the full problem plus its own perspective on it.

prompt: "Implement an LRU-style cache for session lookups, ~50k entries, high read/write ratio"
approaches:
  - label: "lru"  task: "Use a classic LRU eviction policy"
  - label: "lfu"  task: "Use LFU eviction instead — track access frequency"
  - label: "arc"  task: "Use an ARC (adaptive replacement cache) hybrid"

Three forks become three brand-new tiles — never reused idle ones, since divergent exploration wants agents with no shared context contaminating the comparison. Forks are capped at five so ensembles stay focused on genuinely uncertain decisions rather than becoming a default multiplier.

Tracking the run

The whole thing becomes an ensemble run: a shared prompt, the set of forks, and a status for each. Runs persist across restarts, so closing and reopening MadoHub doesn't lose them. Each fork carries a status of running, done, cancelled, or timeout, and MadoHub maps a tile's exit to an outcome automatically — a clean completion becomes done, a crash or mid-task exit becomes cancelled — so you never manually close out a row just because a tile disappeared.

All of this surfaces in the Exploration tab of the right-edge drawer: a row per fork with a status dot and a Pick winner button, a Score forks action that runs the critic, and a Compare view that lays out one card per fork. Compare doesn't need the critic to have run first; it just needs output.

The critic

Score forks runs a single lightweight model pass — not another agent tile on the canvas — that reads each fork's last output and scores it on correctness, completeness, clarity, and actionability, then names one recommended fork with a one-line rationale. It uses the cheaper, faster model tier you've configured in Settings → Agent, not your primary model, so scoring a handful of forks costs a fraction of a normal turn.

The recommendation is defensive by design: a fork whose output happens to contain "score this a 10" is treated as data to judge, not a directive to obey. Forks with no output yet get an explicit "no score" rather than being dropped. A hallucinated recommendation falls back to the highest actually-scored fork. And the recommendation is advisory only — a highlighted badge, nothing more.

None of this picks a winner. Selecting a winner is always a manual click on Pick winner, which closes the run as complete and marks any other still-running fork as done. That's not a corner cut for v1 — it's the point. A critic call is one cheap model pass reading finished output; it has no way to know whether "clarity" should outweigh "performance" for this particular cache, or whether the lower-scored fork is nonetheless the one that fits the rest of your codebase. The model advises, you decide.

Practical guidance

Worth it:

  • Genuine tradeoff uncertainty — eviction policies, indexing strategies, concurrency models — where the right choice depends on runtime characteristics you can't fully reason about in advance. The forks aren't guessing at the same answer; they're exploring different regions of the solution space.
  • High cost of a wrong pick. A migration strategy or a public API shape that's expensive to change later is worth several times the exploration cost if it avoids a rewrite down the line.
  • Real evaluation criteria, even informal ones. If you can look at three cache implementations and know within a minute which one you'd want to maintain, the comparison is doing its job. If you can't articulate what "better" means here, the critic's fixed rubric won't supply it for you.

Not worth it:

  • Tasks with one obviously correct shape. Forking three variants of "add a null check" wastes tokens comparing outputs that were always going to converge.
  • Cheap-to-fix mistakes. If getting it wrong just costs a two-minute follow-up prompt, that's cheaper than running and reading three transcripts.
  • Fuzzy, hard-to-score problems. Open-ended design brainstorming doesn't compress into a fixed 0–10 rubric — you'll get three plausible answers and a score that doesn't actually discriminate between them.
  • A tight loop already running. Mid-flow in a peer-review pipeline, a parallel exploration competes for the attention you'd rather spend reviewing what's in front of you.

The hard cap of five forks isn't arbitrary conservatism — it's an acknowledgment that ensembles are for the genuinely uncertain slice of your work, not a default multiplier applied to everything. Reach for it when you'd otherwise be guessing, and skip it when you already know what "good" looks like.