Skip to main content
Back to Blog
Blog

Under the Hood: How JSONL Session Resolution Powers Agent Routing

A deep dive into how MadoHub reads Claude Code's session logs to extract clean, structured responses for intelligent routing.

The hardest problem in agent routing isn't generating the prompt for the target agent. It's extracting a clean response from the source agent.

Terminal output is messy. ANSI escape codes, progress spinners, line wrapping, color sequences — raw terminal scraping gives you a wall of noise. You can strip the control codes, but you still end up with formatting artifacts, partial lines, and mixed content.

MadoHub solves this differently. Instead of scraping terminal output, we read Claude Code's structured session logs — the JSONL files that contain clean, semantic representations of every conversation turn.

This post explains how that works.

The problem with terminal scraping

When Claude Code runs in a terminal, its output includes:

  • ANSI escape sequences for colors, cursor positioning, and formatting.
  • Unicode markers like ❯ (prompt), ✽ (thinking), ⏺ (tool use).
  • Progress indicators that overwrite previous lines.
  • Permission dialogs with interactive elements.
  • Multi-line code blocks with syntax highlighting escape codes.

Stripping ANSI codes gets you closer to clean text, but not all the way. Progress indicators leave partial lines. Code blocks lose their structure. The boundary between "agent output" and "shell noise" is ambiguous.

For simple routing (e.g., "send the last 10 lines"), terminal scraping is adequate. But for intelligent routing — where an AI model needs to understand what the agent actually did and generate a contextual prompt — you need clean, structured data.

How Claude Code stores sessions

Claude Code writes structured session logs in JSONL format (one JSON object per line). The resolution path is:

tmux pane → ps process-tree BFS
  → agent PID (first descendant whose command matches claude/codex/amp/opencode)
    → ~/.claude/sessions/{PID}.json
      → sessionId field
        → ~/.claude/projects/*/{sessionId}.jsonl
          → conversation turns (user/assistant messages)

Each conversation turn contains the message role, content, and metadata — without any terminal formatting. This is the same data Claude Code uses internally to maintain conversation context.

The resolution chain

MadoHub resolves the JSONL file through a multi-step chain:

Step 1: Get the PTY process ID

Each terminal tile is backed by a tmux pane, and MadoHub tracks that pane's PID from the moment the terminal is created. This PID is the starting point for locating the actual agent process, which typically runs as a descendant of the pane's shell.

Step 2: Find the agent PID and its session file

Claude Code maintains a sessions directory at ~/.claude/sessions/, but each file is named after the agent's process ID — not a session ID. To find that PID, MadoHub walks the process tree from the tmux pane (a ps -eo pid=,ppid=,comm= breadth-first search over descendants) until it reaches the first process whose command name matches a known agent (claude, codex, amp, opencode). MadoHub then reads ~/.claude/sessions/{PID}.json for that PID, which contains a sessionId field pointing at the actual conversation log.

Step 3: Locate and read the JSONL

MadoHub takes the sessionId from the session file and scans every subdirectory of ~/.claude/projects/ for a file named {sessionId}.jsonl — that file is the transcript. MadoHub parses it and extracts the last assistant message — the most recent response from Claude Code.

This message is clean text: no ANSI codes, no terminal artifacts, no progress indicators. It's exactly what Claude Code "said," in the format it intended.

Step 4: When resolution fails

If the process-tree walk can't find a matching agent process, or the session file has already disappeared (e.g., the process exited), MadoHub has no JSONL path available for that turn. It falls back to terminal output extraction instead — the same ANSI-stripping approach used for agents that don't produce session logs at all.

What the clean response enables

With a clean, structured response, the routing transform can:

Understand context: The AI routing model receives the actual semantic content — "I refactored the auth middleware to use JWT tokens instead of session cookies. Changed 3 files: src/auth/middleware.ts, src/auth/tokens.ts, and src/config/auth.ts" — instead of a mess of colored terminal output.

Generate precise prompts: "Review the JWT token migration in src/auth/middleware.ts, src/auth/tokens.ts, and src/config/auth.ts. Verify that the token refresh flow handles expiration correctly and that the middleware properly validates token signatures."

Preserve code blocks: If the agent's response includes code snippets, they're preserved with their original formatting — not corrupted by terminal line wrapping and ANSI color codes.

Caching and performance

JSONL files can grow large over a long session. MadoHub caches the last resolved response per tile ID to avoid re-parsing the entire file on every routing trigger. The cache is invalidated when a new completion is detected.

The resolution chain is also optimized for the common case: the process-tree walk and session file read are fast, and resolution succeeds the overwhelming majority of the time. Scanning ~/.claude/projects/ for the matching {sessionId}.jsonl file is the more expensive step, since it has to search every project subdirectory.

Limitations

JSONL-based state detection is no longer Claude-Code-only — Codex now has its own JSONL-backed status pipeline (task_complete → completed, turn_aborted/error → failed, user_message/task_started/agent_message → working, ri_message → stopped). Where the agents still differ is response-text extraction: the clean, structured session-log format described in this post — reading the last assistant message straight out of the transcript — is specific to Claude Code. For Codex and shell commands, MadoHub still falls back to terminal output extraction with ANSI stripping to get the actual response text.

The raw and full-output transform types work with any agent because they use terminal output directly. The ai-routing and summary transforms produce the best results with Claude Code because they receive the clean JSONL content.

Why this matters

The quality of routing output is directly proportional to the quality of the extracted response. Garbage in, garbage out. By reading structured session logs instead of scraping terminal output, MadoHub gets the highest-fidelity representation of what the source agent actually produced.

This is a design decision with compounding benefits. Better extraction → better transform prompts → more accurate target agent instructions → fewer wasted cycles → faster workflow convergence. The investment in clean extraction pays off at every subsequent step of the routing pipeline.