Multi-model routing measured against single models and frontier baselines.
Conductor v3 final (Aug 2026): DeepSeek V4 Flash + GPT-5.6 Luna + Qwen 3.8 Max. All results downloadable above.
Final cost-performance report for the v3 four-leg configuration (economy = DeepSeek V4 Flash, strong = GPT-5.6 Luna @ xhigh, escalation = Qwen 3.8 Max, media = images → Luna / video → Qwen 3.8 Max; fallback Luna → Qwen 3.8 Max → DeepSeek V4 Flash). AIME / GPQA / Everyday / LCB-100 are live final-ladder runs; MMMU carries over the media-leg measurement.
—
Equal-weight composite over AIME 2025, GPQA Diamond, LCB-100 and Everyday — only models that ran all four suites are ranked. Cost = mean $/task across the four suites.
Conductor v2 is not ranked here: it only has real runs on AIME and Everyday — it never ran the GPQA or LCB-100 pools. The scatter below uses the Everyday suite only.
Each point is one model: per-task cost (log scale) against the equal-weight composite over all four suites — AIME 2025, GPQA Diamond, LCB-100 and Everyday. Only models that completed all four are plotted. The dashed line is the Pareto frontier.
—
Accuracy on a lightweight everyday-task set (trivial QA, yes/no, easy MCQ, grade math, data tasks, knowledge) — 94-task common set that every model, including Conductor v2 and v3, actually ran; per-task cost annotated.
—
Tiers E1–E5 + M1 (by difficulty, lowest→highest). Summary row: accuracy, cost per run, total cost.
Everyday-suite view of the same axes: every point is a measured run — cost per task (log scale) against accuracy. This chart includes models that only ran the Everyday suite; routers sit near the frontier, the oracle marks the upper bound of this routing family.
—
Hard-coding on the LCB-100 pool. Main column = accuracy on the task set every leg completed (same tasks, same grading) — full-pool numbers differ in n because sandbox timeouts are excluded, which drops mostly hard problems and inflates cheap-model headline scores.
Sorted by same-set accuracy. Conductor v3 final ladder: 90%+ on the 99-problem pool at ~$0.02/problem — easy 25/25, all 6 escalations to Qwen 3.8 Max recovered.
Graduate-level science QA (Physics, Chemistry, Biology). 198-task pool × 13 legs, Conductor v3 included.
| Leg | Rate | Passed | $/q | Total $ |
|---|
198-task pool. Sorted by accuracy.