Starchild Benchmarks

Conductor Mode — Benchmark Report

Multi-model routing measured against single models and frontier baselines.

Routing comparison: mixed-150 public set Model bench: Aug 2026 Generated: Updated:

Conductor v4.0 (Sep 2026): DeepSeek V4 Flash + GLM-5.3-Flash + Gemini 3.8 Flash + Muse Spark 1.3. All four composite suites re-run under the 65-line condensed rules: composite 94.8% (AIME 100.0 / GPQA 93.4 / LCB-100 95.6 / Everyday 90.2), Everyday cost down 66% ($0.00093/task). All results downloadable above.

Conductor v4.0 Final Report — new Conductor rules

Final cost-performance report for the v4.0 four-leg configuration (cheap = DeepSeek V4 Flash, normal = GLM-5.3-Flash, strong = Gemini 3.8 Flash, escalation = Muse Spark 1.3; media = images → GLM-5.3-Flash / video → Gemini 3.8 Flash). AIME / GPQA / Everyday / LCB-100 are live final-ladder runs under 65-line production rules; MMMU carries over the media-leg measurement.

★2

Composite weighted ranking (4 suites)

Equal-weight composite over AIME 2025, GPQA Diamond, LCB-100 and Everyday — only models that ran all four suites are ranked. Cost = mean $/task across the four suites.

Conductor v2 is not ranked here: it only has real runs on AIME and Everyday — it never ran the GPQA or LCB-100 pools. The scatter below uses the Everyday suite only.

§1

Cost vs quality

Each point is one model: per-task cost (log scale) against the equal-weight composite over all four suites — AIME 2025, GPQA Diamond, LCB-100 and Everyday. Only models that completed all four are plotted. The dashed line is the Pareto frontier.

§2

Everyday tasks

Accuracy on a lightweight everyday-task set (trivial QA, yes/no, easy MCQ, grade math, data tasks, knowledge) — 94-task common set that every model, including Conductor v2 and v3, actually ran; per-task cost annotated.

Sorted by accuracy on the Everyday task pool. Cost is billed OpenRouter price per task.

Everyday-suite view of the same axes: every point is a measured run — cost per task (log scale) against accuracy. This chart includes models that only ran the Everyday suite; routers sit near the frontier, the oracle marks the upper bound of this routing family.

§2b

LiveCodeBench 100 — Model comparison

Hard-coding on the LCB-100 pool. Main column = accuracy on the task set every leg completed (same tasks, same grading) — full-pool numbers differ in n because sandbox timeouts are excluded, which drops mostly hard problems and inflates cheap-model headline scores.

Sorted by same-set accuracy. Conductor v4.0 ladder: 95.6% on the pool at $0.0299/problem — easy 21/21, medium 26/26, hard 47/51. Driven by Gemini 3.8 Flash + Muse Spark 1.3.

§4

GPQA Diamond

Graduate-level science QA (Physics, Chemistry, Biology). 197-task pool, Conductor v4.0 included.

LegRatePassed$/qTotal $

198-task pool. Sorted by accuracy.