Results

Repeated-trial reliability

Coverage-aware distributions for latency, throughput, schema validity, and routing recommendations.

Repeated-trial reliability

Each case retains up to its latest three valid, prompt-version-compatible trials. Standard deviation is sample standard deviation; incomplete coverage is reported as N/A. Historical hallucination/* rows used the strict hallucination-v2 evidence contract. New runs record natural and guardrailed tracks separately.

Qwen (qwen)

  • Coverage: 39/39 trials across 13/13 cases; 13/13 cases have three trials
  • TTFT: mean 2.730s; median 0.910s; SD 7.964s; range 0.202–47.611s
  • Throughput: mean 14.964 tok/s; median 17.708 tok/s; SD 7.174 tok/s; range 0.084–25.144 tok/s
  • Output tokens: mean 203.4; median 123.0; SD 168.1; range 4.0–492.0
  • Schema success where applicable: 15/15 (100.0%)
  • Peak system memory: 16.80 GB; peak swap: 0.40 GB
CaseCoverageQuality trialsTTFT trials (s)Throughput trials (tok/s)Changed?
coding/implement_function3/30.750 / 0.750 / 0.7503.742 / 0.910 / 0.4802.036 / 15.630 / 21.747no
coding/code_review3/30.667 / 0.667 / 0.6672.339 / 0.300 / 0.2052.092 / 15.948 / 21.224no
reasoning/architecture_tradeoffs3/30.667 / 0.667 / 0.6670.202 / 0.304 / 0.30824.258 / 15.912 / 17.909no
reasoning/incomplete_debugging3/31.000 / 1.000 / 1.0000.202 / 0.297 / 0.23125.144 / 15.907 / 22.892no
repository/architecture-navigation3/31.000 / 1.000 / 1.0001.132 / 1.628 / 1.13222.726 / 14.999 / 22.219no
repository/appropriate-no-change3/31.000 / 1.000 / 1.0001.314 / 1.345 / 1.35818.727 / 18.716 / 18.683no
hallucination/sports-absent-quarterback3/31.000 / 1.000 / 1.0000.445 / 0.683 / 0.45518.440 / 12.034 / 18.298no
hallucination/api-absent-firmware3/31.000 / 1.000 / 1.0000.210 / 0.306 / 0.21117.708 / 11.664 / 17.680no
camera/snow3/30.750 / 0.750 / 0.7501.068 / 0.820 / 0.62113.863 / 18.518 / 19.042no
camera/lookalike-object3/30.950 / 0.950 / 0.9501.111 / 0.838 / 0.61213.869 / 18.366 / 19.226no
sports/nfl-adversarial-noisy-upset3/31.000 / 1.000 / 1.0002.025 / 1.091 / 1.21114.162 / 19.525 / 19.270no
rag/direct_answer3/31.000 / 1.000 / 1.00019.003 / 2.556 / 2.6221.008 / 5.749 / 5.637no
rag/unanswerable3/31.000 / 1.000 / 1.00047.611 / 3.056 / 2.4930.084 / 1.203 / 1.494no

Gemma (gemma)

  • Coverage: 39/39 trials across 13/13 cases; 13/13 cases have three trials
  • TTFT: mean 4.020s; median 0.910s; SD 9.483s; range 0.341–46.307s
  • Throughput: mean 11.055 tok/s; median 13.506 tok/s; SD 5.395 tok/s; range 0.086–15.937 tok/s
  • Output tokens: mean 306.8; median 300.0; SD 249.1; range 4.0–800.0
  • Schema success where applicable: 6/15 (40.0%)
  • Peak system memory: 16.50 GB; peak swap: 0.40 GB
CaseCoverageQuality trialsTTFT trials (s)Throughput trials (tok/s)Changed?
coding/implement_function3/30.750 / 0.750 / 0.7500.835 / 0.971 / 0.61014.442 / 11.238 / 14.039no
coding/code_review3/31.000 / 1.000 / 1.0000.368 / 0.398 / 0.37015.044 / 15.717 / 15.314no
reasoning/architecture_tradeoffs3/31.000 / 1.000 / 1.0000.358 / 0.352 / 0.38715.937 / 15.754 / 13.862no
reasoning/incomplete_debugging3/31.000 / 1.000 / 1.0000.358 / 0.557 / 0.34115.350 / 10.026 / 15.439no
repository/architecture-navigation3/30.000 / 0.000 / 0.00013.066 / 2.579 / 2.1872.765 / 13.506 / 13.930no
repository/appropriate-no-change3/31.000 / 1.000 / 1.00012.597 / 2.358 / 2.3011.378 / 11.534 / 10.588no
hallucination/sports-absent-quarterback3/30.000 / 0.000 / 0.0000.914 / 1.459 / 1.00515.075 / 9.604 / 13.328no
hallucination/api-absent-firmware3/30.000 / 0.000 / 0.0000.570 / 0.897 / 0.63915.852 / 10.027 / 13.467no
camera/snow3/30.000 / 0.000 / 0.0000.817 / 0.910 / 0.83715.558 / 14.631 / 15.373no
camera/lookalike-object3/31.000 / 1.000 / 1.0000.819 / 0.826 / 0.81314.825 / 14.635 / 14.571no
sports/nfl-adversarial-noisy-upset3/30.000 / 0.000 / 0.0001.540 / 1.733 / 1.6653.340 / 13.352 / 12.164no
rag/direct_answer3/31.000 / 1.000 / 1.00038.323 / 3.998 / 4.0240.497 / 3.498 / 3.461no
rag/unanswerable3/31.000 / 1.000 / 1.00046.307 / 3.813 / 3.8910.086 / 0.983 / 0.968no

GPT-OSS (gptoss-final)

  • Coverage: 13/39 trials across 11/13 cases; 0/13 cases have three trials
  • TTFT: mean 0.897s; median 0.672s; SD 0.702s; range 0.330–2.821s
  • Throughput: mean 22.739 tok/s; median 21.711 tok/s; SD 12.034 tok/s; range 2.926–36.665 tok/s
  • Output tokens: mean 158.3; median 42.0; SD 159.9; range 4.0–400.0
  • Schema success where applicable: 3/3 (100.0%)
  • Peak system memory: 18.49 GB; peak swap: 2.37 GB
CaseCoverageQuality trialsTTFT trials (s)Throughput trials (tok/s)Changed?
coding/implement_function1/31.0000.67235.666no
coding/code_review1/31.0000.33036.665no
reasoning/architecture_tradeoffs1/31.0000.45730.923no
reasoning/incomplete_debugging1/30.5000.38428.995no
repository/architecture-navigation1/31.0001.05433.463no
repository/appropriate-no-change1/31.0001.10919.276no
hallucination/sports-absent-quarterback2/31.000 / 1.0000.511 / 2.82116.863 / 4.707no
hallucination/api-absent-firmware2/31.000 / 1.0000.381 / 0.37821.549 / 21.711no
camera/snow0/3N/AN/AN/AN/A
camera/lookalike-object0/3N/AN/AN/AN/A
sports/nfl-adversarial-noisy-upset1/31.0000.73935.316no
rag/direct_answer1/31.0001.6077.544no
rag/unanswerable1/31.0001.2162.926no

Deterministic routing

A five-percentage-point margin is required; executable full-suite pass rate drives coding. N/A means the required benchmark data is unavailable.

WorkloadQwenGemmaGPT-OSSRecommendation
coding100.0%80.0%100.0%No meaningful difference
reasoning83.3%100.0%N/AGemma
repository100.0%50.0%N/AQwen
multimodal93.4%84.1%N/AQwen
vision93.4%84.1%N/AQwen
camera91.4%91.1%N/ANo meaningful difference
rag99.9%95.0%94.1%No meaningful difference
sports100.0%0.0%N/AQwen
hallucination100.0%0.0%N/AQwen
long_context68.3%41.6%100.0%GPT-OSS
default50.0%10.0%N/AQwen