Practical evaluation · versioned contracts

Benchmarks built to fail honestly.

A request that never reaches inference is not a model failure. A fast answer is not automatically a good one. Every layer is measured separately.

Infrastructure ≠ qualityTransport and parsing failures stay separate
Unsupported ≠ zeroMissing capabilities are reported as N/A
Small samples stay labeledNo confidence theater

Real work, bounded claims.

Deterministic checks carry the correctness signal where possible. Qualitative outputs and raw JSONL remain available for human review.

01

Coding

Implementation, repair, review, refactoring, and executable patch validation.

02

Grounding

Natural and guardrailed hallucination resistance against fixed local evidence.

03

Repository

Navigation, bounded changes, tests, and appropriate no-change decisions.

04

Vision

Synthetic diagrams, camera scenes, photographs, charts, and mixed inputs.

05

RAG

Answerable and unanswerable retrieval at approximately 2K, 8K, and 16K.

06

Sports

Fictional prediction, recommendation gates, arithmetic, and ledger audits.