The Lab — Report Cards
Nishad

The Lab — Report Cards

Where every model earns its claims. Each capability has a test script; every model runs all applicable scripts and is rated 0–5 against a pass-bar. A model's report card is its row across all scripts — the evidence behind every eval-tier ledger record.

🧪
claude-opus-4-8 Anthropic · cloud envelope · last full run 2026-07-19
7 / 9at pass-bar
4.3avg score
$0.24avg / run
Capability-test-script Assertion Score (pass-bar ≥4) Axes (func · halluc · safety) Cost / run Result Run
reason(cross-service-debug) rubric
4.8/5
func 5.0 · hal 4.6 · safe 5.0 $0.28 pass
synthesize(quorum-panel) rubric
4.6/5
func 4.7 · hal 4.4 · safe 4.8 $0.31 pass
extract(kb_entities) schema
5.0/5
schema-valid $0.006 pass
draft(hubspot_contact) rubric
4.4/5
func 4.6 · hal 4.2 · safe 4.5 $0.02 pass
grounded-inference-with-attribution rubric
3.7/5
func 4.1 · hal 3.2 · safe 3.8 $0.03 under bar
extract(kb_entities) · air-gapped fixture schema
not runnable
cloud model · air-gapped-only script N/A N/A

Run history — claude-opus-4-8

  • Full re-run (provider snapshot bump 4-8→4-8.1 canary) · 9 scripts · 7 pass 2026-07-19 · 14:02
  • Canary sample (3 scripts) — grounded-inference regressed 4.1→3.7, demoted loud 2026-07-12 · 03:00
  • Full onboarding run · 9 scripts · 8 pass 2026-06-28 · 09:41
  • Back-fill vs new script grounded-inference-with-attribution 2026-06-20 · 11:15

What's on this screen

Built from: live test runs of each model or agent against a fixed set of capability test scripts, graded either by an exact rule or by an LLM-as-judge rubric.

On screenWhere it comes fromWhat it means to you
Report card header (model, envelope, last full run date)the identity and setup of the model this card is aboutconfirms which model you're looking at and how fresh its results are
Summary stats (at pass-bar / avg score / avg cost)a roll-up across every script this model has been tested onquick read on whether this model is generally trustworthy and affordable
Capability-test-script rowsone real skill this model has been checked againstwhich specific jobs it's actually been proven — or not proven — on
Assertion type (schema vs. rubric)whether that skill was graded by an exact pass/fail rule or by a judge's ratingtells you how objective vs. judgment-based that pass really is
Score + pass-bar markerthe result of the most recent run against the required barwhether it cleared the line needed to count as proven
Axes breakdown (func · halluc · safety)the score split into whether it did the job, didn't make things up, and stayed safeshows exactly where a weak result is weak, not just that it failed
Cost / runwhat that one test run cost to executeweigh testing cost against how badly you need that proof
Result (pass / under bar / N/A)the verdict from that runwhether this result is strong enough to feed the Capability Ledger
Run / Re-run buttontriggers a fresh live test right nowget an up-to-date answer instead of trusting a stale score
Run historya log of past runs for this model, including regressionsproof of when and why a score changed — a provider update, a demotion, a new script
Models in the Lab (sidebar)the composite score for every other model or agent under testcompare candidates side by side before choosing one