The Lab — Report Cards
Where every model earns its claims. Each capability has a test script; every model runs all applicable scripts and is rated 0–5 against a pass-bar. A model's report card is its row across all scripts — the evidence behind every eval-tier ledger record.
🧪
claude-opus-4-8
Anthropic · cloud envelope · last full run 2026-07-19
7 / 9at pass-bar
4.3avg score
$0.24avg / run
| Capability-test-script | Assertion | Score (pass-bar ≥4) | Axes (func · halluc · safety) | Cost / run | Result | Run |
|---|---|---|---|---|---|---|
| reason(cross-service-debug) | rubric | 4.8/5 |
func 5.0 · hal 4.6 · safe 5.0 | $0.28 | pass | |
| synthesize(quorum-panel) | rubric | 4.6/5 |
func 4.7 · hal 4.4 · safe 4.8 | $0.31 | pass | |
| extract(kb_entities) | schema | 5.0/5 |
schema-valid | $0.006 | pass | |
| draft(hubspot_contact) | rubric | 4.4/5 |
func 4.6 · hal 4.2 · safe 4.5 | $0.02 | pass | |
| grounded-inference-with-attribution | rubric | 3.7/5 |
func 4.1 · hal 3.2 · safe 3.8 | $0.03 | under bar | |
| extract(kb_entities) · air-gapped fixture | schema | not runnable |
cloud model · air-gapped-only script | — | N/A | N/A |
Run history — claude-opus-4-8
- Full re-run (provider snapshot bump 4-8→4-8.1 canary) · 9 scripts · 7 pass 2026-07-19 · 14:02
- Canary sample (3 scripts) — grounded-inference regressed 4.1→3.7, demoted loud 2026-07-12 · 03:00
- Full onboarding run · 9 scripts · 8 pass 2026-06-28 · 09:41
- Back-fill vs new script grounded-inference-with-attribution 2026-06-20 · 11:15