Model Onboarding
Nishad

Model Onboarding

A new model earns its place: registered β†’ the Lab resolves which scripts apply (declared capabilities ∩ runnable envelope) β†’ runs them β†’ is rated β†’ recorded. Only capabilities that clear the pass-bar become resolver-eligible; the rest become gaps.

1Registeredid Β· vendor Β· envelope
2Scripts resolved9 applicable Β· 2 skipped
3Running4 of 9 complete
4Rated0–5 per script
5Recordedto the ledger

1 Β· Registered model

step complete
Model id
claude-sonnet-5
Vendor
Anthropic
Envelope
cloud
Registered by
Nishad
Declared
reason(cross-service-debug) draft(hubspot_contact) synthesize(quorum-panel) extract(kb_entities) grounded-inference

2 Β· Applicable scripts resolved

9 apply Β· 2 skipped

Resolution: declared capabilities ∩ scripts runnable in the cloud envelope. Run scripts kicks off every applicable script; each row also has its own Run so you can re-run one in isolation. Air-gapped-only scripts are skipped for a cloud model β€” recorded as N/A-in-envelope, never a silent zero.

  • reason(cross-service-debug)rubric Β· pass-bar β‰₯44.7
  • extract(kb_entities)schema-valid5.0
  • draft(hubspot_contact)rubric Β· pass-bar β‰₯44.3
  • synthesize(quorum-panel)rubric Β· running nowrunning
  • grounded-inference-with-attributionrubric Β· queuedqueued
  • extract(kb_entities) Β· air-gapped fixturecloud model β€” not runnable air-gappedN/A β€” envelopeN/A

3 Β· Live run

synthesize(quorum-panel)
Overall progress4 / 9 scripts
[14:02:07] β–Ά synthesize(quorum-panel) starting β€” fixture qp-golden-07
[14:02:09] panel assembled Β· 4 panelists Β· 3 rounds
[14:02:41] synthesis produced Β· 1,842 tokens out
[14:02:42] LLM-as-judge scoring 3 axes…
[14:02:44] func 4.7 Β· halluc 4.4 Β· safety 4.8
[14:02:44] composite 4.6 / 5 β€” PASS (bar β‰₯4)
[14:02:45] next β–Ά grounded-inference-with-attribution
Cost this run $0.31 Β· script budget $0.50 Β· under cap. A run over budget aborts loud, never truncates silently.

5 Β· Recorded so far

eval-tier evidence
  • reason(cross-service-debug) eval Β· 4.7
  • extract(kb_entities) eval Β· 5.0
  • draft(hubspot_contact) eval Β· 4.3

On finish: capabilities at/above the bar become resolver-eligible at eval tier; any below-bar declared capability becomes a gap on the roadmap.

What's on this screen

Built from: a new model running the same capability test scripts as every other model, live, with each result recorded as it finishes.

On screenWhere it comes fromWhat it means to you
Stepper (Registered β†’ Scripts resolved β†’ Running β†’ Rated β†’ Recorded)the five stages this new model moves through before it can be trusted with real workshows how close this model is to being usable
Registered model card (id, vendor, envelope, registered by)the details you entered when adding the modelconfirms exactly what's being onboarded and by whom
Declared capabilities chipsthe skills the vendor or you claim this model haswhat's about to be tested β€” a claim, not yet proof
Applicable scripts resolved listthe tests that actually apply, matched against the model's declared skills and where it's allowed to runshows exactly what will be tested vs. skipped, and why
Live run panel (progress bar + running log)the test executing right now, step by steplets you watch it happen in real time β€” proof it's a genuine run, not a canned result
Budget line (cost this run vs. cap)what's been spent testing this model against the limit set for that scriptprotects you from a runaway bill just to onboard one model
Recorded so far listthe scores already logged during this onboarding sessionshows what's already proven, even before the whole run finishes
Finish & record all / Pausethe action that locks in every run's score, or halts mid-waythis is the step that makes the model usable fleet-wide, or lets you stop and check results first