Capability Ledger
Nishad

Capability Ledger

Every evidenced claim in the install — organized capability-first. Each row is one capability; expand it to see every subject that holds it (models and agents), each with its own evidence tier, confidence, deployment envelope, and cost. One capability can have many holders — some at production tier, some Lab-only, some below the bar.

6 capabilities · 13 holders · 8 below the 95% gate
draft(hubspot_contact) via hubspot-connector to rubric≥4/5 Draft a HubSpot contact record from an inbound lead, judged by rubric. production best held 2 holders
Subject (holder)Deployment envelopeEvidence tierConfidenceCost / unit
codex-1 agent · model×connector×playbook cloud production
42 clean cycles
0.97
$0.011 / contact
sonnet-1 agent on-prem-connected eval
Lab run #3980
0.79
below 95% gate — Lab-only
$0.02 / contact
reason(cross-service-debug) to rubric≥4 Reason across logs, DB, and code to a root cause, judged by rubric. production best held 2 holders
Subject (holder)Deployment envelopeEvidence tierConfidenceCost / unit
claude-opus-4-8 model · isolated cloud production
18 clean cycles
0.96
$0.28 / task
gemma-4-27b model · air-gapped-runnable air-gapped eval
Lab run #4290 (Thor node)
0.58
below 95% gate
$0.0006 / task
synthesize(quorum-panel) to rubric≥4 Synthesize a multi-model deliberation into a decision-grade answer. production best held 2 holders
Subject (holder)Deployment envelopeEvidence tierConfidenceCost / unit
opus-1 agent cloud production
0.95
$0.31 / synthesis
sonnet-1 agent on-prem-connected eval
Lab run #4051
0.77
below 95% gate — Lab-only
$0.09 / synthesis
grounded-inference-with-attribution via kb-graph, cited Answer from the KB graph with every claim cited to a source — no ungrounded assertions. production best held 3 holders
Subject (holder)Deployment envelopeEvidence tierConfidenceCost / unit
sonnet-1 agent cloud production
31 clean cycles
0.96
$0.02 / answer
sonnet-1 agent · in-network on-prem-connected eval
Lab run #4127
0.71
below 95% gate — Lab-only
$0.02 / answer
gemma-4-27b model · air-gapped-runnable air-gapped eval
Lab run #4331 (Thor node) · KB-only
0.66
below 95% gate
$0.0005 / answer
extract(kb_entities) to schema-valid Pull structured entities out of a document into a valid schema. production best held 3 holders
Subject (holder)Deployment envelopeEvidence tierConfidenceCost / unit
claude-opus-4-8 model cloud production
0.98
$0.006 / doc
gemma-4-27b model · in-network on-prem-connected eval
Lab run #4188
0.81
below 95% gate — Lab-only
$0.0008 / doc
gemma-4-27b model · air-gapped-runnable air-gapped eval
Lab run #4310 (Thor node)
0.62
below 95% gate
$0.0004 / doc
test-execute(ui-playwright) to evidence-verified Run a UI test in a real browser and produce runtime evidence it actually executed. declared best held 1 holder
Subject (holder)Deployment envelopeEvidence tierConfidenceCost / unit
haiku-1 agent cloud declared
production evidence revoked
0.30
DEMOTED — fake-green (TECH-652); N=5 re-rating
95% honesty gate. A capability lists every holder — a holder below the confidence bar is shown as such (Lab-only, demoted, or not-scored), never dressed up as a confident match. A cloud-only capability is N/A for an air-gapped install until a local candidate earns it.

What's on this screen

Built from: results the Lab has recorded from testing each model or agent against real capability test scripts, organized by the skill they were tested for.

On screenWhere it comes fromWhat it means to you
Capability statement + descriptiona real skill the fleet needs, phrased plainly under the jargonwhich job this is — the plain sentence tells you what it actually does
Filters (capability / subject / envelope / evidence tier / search)the same holder records, narrowed to what you typefind the model or agent you care about fast, instead of scrolling everything
Best tier badge + holders countthe strongest proof any model or agent has earned for this skill, and how many have triedat a glance, is this skill solidly proven or still shaky
Subject (holder)the specific model or agent that was testedwho to trust — or not — with this kind of work
Deployment envelopewhere that model or agent is allowed to run (cloud, in-network, or fully offline)whether this option even works for your install's rules
Evidence tier (declared / eval / production)how thoroughly that claim has been checked — a stated claim, a lab-tested one, or one with a real track recordhow much weight to put on the claim before relying on it
Confidence score + "below 95% gate" tagthe result of the most recent test run for that holderwhether it's reliable enough for real work, or Lab-only for now
Cost / unitwhat it costs each time that model or agent does the jobweigh cost against confidence when choosing who does the work
95% honesty gate notethe ledger's rule that a weak result is always shown, never hiddenyou can trust that silence means "not proven," not that it was left out