Evaluation
Stage the action. Hold when it is not proven.
Blackbook is the assurance protocol for promoting a model, a plan or an actuator from candidate to live action. Ambiguity holds. Verification precedes use. Blackbox keeps the receipt.
Overview
Doctrine, authority, protocol, evidence.
Evaluation here is how Mediator stages deterministic fail-closed action on robotics, AI and deep reasoning. Blackbook supplies the law: structure precedes action, verification precedes use, stop precedes overreach. AOS binds intent, authority, source and falsifiers before a run settles. Reaper executes authorized cyber evaluation. Ghoul keeps identity, lifecycle and recovery intact while cognition changes. Blackbox Systems is the evidence machinery. Protocol without evidence is theater. Vocabulary lives under AI resources. Staging lives here.
Blackbook
The operations manual.What goes, who may act, when they must not, and what happens when the book cannot be enforced.
AI and machine learning
Authorized evaluation of models that can act: tool grants, code generation, N-day class findings, ATT&CK-shaped mapping.Reaper owns the test.
Deep reasoning
Multi-step plans, persistent agents and shared world-state.A plan is not issuance. Confidence is not verification.
Robotics and physical AI
Fail-closed promotion from simulation to actuator.Blackwell video is in hand. NVIDIA Cosmos is the next world-model delta.
Cryptographic evidence
The evidence layer of the stack.Hashes, signatures, manifests and independent verification as separate properties.
Evaluation graphs
How evaluation is drawn: nouns, edge kinds, verbs, zoom, and when a picture must yield to a receipt.
Notes
Fractionalization, Graphite families, effort governor, integrity map and NVIDIA AISimulate cutover.
AI resources
Public vocabulary with operating boundaries: AI classes, generative models, agentic systems and MCP.
Reaper
Authorized testing with written Rules of Engagement, reproducible findings and bounded retest.
Proof
Capability described. Result not invented.
These pages are evaluation lanes and use cases we can build into. They are not claimed publications and not a substitute for a written engagement.