Agent competency assurance
Model capability is not the same thing as deployed-agent competence.
Assure the deployed system, not merely the model.
An agent's real competence depends on the model plus its harness, tools, memory, permissions, environment, supervision and recent performance. The useful analogy is regulated human competency assurance: define the role, required competence, evidence standard, authorisation and currency.
Capability ≠ competence ≠ authorisation ≠ currency.
What supports it
Agent evaluations increasingly need to inspect the full deployed system — prompts, retrieval, tools, traces and environment — rather than a model score alone.
Where the analogy breaks
Agents can be copied, updated and instrumented differently from people. Their evidence may be more continuous, but configuration drift can invalidate assurance quickly.
What to watch
Whether enterprise governance begins defining task-specific agent permissions and evidence thresholds in the same way regulated work defines competency and authorisation.
What would change our view?
If robust model-level evaluation proved sufficient across changing tools and environments, the case for configuration-specific assurance would weaken.
