Current official AI guidance points to a practical control gap: a model can be technically capable while its intended workflow, evidence, human oversight and change history remain insufficiently defined for a dependable operating decision.
Observed fact: an AI model is not the same thing as an operating use case
Applied-AI discussions often begin with a model name, benchmark result or interface demonstration. Those are relevant inputs, but they do not by themselves define the work the system will perform, the information it may receive, the decision it will influence or the consequence of an error. NIST's updated AI Risk Management Framework Playbook retains a Govern, Map, Measure and Manage structure and describes these functions as a practical route for incorporating trustworthiness considerations into AI activities [1]. The important factual point is that risk management is framed around the surrounding activity, not only around model selection.
That distinction matters when a prototype moves into a routine workflow. A model can produce a plausible output in a controlled demonstration while the production process still lacks an authorised input source, retention rule, escalation path or named reviewer. It may also depend on vendor configuration, retrieval content, prompts or integrations that change after the initial test. A deployment case therefore needs to describe the bounded task, the operating environment and the controls that apply if the system behaves unexpectedly. This is a control question, not a claim that any particular model is unsuitable [1].
Observed fact: documentation and downstream information are part of the regulatory direction
The European Commission's General-Purpose AI Code of Practice provides a current official example of this wider evidence expectation. Its Transparency and Copyright chapters are intended to support providers of general-purpose AI models with obligations concerning technical documentation, information for downstream providers, copyright policy and a public summary of training content [2]. The Code is a voluntary compliance tool, and its scope is not identical to every organisation that deploys an AI feature. Even so, the categories are useful because they connect a model to evidence that another party can review and use.
For an operating team, downstream information changes the diligence question. Instead of recording only that a vendor uses a named model, the record can identify the supplied technical information, permitted purpose, known limitations, interface dependencies, data-handling conditions and the point at which a material change must be reassessed. That is not a substitute for jurisdiction-specific legal advice or for contractual diligence. It is a practical way to avoid treating an opaque marketing description as sufficient operating evidence [2].
A deployment record should make the workflow testable
The OECD's 2026 Due Diligence Guidance for Responsible AI translates high-level principles into a risk-based process: embed responsible AI in policies and management systems, identify and assess impacts, prevent or mitigate adverse impacts, track results, communicate how impacts are addressed and provide or cooperate in remediation where appropriate [3]. This is an observed description of the OECD framework, not evidence that every deployed system has the same level of risk. It does show why a single policy statement or supplier questionnaire cannot establish that a workflow is under control.
A compact deployment record can make the framework operational. It can state the workflow boundary; categories of permitted and prohibited inputs; expected output and failure modes; the evidence used in evaluation; a human review or escalation rule; accountable owner; supplier and integration dependencies; and the revalidation trigger for model, prompt, data or process changes. Versioning this record makes later comparisons possible: a reviewer can see whether performance, scope or safeguards changed rather than assuming that a prior pilot approval remains applicable [1][3].
- Scope: identify the business task, users, decision consequence and exclusions before measuring performance.
- Evidence: retain representative evaluation cases, their provenance, acceptance thresholds and known blind spots.
- Ownership: name the workflow owner, technical maintainer, reviewer and escalation recipient rather than assigning responsibility to the model.
- Change control: define what model, data, prompt, integration or policy change requires a renewed assessment.
Oakhampton inference: evidence should travel with the use case
Oakhampton's inference is that the useful unit of governance is the deployment case rather than the model inventory alone. A model inventory can show what technologies an organisation has approved, but it cannot on its own show which documents are processed, which operational decision is influenced, what review occurs or whether a later configuration change altered the result. Joining the model record to a concise, versioned workflow case gives decision-makers a traceable route from stated purpose to input controls, evaluation evidence, accountability and remedial action [1][2][3].
This approach also separates facts from assumptions. Facts can include the provider's published documentation, the tested workflow version, observed evaluation results and the named control owner. Assumptions can include the expected volume, the stability of retrieval content, the availability of human review and the relevance of a test set to future cases. Keeping those labels distinct prevents a deployment from inheriting confidence merely because a demonstration was persuasive. It also makes review more efficient because challenged assumptions have an obvious place in the record.
Remaining uncertainty belongs in the release decision
Public guidance does not settle whether a specific system is reliable for a particular transaction, customer interaction, safety-sensitive process or regulated decision. NIST's Playbook is voluntary guidance, the Commission's Code applies within a defined general-purpose AI framework, and the OECD guidance is risk-based rather than a universal technical test [1][2][3]. Their existence should not be presented as certification of a tool or as a forecast that every deployment will create the same legal, operational or commercial exposure.
The bounded conclusion is operational. Treat a proposed AI capability as ready for wider use only to the extent that its actual workflow, evidence, limits, owners and change triggers can be identified and reviewed. Where any of those remain unverified, record the gap and retain an appropriate human or process control while the evidence is improved. That preserves room for legitimate experimentation without converting an early demonstration into an unsupported claim of dependable operating performance.
Sources
- NIST AI RMF PlaybookNational Institute of Standards and Technology · 10 June 2026
- General-Purpose AI Code of PracticeEuropean Commission · 10 July 2025
- OECD Due Diligence Guidance for Responsible AIOrganisation for Economic Co-operation and Development · 19 February 2026