Governance, Risk & Responsible Use: A Crash Course
Four questions. Three possible answers. One habit that keeps useful AI work safe enough to survive.
Before AI does meaningful work, ask four questions:
- The Case: Can AI do this work?
Yes · Yes, with a defined human review · No - The Data: Can this information go in?
Yes · Check first / add a control · Not through this route - The Capability: Can I turn this tool, Skill, connector, or action on?
Enable · Escalate for review · Decline - The People: Could this affect someone unfairly or require disclosure?
Decide and document · Disclose · Escalate the question
The middle answer is where most professional judgment lives, because it requires you to name a reviewer, control, approved route, disclosure, or escalation.
How to use this course
This course is written for three kinds of reader at the same time.
- If you are new to AI governance, read the main text and the examples. You can skip every expandable Professional detail block on your first pass.
- If you are a knowledge worker or manager, read the main text plus the Governance Record and evidence sections. Your goal is to turn the framework into a repeatable work habit.
- If you build agents, integrations, or automated workflows, read the For agent builders notes. They translate the same questions into permissions, actions, monitoring, and system design.
The core path takes roughly 35–45 minutes, plus about 15 minutes to create a Governance Record for one real workflow.
Do not paste confidential, personal, regulated, privileged, or otherwise restricted work material into an AI tool just to complete an exercise in this course. Use a fictional example, synthetic data, or an entry point your organisation has already approved for that data.
An AI assistant can help you apply a framework. It is not your policy owner, security reviewer, compliance function, legal adviser, or final authority on whether a particular data type or tool is approved.
📚 Teaching Aid
View Full Presentation for Governance, Risk & Responsible Use
Start with an ordinary Tuesday
The following is a fictional composite scenario based on a common pattern in workplace AI adoption.
A logistics company has a good year with AI. Operations cuts reporting time. Marketing drafts campaigns faster. Finance automates routine matching. The tools become ordinary enough that people stop thinking of each use as an "AI decision."
Then a project manager uploads a customer spreadsheet containing names and account numbers into an AI chat she uses every day. She is not trying to break a rule. She is trying to answer a director's question before lunch.
The company responds by freezing AI use while it investigates. Teams that did nothing wrong lose useful workflows because one routine decision created a risk the organisation could not ignore.
That is the problem governance is meant to solve.
Governance is not mainly a policy document. A policy cannot look at the file on your screen. It cannot decide whether the task needs the customer names. It cannot notice that a new connector can now send messages as well as read them.
Those decisions happen where the work happens.
The goal of this course is therefore not to make you a compliance specialist. It is to give you a small decision procedure that ordinary workers can run quickly, managers can review, and builders can turn into controls.

What you should be able to do by the end
You should be able to:
- classify an AI use case as fully appropriate, appropriate with human review, or inappropriate,
- define a real human-review gate as who checks what, and when
- classify data before uploading it, choose an approved route, and distinguish redaction from genuine anonymisation
- evaluate a Skill, connector, or tool by source, reach, fit, outside content, and actions
- spot fairness and disclosure issues, especially when AI filters or ranks people
- document the decision in a one-page Governance Record
- respond quickly and transparently when a mistake, near miss, or harmful output occurs.
Part 1: The four questions
1. The Case: Can AI do this work?
Start here. If AI should not be doing the work, the later questions about data and tools do not matter.
You do not need a complicated risk matrix for most everyday cases. Ask four questions.
| Screen | Plain-English question |
|---|---|
| Reversibility | If the output is wrong, can we catch and undo it before harm happens? |
| Consequence of error | What happens if it is wrong? Is the cost trivial, expensive, harmful, regulated, or irreversible? |
| Human judgment or empathy | Does the task depend on relationship, care, original judgment, or context a person should personally own? |
| Accountability | Who is answerable for the outcome, and can that person meaningfully review and own what the AI produced? |
1.1 Choose one of three answers
Fully appropriate. The work is low consequence, reversible, and easy to review in the normal course of work. AI can do it without a special governance gate.
Examples: drafting an internal FAQ from approved source material, restructuring notes, creating first-pass meeting summaries, or generating alternative headings.
Appropriate with human review. AI is useful, but a specific human check is required before the output is used, sent, published, or acted on.
Examples: drafting customer replies about billing, preparing a management summary from financial data, or creating a shortlist that influences who receives attention next.
Inappropriate. The task should remain human-owned because the consequence, irreversibility, or human responsibility cannot be repaired by review.
Examples include making a final medical, legal, disciplinary, or other professional determination where a qualified person must make and own the judgment. AI may sometimes support the work, but it should not become the final decision-maker simply because it is capable of producing an answer.
1.2 Find the deciding factor
Professionals often use formal language such as load-bearing criterion. In plain English, this means the deciding factor.
Ask:
If one of the four answers changed, which change would move this use case into a different classification?
That is the factor carrying the decision.
For a condolence message, the decisive factor may be the human relationship even though the message is easy to rewrite. For a customer billing reply, accountability may be decisive because the company remains responsible for what it tells the customer about their money.
Naming the deciding factor improves the quality of the discussion. Instead of "this feels risky," you can say:
"This is appropriate with review because the organisation remains accountable for the factual claim."
That sentence can be checked, challenged, and improved.
1.3 A human in the loop is not a gate
The phrase "a human will review it" sounds responsible and often means nothing in practice.
A defined gate has three parts:
- Who reviews: the role that actually holds the responsibility
- What they verify: the specific risk the review is meant to catch
- When they review: before the output becomes hard to undo.
Compare:
"Someone will check it."
with:
"The account manager checks delivery facts against the dispatch record before the report is sent to the customer."
The second statement is a control. The first is an intention.

If you cannot write the review step in who / what / when form, classify the workflow as not ready to run yet rather than pretending "human oversight" exists.
Professional detail: consequence and accountability are different
A high-consequence task is not automatically prohibited. Sometimes a well-designed gate restores sufficient control. Conversely, a low-consequence task can still require a human because the human relationship is the point of the task.
Do not score the four criteria mechanically and add them up. Use them to expose the reason the case belongs in one category rather than another.
Also remember: accountability cannot transfer to a tool. Delegating a task does not delegate accountability. A manager, clinician, lawyer, employer, or company does not become less responsible because an AI system drafted part of the work.
Worked examples
| Use case | Classification | Why | Gate, if needed |
|---|---|---|---|
| Turn approved policy documents into an internal FAQ draft | Appropriate | Reversible, low consequence, authoritative sources exist | Normal editorial review |
| Draft a customer response to a billing complaint | Appropriate with review | Company remains accountable for facts about the customer's account | Support agent verifies account facts and tone before sending |
| Generate a final professional determination | Inappropriate as the final decision-maker | Professional accountability and consequence cannot be transferred to the tool | Human professional makes and owns the determination |
| Summarise candidate applications to help organise reading | Appropriate with strong review and policy checks | Consequential to applicants and sensitive to unfair filtering | Hiring owner reviews inclusion and exclusion decisions before action |
2. The Data: Can this information go in?
A perfectly reasonable AI task can still be run on data that should not be placed in that tool or route.
Use this order:
Classify the data → ask whether the task needs the identifying details → confirm the route → choose the control.
Do not start with a privacy feature and work backwards.
2.1 Use three practical tiers
Your organisation may use different labels. Translate these to your policy rather than replacing it.
Green: generally permitted. Public material, genuinely anonymised or aggregated data, and internal material approved for broad use in the relevant AI environment.
Yellow: check first or add a control. Internal-only documents, personal contact information, customer or employee identifiers, confidential drafts, unannounced deal or product information, or other material whose use depends on the approved route and configuration.
Red: do not use through an unapproved route. Credentials, secrets, highly regulated or specially protected data, privileged material, or third-party confidential information you are not authorised to disclose.

If you are unsure between two tiers, use the more sensitive one until somebody authorised to interpret the policy confirms otherwise.
2.2 Ask the most useful data question
Before deciding a sensitive task is impossible, ask:
Does the task actually need the identifiers, or only the pattern?
This is the key branching question.
If you are analysing spending trends, the model may not need customer names. If you are reconciling a particular account, the identifier may be essential.

2.3 Redaction is useful, but it is not magic
When the task does not need the sensitive fields, remove them before the data reaches the model.
Redaction fails in two common ways.
Failure 1, partial redaction: you removed the obvious identifier but left enough clues to identify the person. A rare job title, exact date, region, age, or combination of ordinary facts can identify somebody as reliably as a name.
Failure 2, redaction that breaks the task: you removed information the task actually needs. The result is safe-looking but useless. If individual account identifiers are necessary to reconcile transactions, redaction is not the control for that task.
There is also an important distinction:
- Pseudonymised data replaces direct identifiers with labels while somebody still holds a way to map them back.
- Anonymised data cannot reasonably be traced back to the person under the standard your organisation uses.
Do not lower the data tier simply because names were replaced with "Person A" or "Customer 17."
2.4 The route matters as much as the tool name
An entry point is the specific route by which material reaches an AI system: a normal chat, a project/workspace, an incognito or temporary conversation, a file upload, a connector, an API, a desktop agent, or another interface.
Different routes can have different retention, access, logging, admin, or contractual properties.
So the useful question is not:
"Is this AI product approved?"
It is:
"Is this route approved for this data and this purpose?"
For sensitive or regulated data, that approval should come from the relevant administrator, policy owner, security function, compliance team, or contract. It does not come from a guess at the keyboard.
2.5 Privacy and workspace controls answer narrower questions
Features such as temporary conversations, memory controls, retention settings, sandboxed execution, encryption, projects, or organisation-managed workspaces can reduce particular risks. They do not answer the earlier authorisation question by themselves.
Think of controls this way:
| Control | What it may help with | What it does not decide |
|---|---|---|
| Temporary/incognito conversation | History or memory persistence | Whether the data was permitted to enter the system |
| Memory controls | Cross-session reuse of information | Whether the original upload was allowed or how long all underlying data is retained |
| Code-execution sandbox | Limiting where code or file-processing can operate | Whether the data was authorised for that environment |
| Project/workspace | Keeping shared knowledge, instructions, files, and related work together | Whether every file or data type is approved for that workspace |
| Organisation-managed account and admin controls | Roles, feature controls, retention, exports, audit, or compliance options depending on the plan and configuration | Blanket approval for every workflow, data type, or audience |
A sandbox is an execution boundary, not an approval boundary. It can reduce what code or file-processing can reach outside an isolated environment, but the information still entered a system and outputs can still leave it. Ask separately whether the data is allowed in that route and whether the resulting files or actions are allowed out.
2.6 Claude routes and controls: six distinctions
Use these as routing distinctions, not as a substitute for your organisation's policy or contract.
- Memory is not retention. Memory affects what Claude can carry forward across work. Chats and memory-related data still follow the applicable retention and export controls.
- Incognito is not zero retention. Incognito chats do not appear in normal chat history or memory, but Anthropic currently documents a 30-day default retention period, with longer custom retention possible on Enterprise. Incognito is currently available only outside Projects.
- A sandbox is not data authorisation. Execution isolation can reduce the reach of code. It does not prove that a sensitive file was allowed into that environment.
- A Project is a route, not a permission shortcut. Project knowledge can be reused across chats, and Team/Enterprise Projects can have different visibility or sharing settings. Check the files, audience, configuration, and policy.
- Organisation-managed does not mean universally approved. Team and Enterprise add administrative controls. Enterprise adds further retention, audit, and compliance options. Approval still depends on the particular data, route, purpose, configuration, agreement, and organisational rules.
- Private in the interface does not mean absent from organisational records. Team/Enterprise provide organisational data-export mechanisms, and Enterprise adds additional compliance and audit mechanisms. Treat administrative visibility and exportability as part of the route assessment.
Product behaviour changes quickly. If a decision depends on one of these features, verify current Anthropic documentation and your organisation's configuration.
Worked examples
Meeting notes with no confidential material. Usually green or ordinary internal use, subject to your company's approved AI environment.
Internal sales pipeline with customer names. Usually yellow. Ask whether names are needed. If not, remove them. If yes, confirm the approved route.
Credentials or API secrets. Red. Do not paste them into an ordinary AI conversation. Use your organisation's approved secret-handling mechanism.
Patient, financial, privileged, or other specially protected records. Treat as red until your organisation has explicitly approved the specific route and use case. Do not infer permission from the fact that the AI product has an enterprise plan or security page.
3. The Capability: Can I turn this on?
AI systems increasingly do more than generate text. They can use Skills, tools, plugins, connectors, web pages, files, code, browsers, email, calendars, CRMs, and other services.
The question is no longer just "do I trust the model?" It is:
What authority am I adding to this session or agent?
3.1 Treat extensions like software, not like prompts
A Skill may contain instructions, scripts, resources, dependencies, or procedures. A connector may expose data or actions in another system. A plugin may combine both.
Do not judge these by the friendliness of the description or the fact that a colleague shared them.
One property of Skills matters more than any other, and it is the one people get wrong.
A Skill does not carry its own permission list. A connector has scopes that you grant and can narrow. A Skill has no such dial. It runs with whatever access the session it runs in already has, so its reach is everything that session can reach, not only what its stated task needs. Where a Skill's definition carries a list of tools, that list pre-approves them so they run without prompting you. It widens what happens silently. It does not narrow what the Skill can touch.
This is why the second check below asks what the Skill could reach, rather than what it is asking for. It is not asking for anything.
Run five checks:
- Source: who created or published it?
- Reach: what data, files, systems, tools, or credentials could it touch in the environment where it runs?
- Fit: is that reach proportionate to the job you actually need done?
- Outside content: will it read web pages, incoming email, customer files, shared documents, or other content you did not author?
- Actions: can it send, pay, delete, publish, edit, approve, or otherwise cause a hard-to-reverse change?
The first three tell you whether the capability belongs in the workflow. The last two tell you how dangerous a mistake or hostile instruction could become.
3.2 Use three outcomes
Enable. You know the source, the reach is proportionate, the task fits, and the actions are controlled.
Escalate. The capability may be useful, but you cannot establish something important: the source is uncertain, the reach is broad, the tool wants unnecessary access, or the security implications are beyond your role.
Decline. The capability is clearly disproportionate, untrustworthy, or unnecessary for the job.

Escalating is not "asking permission to be careful." A good escalation names the exact unresolved question:
"This tool is useful, but it requests write access to our shared drive even though the workflow only requires reading. Can security review whether that access can be narrowed?"
That is far better than "Can I use this?"
3.3 Trusted tool, untrusted content
A capability can come from a trusted publisher and still read content written by somebody you do not trust.
An email, web page, customer upload, or shared document can contain instructions that attempt to steer the model. This family of attacks is commonly called prompt injection.
The key insight for non-specialists is simple:
Who wrote the tool and who wrote the content are two different trust questions.
The risk becomes much higher when an AI workflow both reads untrusted content and can take consequential actions.

For ordinary knowledge work, a useful default is:
- let AI read only what it needs
- prefer read-only access when it is enough
- keep send, publish, pay, delete, approve, or irreversible edits behind a defined human gate until the workflow has earned a higher level of autonomy.
At system level, "reach" becomes architecture: scoped credentials, tool allow-lists, typed actions, confirmation policies, network restrictions, isolation, audit logs, and clear separation between reading content and executing actions.
The governing principle is least privilege: give an agent the narrowest authority that lets it complete the job, then revisit that authority when the job changes.
Professional detail: scanning helps, but it is not approval
Security scanning can raise the floor by detecting some malicious Skills or plugins. It does not establish that the capability is appropriate for your data, environment, or intended use.
Anthropic's current Enterprise documentation, for example, describes skill/plugin scanning as a beta check for malicious content on certain new uploads/edits and explicitly notes that a pass is not a guarantee of safety in every respect.
Keep the source/reach/fit review even when a scanner passes the package.
Worked examples
Document-formatting Skill from a trusted publisher. Source is known, purpose matches the job, and it does not introduce unnecessary authority. Enable in an approved environment.
Third-party "analytics booster" with broad instructions and unknown source. Useful idea, unresolved trust and reach. Escalate with the specific concerns instead of experimenting on real company data.
Read-only connector to a reporting database. If the route is approved and the access is limited to the tables required, enable. Do not grant write access "just in case."
Agent that reads incoming email and can automatically send external replies. Outside content plus consequential action. Require stronger controls and a human gate unless the organisation has deliberately designed, tested, and approved a higher-autonomy workflow.
4. The People: Could this affect someone unfairly or require disclosure?
The first three questions mostly protect the organisation and its information. This one makes you look outward.
Ask:
- Who is affected, including people who never see the output?
- What could go wrong for them?
- Would they be able to notice or challenge it?
- What would a fair process look like?
- Is disclosure required, or would AI involvement reasonably matter to them?
4.1 Look at who was excluded
One of the easiest AI risks to miss occurs when the system narrows a set and humans inspect only the survivors.
Examples:
- candidate shortlists
- fraud or risk flags
- support tickets selected for escalation
- leads selected for attention
- documents selected as relevant
- people or cases ranked for priority.
If AI systematically removes a group, nobody notices if the review looks only at what remains.
A practical control is to sample exclusions as well as inclusions.
This is especially important in high-impact decisions involving employment, access, benefits, education, finance, health, or other significant opportunities. Your organisation may also have legal or policy restrictions on automated decision-making in these areas. The framework here does not replace them.
4.2 Disclosure: use rules first, judgment second
Ask in this order.
First: is disclosure required? Check law, policy, contract, professional rules, client commitments, and organisational standards. If one of them requires disclosure, the decision is already made.
Second: if no rule answers it, would AI involvement reasonably change this person's understanding of the work or the relationship?
AI that reformats a table is different from AI that drafts what an employee will understand as a manager's personal assessment. Routine tooling does not always need a label. Consequential or relational work often deserves more transparency.
When you remain genuinely uncertain, escalation is better than inventing a rule.
4.3 Escalate the question, not a verdict
Weak escalation:
"I think this is fine. Can you approve it?"
Strong escalation:
"Here is the workflow, who is affected, the control we added, and the point the framework does not settle: whether our disclosure obligation applies to this audience."
A strong escalation gives the reviewer something specific to decide.
Worked example: performance-review drafting
A manager wants AI to turn their own notes into draft performance-review summaries.
Case: appropriate with human review. The manager must own the assessment and verify every statement against the underlying notes.
Data: employee information is sensitive internal material. Use only an approved route and minimise unnecessary personal details.
Capability: ordinary drafting may need no new connector. If a system pulls data automatically from HR tools, access and scope require much stronger review.
People: employees are directly affected. The manager should check consistency across employees, not merely whether each individual paragraph sounds plausible. Disclosure obligations depend on policy, contract, and context.
The point is not that AI makes performance reviews automatically wrong. The point is that the workflow needs controls shaped around the way the output can affect people.
Part 2: Put the four questions together
5. One ordinary workflow, start to finish
Ayesha leads operations at a regional logistics company. Her team manually prepares a weekly service-exception report for large customers. She proposes using AI to draft the report from dispatch records, then having the account manager send it.
5.1 The Case
- Reversibility: high before sending. Low after sending.
- Consequence of error: meaningful because delivery failures may affect contracts and customer trust.
- Human element: moderate. The summary is mechanical, but tone matters on disputed accounts.
- Accountability: the company remains responsible for every factual claim.
Classification: appropriate with human review.
Deciding factor: accountability.
Gate: The account manager checks delivery facts against the dispatch record and reviews the tone on disputed accounts before the report is sent.
5.2 The Data
Inputs include customer names, shipment information, delivery windows, failure reasons, and some free-text operational notes.
This is at least yellow under the simple model because it is internal/customer data and contains identifiers.
Does the task need the identifiers? Yes: a customer report must identify the customer and shipments.
So the team does not solve the issue by stripping every identifier. It confirms that the specific workspace and integration route are approved for this information.
5.3 The Capability
The team wants a connector to the dispatch system.
- Source: internal platform team.
- Reach: only the reporting tables required.
- Fit: strong. It eliminates a manual export step.
- Outside content: some free-text fields contain text copied from customer communications.
- Actions: the workflow does not need permission to send the report.
Outcome: enable read-only access after scope confirmation. Keep sending with the account manager.
5.4 The People
Customers are affected, but so are drivers and operations staff described in failure notes.
A summary could accidentally turn "weather delay" into "driver failure," or apply harsher language to some teams than others.
Add a fairness/accuracy check: attribution must match the dispatch record, and the team should periodically compare phrasing across reports for consistency.
Disclosure: check contractual and organisational requirements first. If they are silent, AI drafting of a routine operational report may not require a special customer notice, but employee-facing consequences deserve separate consideration.
5.5 The Evidence
A control is stronger when you can tell whether it is working.
For the first quarter:
- Success measure: account managers complete the factual/tone check before sending
- Failure threshold: any material AI-generated factual error reaches a customer
- Monitor: Ayesha reviews a sample monthly
- Residual risk: a consistent wording bias across all reports could survive individual checks, so the sample review compares across customers as well as within each report.

One modest workflow has now produced a classification, a gate, a data route, a permission boundary, a fairness check, and a monitoring plan. That is governance as work design, not paperwork added after the fact.
6. The Governance Record: one page that survives the meeting
The four questions are useful in your head. They become organisationally useful when the answers are written down.
A Governance Record is a one-page record for one workflow.

Block 1: The Case
- Classification: appropriate / appropriate with review / inappropriate
- Deciding factor
- Gate in who / what / when form
Block 2: The Data
- Data tier
- Does the task need identifiers?
- Approved route
- Fields removed or minimised
Block 3: The Capability
- Skills, connectors, tools, or actions enabled
- Source and reach check
- Outside content it may read
- Actions it must not take without review
Block 4: The People
- Who is affected
- Fairness check
- Disclosure decision
- Open question or escalation, if any
Evidence line
- Success measure
- Failure threshold
- Monitoring owner and cadence
- Residual risk: what could still go wrong if every control works as designed?
Footer
- Workflow owner
- Date
- Re-check if: model, feature, data, audience, permission, policy, vendor term, or business consequence changes
A blank field is better than a guessed one. If the owner or approved route is unknown, write OPEN QUESTION and route it to the right person.
7. What to do when something goes wrong
Governance is not the promise that nobody will ever make a mistake. It is the ability to surface mistakes early enough to contain them.
Common incident types include:
- Data: information entered a route that was not approved for it
- Output: an AI-assisted output went out with a material error or harmful framing
- Action: an AI tool sent, changed, deleted, approved, or shared something it should not have
- Security: untrusted content steered an AI system or tool in an unintended way
- Fairness: a ranking, filtering, or screening process appears to have disadvantaged people and the issue was not caught by review.
The first-hour procedure
- Stop the spread. Do not forward, repost, or create unnecessary new copies.
- Record the facts. What happened, what data/output/action was involved, which route or system, when it happened, and whether anything went onward.
- Report promptly through your organisation's incident path. For sensitive or regulated cases, timing may matter legally, so do not wait until you have a perfect explanation.
- State the facts plainly. Avoid speculation and self-defence. The incident owner needs a reliable account.
- Follow the incident owner's instructions. Do not independently decide questions such as deletion, notification, disclosure, or evidence preservation when they may have legal, contractual, security, or forensic consequences.

Do not quietly delete evidence and hope the issue disappears. Deleting your visible copy may not delete organisational or vendor records, and it can interfere with investigation or required preservation.
Report near misses too. A near miss reveals where the process is confusing or the approved route is too hard, without the cost of a full incident.
How the organisation treats the first person who reports a good-faith mistake determines whether the next ten mistakes are reported or hidden. Accountability matters, and so does creating a reporting path people can actually use.
Part 3: Make the habit survive real work
8. Why governance drifts
High-stakes decisions often receive attention because everybody knows they are important. Routine decisions are more dangerous to governance because each one feels too small to count.
A person uses a slightly easier tool. A human review becomes a glance. A connector keeps permission after the project changes. A sensitive field starts appearing in a dataset that used to contain none.
No single moment looks like a programme-level failure. The accumulated distance between the intended process and the actual process is the risk. That distance has a name worth using in a review: the Diligence gap.

8.1 Run a usage audit on the work people actually do
A lightweight team audit should answer:
- Which AI workflows are actually in use?
- What data is actually going through them?
- Did defined human gates really run?
- What new tools, Skills, connectors, or permissions were added?
- What changed in the model, feature, route, data, or audience?
A reasonable starting point for many teams is a small quarterly sample, with more frequent review during a new rollout or after gaps are found. Do not turn this into a universal compliance requirement: frequency should match the risk, policy, and sector.
8.2 Audit the process, not the person
If a governance audit becomes a hidden performance review, people will show auditors only the cleanest work. The organisation then loses visibility into the behaviour it was trying to understand.
The purpose is to find system gaps: confusing rules, unnecessary friction, weak gates, stale permissions, or practices that evolved without review.
8.3 The friction rule
If the approved path takes ten steps and the unapproved path takes one, ordinary people under deadline pressure will discover the one-step route.
That is not a reason to excuse policy violations. It is a reason to treat usability as part of the control system.
When you find repeated workarounds, ask:
"What makes the approved way harder than the unsafe way?"
Fixing that friction may produce more risk reduction than another reminder email.
8.4 When there is no AI policy
If your organisation has no AI policy, do not invent one and present it as official.
Instead, create an interim proposal for the appropriate owner to endorse. A useful one-page starting point can contain:
- approved AI products and specific routes for work data
- a short list of data types that require confirmation before use
- the human-gate rule for external or consequential outputs
- the review rule for new Skills, connectors, and high-authority tools
- one named role or channel for questions and incidents.
In regulated environments, that interim guide does not substitute for legal, compliance, security, privacy, or contractual approval.
8.5 Re-check when the workflow changes
Calendar reminders help, but change is the stronger trigger.
Re-check the Governance Record when any of these changes:
- model or model family
- AI product feature or retention behaviour
- connector, Skill, plugin, permission, or action
- data type or sensitivity
- source of outside content
- audience or population affected
- consequence of error
- law, policy, contract, or vendor terms.
"Approved last year" is not a reason to skip review if the thing that was approved is no longer the same thing.
9. For professionals: evidence, residual risk, and review quality
You can use the four-question framework without this section. This section is for managers, governance leads, risk owners, and people who need to defend the design later.
9.1 Add evidence to every important control
A control stated only as "the manager reviews it" is a belief unless you have some way to see that the review occurs and catches the intended problem.
For meaningful workflows, write four things:
- Success measure: what would indicate the control is operating?
- Failure threshold: what outcome should trigger intervention or suspension?
- Monitoring owner: who looks at the evidence, and how often?
- Residual risk: what can still go wrong even when the control works perfectly?
Residual risk is not a confession that the design failed. Every real control leaves something outside its boundary. Naming that remainder is what lets the appropriate owner decide whether it is acceptable.
9.2 Review quality matters more than review volume
A human gate does not make a workflow safe merely because a person clicked "approve."
Good review has:
- access to the source material needed to verify the output
- enough time to perform the check
- a specific risk to look for
- authority to reject or correct the output
- a workflow position before the outcome becomes hard to reverse.
If reviewers routinely approve everything, investigate whether the control is unnecessary, badly placed, poorly specified, or impossible to perform, not simply whether the reviewers need another reminder.
9.3 Sample the edge cases
Random samples are useful, but high-risk systems should also sample cases where failure is more likely or more expensive: unusual inputs, minority categories, escalations, low-confidence outputs, exclusions, or cases near thresholds.
For AI systems that change over time, evaluation should eventually move beyond manual sampling into structured test sets, baselines, monitoring, and regression checks. Those belong in the evaluation material of the Agent Factory book. The Governance Record is the bridge to them.
10. For agent builders: translate the four questions into architecture
The same framework changes shape when AI acts through software rather than waiting for a person to copy and paste.
The Case becomes an autonomy level
Do not ask only "can the agent do this?" Ask "what level of autonomy has this task earned?"
A useful progression is:
- draft only
- propose an action for confirmation
- act within a narrow reversible boundary
- act autonomously only where reliability, monitoring, and rollback justify it.
High-impact or irreversible actions should have stronger confirmation and recovery design.
The Data becomes custody and minimisation
Design which data the agent may read, where it can be stored, what enters prompts or logs, how long it is retained, and whether sensitive fields can be excluded before the agent sees them.
Do not rely on users remembering to redact what the architecture can remove automatically.
The Capability becomes scoped authority
Prefer:
- narrow credentials
- read-only where sufficient
- explicit tool allow-lists
- domain/network restrictions where relevant
- typed tool schemas and validation
- confirmation for consequential actions
- audit trails for tool use
- separate identities for agents rather than shared human credentials
- revocation and rotation paths.
The People becomes evaluation and recourse
If an agent affects people, design for:
- fairness testing across relevant groups and cases
- review of exclusions and false negatives, not only successful outputs
- meaningful disclosure where required
- a way to appeal, correct, or reverse consequential outcomes
- monitoring after deployment, not only before launch.
The practitioner framework and the system architecture should tell the same story. If the Governance Record says "human approves before send" but the API token can send without confirmation, the architecture has already overruled the policy.
11. Six common failure patterns
1. Forcing every question into yes or no
This produces both recklessness and unnecessary abandonment. Always check whether a workable middle option exists.
2. Writing "human in the loop" without defining the loop
Name who / what / when or admit the gate does not yet exist.
3. Treating privacy mode as data authorisation
A temporary or incognito mode may reduce persistence. It does not decide whether the data was allowed to enter the system.
4. Assuming an extension can only do what its description says
Check source, reach, fit, outside content, and actions. Friendly intent is not a permission boundary.
5. Reviewing only the people or cases that AI selected
If AI filters, ranks, or triages, sample what it excluded as well.
6. Abandoning useful work before asking whether sensitive fields are necessary
If the task needs only the pattern, data minimisation can turn a difficult problem into an ordinary one. If the task truly needs the identifiers, use the approved route rather than pretending redaction can solve it.
Part 4: Use it tomorrow
12. The one-minute checklist
Before AI does meaningful work, run this:
| Question | Quick check | Middle answer requires |
|---|---|---|
| Case | Can AI do this task responsibly? | A defined human gate: who / what / when |
| Data | Can this information go through this route? | A control, data minimisation, or confirmed approved route |
| Capability | Should this tool or authority be enabled? | A specific security/admin review question |
| People | Could this affect people or require disclosure? | A fairness check, disclosure, or escalation |
Then ask one final question:
What changed since the last time we decided this was okay?
If nothing changed, continue. If something material changed, re-check the relevant block.
13. Build a Governance Record in 15 minutes
Choose one real workflow you own or influence. Use fictional or approved data while working through the prompts.
Step 1: The Case
Write, in your own words:
- what AI does
- what the human still does
- whether the workflow is appropriate, appropriate with review, or inappropriate
- the deciding factor
- the gate in who / what / when form.
You may ask an AI assistant to challenge your reasoning, but do not ask it to "approve" the workflow. A useful prompt is:
Here is a workflow I am thinking of handing to AI: [describe it, using a fictional or already-approved example].
Push back on my reasoning. Screen it on reversibility, consequence of error, how much human judgment it needs, and who stays accountable. Tell me whether it looks fully appropriate, appropriate with a human review, or inappropriate, and say which single factor is actually deciding that.
If it needs review, write the gate as who checks what, and when. Do not tell me the workflow is approved or compliant, and list anything you would need me to confirm with my organisation.
Step 2: The Data
Record:
- what data enters
- the tier under your organisation's policy
- whether identifiers are actually needed
- which fields can be removed
- the approved route if sensitive fields remain.
For redaction practice, use synthetic data that resembles the shape of the real material. Do not paste a partially redacted sensitive document into an unapproved chat to ask whether the redaction was good enough.
Step 3: The Capability
For every Skill, connector, integration, browser, code tool, or action, write:
- source
- reach
- fit
- outside content it reads
- consequential actions it can take
- outcome: enable / escalate / decline.
If you ask an AI assistant to inspect a Skill or plugin package, ask it to identify concerns. Do not ask it to certify the package as safe.
Read this package and tell me what it actually instructs the AI to do, including anything that goes beyond what it says it is for.
Point out any file, network, credential, tool, or external-service access you can see, and explain what would become risky if the session it runs in can reach sensitive data or take consequential actions.
Do not certify it as safe. Finish with one of three lines: obvious concerns found, no obvious concerns found in this review, or not enough information to say.
Step 4: The People
Write:
- who is affected
- what harm or unfairness is plausible
- whether you are reviewing exclusions as well as inclusions
- what disclosure is required by rule
- what disclosure is appropriate by judgment
- what, if anything, needs escalation.
If the fairness question is the one you are least sure about, work it aloud:
Here is a workflow: [describe it, using a fictional or already-approved example].
Take me through four questions one at a time, and wait for my answer before you move on. Who is affected, including people who will never see the output? What could go wrong for them, and would they be able to tell? What would a fair outcome look like? What disclosure is required by a rule, and what would only be my own judgment?
Then tell me whether this is mine to decide and document, or something to escalate. If it should be escalated, draft it as a question with my reasoning attached, not as a request for approval. Do not tell me the workflow is compliant.
Step 5: The Evidence
Add:
- success measure
- failure threshold
- monitoring owner/cadence
- residual risk.
Step 6: Share the record
Send it to the person who owns the workflow, policy, or risk with a simple request:
"This is how I propose to run the workflow. Please correct any assumption about approval, data route, reviewer, or control."
The goal is not to collect signatures for every harmless task. It is to turn non-obvious assumptions into visible decisions before they become incidents.
Here is the whole record as one copyable block. Fill it in, delete nothing, and write OPEN QUESTION wherever you do not yet know the answer.
GOVERNANCE RECORD
Workflow: Owner: Date:
THE CASE
Classification: appropriate / appropriate with review / inappropriate
Deciding factor:
Gate: who what they verify when
THE DATA
Tier: green / yellow / red
Does the task need the identifiers? yes / no
Fields removed:
Approved route:
THE CAPABILITY
Tools, Skills, connectors enabled:
Source: Reach:
Outside content it reads:
Actions it must not take without review:
THE PEOPLE
Who is affected:
Fairness check (including exclusions):
Disclosure decision: Required by / judgment
Open question or escalation:
THE EVIDENCE
Success measure:
Failure threshold:
Monitored by: How often:
Residual risk:
RE-CHECK IF: model, feature, data, audience, permission, policy,
vendor term, or business consequence changes.
14. Quick self-check
-
A customer-facing AI draft is "appropriate with review." What three things are still missing?
Who reviews, what specific risk they verify, and when they review. -
You have a sensitive dataset, but the task only needs aggregate patterns. What question comes before "which privacy setting should I use?"
Whether the task needs the identifiers at all. -
A Skill comes from a trusted source. Is the trust check finished?
No. Check its reach, fit, the outside content it reads, and the actions it can take. -
An AI shortlist looks reasonable. What should you inspect besides the shortlist?
A sample of the excluded cases. -
You accidentally used an unapproved route for sensitive information. What is the first governance objective?
Contain the spread, record the facts, and report promptly through the incident path, not hide or improvise. -
A workflow was approved last year. What kinds of change should trigger a re-check?
Changes to model, feature, permissions, data, audience, policy, contract, or consequence.
15. Plain-English glossary
Appropriate with review. AI may do the task, but a specific human gate must run before the outcome is used.
Defined gate. A review that names who reviews, what they verify, and when in the workflow.
Deciding factor / load-bearing criterion. The factor that is actually carrying the classification. If it changed, the classification would likely change.
Accountability. Who remains answerable for the outcome. Accountability cannot transfer to a tool, so delegating the task does not delegate the answerability.
Accountability transfer. The mistaken belief that responsibility moves to the tool along with the work. It does not.
Partial redaction. The failure where enough identifying detail remains, such as a rare role, an exact date, or a combination of ordinary facts, that a person is still identifiable.
Diligence gap. The distance between what the policy requires and what people actually do. This is where risk lives.
Usage audit. Periodically comparing actual usage against policy, to turn invisible risk into actions you can close.
Data tier. A simple classification of how carefully information must be handled. This course uses green / yellow / red as a teaching model. Use your organisation's actual scheme in practice.
Entry point / route. The specific way data reaches an AI system, such as chat, project, connector, API, or file upload.
Redaction. Removing fields the task does not need before the information reaches the AI system.
Pseudonymisation. Replacing direct identifiers with labels while a way to reconnect the data to people still exists.
Anonymisation. Transforming data so that it meets the relevant standard for no longer being identifiable. The exact test depends on jurisdiction, policy, and context.
Source. Who created or published a Skill, plugin, extension, tool, or integration.
Reach. What the capability could access or affect in the environment where it runs.
Fit. Whether the capability and access are proportionate to the job.
Prompt injection. Malicious or misleading instructions placed inside content an AI system reads, such as a web page, email, or document, in an attempt to steer its behaviour.
Least privilege. Giving a person, service, or agent only the access required for the task.
Residual risk. What can still go wrong after the planned controls work as designed.
Near miss. A governance problem caught before meaningful harm occurred.
Governance Record. The one-page summary of the Case, Data, Capability, People, evidence, owner, and re-check triggers for a workflow.
16. What this course does not replace
This framework does not replace:
- applicable law or sector regulation
- your organisation's privacy, security, acceptable-use, records, HR, legal, procurement, or risk policies
- client contracts, confidentiality terms, or professional obligations
- formal security assessment of software or agent architecture
- structured model evaluation for high-impact systems
- legal advice about whether particular data is personal, anonymous, privileged, regulated, or permitted for a given AI provider.
It gives practitioners a way to recognise which question they are facing and arrive at the right owner with useful reasoning.
To keep the underlying judgment sharp, the capacity to notice when a fluent output is wrong, read How to Think in the AI Era. For the practical checklist on installing and scoping extensions, read Skills & Connectors, with one correction: its guidance on starting read-only and granting narrow scopes describes connectors, which have scopes. Skills do not, as Section 3.1 explains.
If you are collecting credentials, see Certifications for the exams this material prepares you for.
17. Where this leads in The AI Agent Factory
The four questions scale into the rest of the book:
- The Case becomes autonomy design and the human–agent operating model in Human–Agent Teams.
- The Data becomes custody, context, minimisation, and approved routes across agent systems, worked through in General Agents on the Web and Cowork.
- The Capability becomes tool boundaries, scoped credentials, confirmations, typed actions, and audit logs in Designing Agent Experiences and the Mode 2 material.
- The People becomes evaluation, disclosure, fairness checks, monitoring, and recourse, and eventually the governance layer of a System of Record.
- The Governance Record becomes a compact input to design review, testing, launch criteria, and ongoing monitoring.
If you have not yet read AI Fluency, read it alongside this course. AI Fluency helps you decide how to delegate work well. This course helps you decide how to make that delegation defensible and repeatable inside an organisation.
Sources and product notes
The screening approach builds on the AI Fluency Framework by Rick Dakan and Joseph Feller, particularly its ideas around Delegation and Diligence. The three-answer structure, deciding-factor test, route-first data reasoning, and Governance Record are adaptations for this book.
Product-specific behaviour changes faster than governance principles. Claude examples in this course were re-checked on August 31, 2026 against current Anthropic documentation. Before relying on a product feature for a real decision, verify the current source:
- Use incognito chats
- Use Claude's chat search and memory
- What are Projects?
- How can I create and manage Projects?
- Roles and permissions
- Export your organisation's data
- Configure custom data retention controls for Enterprise plans
- Claude Cowork architecture overview
- What are skills?
- Use skills in Claude
- Agent Skills, Claude Platform Docs
- Get started with skill and plugin scanning
The treatment of anonymisation and regulated data is deliberately principle-based. Different jurisdictions and sectors use different legal tests. Remove information the task does not need, then check the remaining dataset and the intended route against the standard that actually governs your organisation.
Flashcards Study Aid
Test Your Understanding
The four questions read easily and run hard. These scenarios drop you into somebody else's decision with the pressure already applied. Twelve come up per sitting, drawn from a larger pool, so a second attempt is not the same paper.
Answer from the reasoning rather than the wording, and for each one, notice which of the four questions you are actually being asked. That identification is half the skill.
The sentence to remember
Classify before you act. If the answer is in the middle, name the commitment. Then write down why.