AI Employee Approval Matrix: What It Can Do Without You
An AI employee approval matrix should let routine, reversible, policy-covered work run automatically. It should pause when an action crosses a clear boundary involving money, legal or clinical judgment, an unusual customer promise, sensitive access, or a change that is hard to undo.
If your AI employee asks you to approve 40 normal messages before lunch, the software did not remove the work. It changed the shape of your inbox.
Two failure modes are requiring approval for every message and record update, which creates another inbox to manage, and telling the system to "use good judgment" without defining edge-case boundaries.
The better approach is to define authority one action at a time. Name what the AI employee can run, what it can prepare, what requires a decision, and what it cannot do. Then attach a finish condition and proof to each action.
The matrix below is a planning starter, not a legal or compliance standard. Adjust it to your policies, systems, customers, and regulated obligations.
Start with this four-state approval matrix
Use these four states for the starter matrix.
Run: complete the action automatically within written policy.Prepare: gather facts and build the finished draft or recommendation, but do not execute it.Ask: pause only the affected action and send the authorized person a decision-ready brief.Blocked: the AI employee cannot perform the action. Route it to the qualified role.
Here is a populated starter for common small-business work:
| Business action | Default state | Move to Ask when | Completion proof |
|---|---|---|---|
| Answer a routine service question from approved information | Run | The answer is missing, conflicting, or would create a new promise | Message sent and conversation logged |
| Ask a lead for missing contact or job details | Run | The person objects, asks for an exception, or shares sensitive information outside the workflow | Required fields captured or request closed |
| Schedule within approved hours, service area, and availability | Run | The request needs an override, special accommodation, or unverified capacity | Calendar and CRM agree; confirmation sent |
| Send a policy-approved reminder or status update | Run | The record shows an opt-out, dispute, complaint, or conflicting status | Current state rechecked; message and outcome logged |
| Correct a duplicate tag or formatting error in a CRM | Run | The change would merge people, delete history, or alter a financial or legal record | Before-and-after values recorded |
| Draft a discount, credit, refund, or payment-plan option | Prepare | Always, unless a written policy delegates a specific amount and scenario | Decision packet contains facts, options, and impact |
| Issue a refund or credit within a written amount-and-scenario delegation | Run | The amount, reason, customer, or transaction falls outside that delegation | Approval identity when required, amount, transaction result, and customer notice logged |
| Publish a new public claim, offer, price, or policy | Ask | Always for new material; approved routine updates may have a narrower rule | Approved version and live destination match |
| Change tool permissions, credentials, automations, or production settings | Ask | Always unless the change is a tested, pre-approved maintenance action | Change log, test result, and rollback state saved |
| Delete customer records or bulk-modify data | Ask | Always; some records may be subject to retention or other requirements | Authorized scope, result count, and exception log saved |
| Make legal, clinical, tax, lending, or licensed professional judgment | Blocked | Route to the qualified professional | Decision source and downstream administrative steps recorded |
| Override company policy or make a novel customer commitment | Blocked | Route to the person who owns that policy or commercial authority | Authorized decision received, communicated, and completed |
Use each row to define a default, boundary, and proof for one exact workflow action.
A plumbing company may let an AI employee qualify a service-area request and offer approved appointment windows, while a restaurant may let one answer standard event-capacity questions and gather details. A law firm or med spa will draw different lines around professional judgment and sensitive information. The law-firm responsibility matrix shows why administrative ownership and licensed judgment must stay separate.
Approval should follow the action, not the tool
Owners often grant access by application: "It can use the CRM, but not accounting." That is too broad in one system and too narrow in another.
A CRM can contain harmless tags, sensitive contact details, financial notes, consent records, and history that should not be deleted. Accounting software can support a read-only cash brief without allowing a payment. Email can be used to send an approved appointment confirmation or to make a novel promise that creates real liability.
Define authority as verb + object + conditions:
- update the lead-status field after verified customer action;
- send the approved missed-call reply to an eligible contact;
- draft a refund recommendation using current transaction facts;
- read invoice status but do not change the balance;
- publish only a version approved for that exact destination.
NIST's AI Risk Management Framework is voluntary guidance built around managing AI risk in context. That is the useful lesson for an owner: there is no universal list of safe actions. The same tool call can be routine in one workflow and material in another.
NIST's security and privacy control catalog also covers access control, audit and accountability, authorization, and monitoring. You do not need to turn a small-business rollout into a federal control program. You do need the operating principles: minimum necessary access, evidence of what changed, and a way to monitor whether the system stayed inside its assignment.
Do not make customer communication one giant approval category
"All customer messages need approval" sounds safe. It also prevents the AI employee from owning routine follow-up.
Separate messages by purpose and authority:
- Routine factual messages can run when they use approved information and the source record is current. Examples include appointment confirmations, requests for missing details, and status updates.
- Policy-sensitive messages can run only inside a defined scenario, approved template family, and verified eligibility rule.
- Novel promises, disputed facts, complaints, concessions, pricing changes, and professional advice move to
AskorBlocked.
This is how an AI employee can own the daily work without speaking beyond the business's authority. Our restaurant review-response SOP uses a similar separation: routine replies follow approved response paths, while privacy issues, threats, allegations, and unusual recovery decisions become exceptions.
Before each outbound action, the AI employee should recheck the source of truth. A lead may have already booked. An invoice may have been paid. A customer may have opted out. An appointment may have moved. Approval of a template last month does not excuse sending it against a stale record today.
Build the boundary around reversibility and impact
Two questions reveal where authority should change.
First: if the action is wrong, can the business reverse it cleanly?
Correcting a lead-stage tag is usually easier to reverse than deleting a record, issuing a refund, publishing a price, or changing production permissions.
Second: who bears the impact?
An internal summary may affect only the person reading it. A customer message can create an expectation. A payment action affects money. A public claim affects the brand. A clinical or legal conclusion can affect someone's rights, health, or case.
Use these factors when assigning a state:
- source reliability;
- customer visibility;
- financial impact;
- sensitivity of the data;
- reversibility;
- novelty;
- professional or regulatory judgment;
- permission level required;
- number of records affected;
- ability to prove the final state.
Do not turn those factors into a fake scientific score. They are discovery prompts. The owner, workflow lead, and qualified advisers still decide the rule.
A true exception should arrive ready for a decision
An alert that says needs review sends the work back to the owner. So does a forwarded email, raw transcript, or list of links.
When an action moves to Ask, the AI employee should finish all unaffected work and prepare the smallest useful decision packet:
Exception brief
- Workflow and record:
- Proposed action:
- Why the normal rule did not apply:
- Facts verified and source timestamps:
- Customer, money, data, or policy impact:
- Options allowed by current policy:
- Recommended option and reason:
- Exact decision needed:
- Decision deadline:
- What remains paused:
- Work that continued:
- What the AI employee will do after the decision:
The last three lines matter. One exception should not freeze the whole workflow.
Suppose a customer accepts an invoice balance but asks for a payment schedule outside policy. The AI employee can verify the account, stop the normal reminder sequence, prepare the request, and keep unrelated invoices moving. After an authorized person chooses an option, the AI employee sends the response, updates the account, sets the next due action, and confirms that the record matches the decision. The plumbing invoice follow-up workflow shows this distinction between routine collection work and true disputes or policy exceptions.
Human approval is a pause in one action. It is not the end of ownership.
Route decisions before the first exception arrives
An Ask rule is incomplete until the AI employee knows who can decide, who covers when that person is unavailable, when the answer is needed, and what happens if nobody responds.
Use this routing starter for the exception families in the matrix:
| Exception family | Decision owner | Backup owner | If nobody responds |
|---|---|---|---|
| Refund, credit, discount, or payment-plan exception | Person who owns the financial policy | Named finance or operating backup | Pause that transaction, send one timed reminder, and keep unrelated accounts moving |
| New public claim, offer, price, or policy | Owner of that public commitment | Named marketing or executive backup | Keep the item in draft; do not publish by timeout |
| Permission, credential, automation, or production change | System owner | Named technical or operating backup | Leave the current configuration unchanged and report the blocked work |
| Record deletion or bulk data change | Data or system owner | Named operating backup | Preserve the records and stop only the proposed bulk action |
| Legal, clinical, tax, lending, or other licensed judgment | Qualified professional for that matter | Another qualified person named by the business | Hold the affected decision and continue administrative work that does not depend on it |
| Policy override or novel customer promise | Person who owns that policy or commercial authority | Named executive backup | Make no new promise; acknowledge receipt if policy permits and wait for authority |
Set a real response deadline for each workflow. "ASAP" is not a routing rule. A same-day refund request, a customer waiting on a schedule exception, and a nonurgent system change do not need the same window.
Copy this authority record for every Ask or Blocked row:
Action | Primary decision owner | Backup owner | Response deadline | Reminder time | If no response | Work that continues | Post-decision finish steps
The safe default for no response is to preserve the current state, not silently approve the action. The AI employee should send the planned reminder or escalation, explain the operational impact, and continue work outside the blocked decision.
Apply least privilege without crippling the job
The AI employee needs enough access to finish the assigned outcome, but no more.
Start with four permission questions:
- Which systems must it read?
- Which exact fields or records must it change?
- Which communications or transactions can it initiate?
- Which sensitive actions need an independent check or approval token?
Read access is not automatically harmless. A broad inbox, drive, or CRM can expose information unrelated to the job. Write access is not automatically dangerous. Updating one approved status field may be the routine finish step.
Limit scope by account, record type, field, action, amount, destination, and time window when the system supports it. Keep credentials out of prompts and plain-text notes. Record the action parameters and result. For sensitive execution, the approval should apply to the exact action the person reviewed, not to a vague category the AI employee can reinterpret later.
OWASP's Agentic AI - Threats and Mitigations presents a threat-model-based reference for emerging agentic threats and mitigations. A prompt that says "be careful" is not an access-control system. Tool permissions, validation, logs, and execution checks must enforce the boundary even when the model receives bad or misleading input.
Build the first matrix in 30 minutes
Do this before a longer rollout:
- Pick one recurring workflow with a visible finish condition.
- List the actions the AI employee must take instead of listing tools.
- Assign
Run,Prepare,Ask, orBlockedto each action. - Add the condition that changes the state and the proof that closes the action.
- Name the primary and backup decision owner for every exception family.
- Set the response deadline, reminder, and no-response behavior.
- Put the workflow into shadow mode so the AI employee proposes and routes work without executing live actions.
Thirty minutes will not settle every policy question. It will expose the missing ones before they turn into vague prompts or surprise approval requests.
Test the matrix in shadow mode for seven days
Do not begin by connecting every system and watching what happens. Test one recurring workflow where the finish condition is visible.
Anthropic's Building effective agents recommends simple designs, ground truth from the environment, checkpoints or human feedback when needed, and extensive sandboxed testing with guardrails. That translates into a practical seven-day rollout:
Day 1: map the real work
Pull a small sample of recent cases. Reconstruct each trigger, decision, action, handoff, and final state. Note where policy was clear and where staff improvised.
Day 2: assign authority by action
Use the four states for every action in the workflow. Name the exact condition that changes the state. Avoid broad labels such as "customer work" or "financial task."
Day 3: define finish conditions
For every Run action, state what must be true before the item is closed. A sent message is not enough if the CRM remains stale or a promised follow-up has no due time.
Day 4: run in shadow mode
Let the AI employee read the inputs and propose actions without executing them. Compare its route, facts, and intended changes with what a trusted employee would do.
Day 5: test ugly cases
Use stale records, duplicate events, missing fields, contradictory notes, opt-outs, complaints, policy exceptions, bulk changes, and unavailable tools. Confirm that one exception does not stop unrelated work.
Day 6: release the narrowest routine lane
Turn on a small set of Run actions with clear rollback and complete logging. Keep Ask and Blocked boundaries enforced outside the model where possible.
Day 7: review evidence, not confidence
Check source timestamps, action parameters, before-and-after values, customer-visible outputs, system agreement, unresolved items, and exceptions. Expand authority only when the evidence supports it.
OpenAI's practical guide to building agents describes layered guardrails and human intervention for high-risk actions and failure thresholds. This supports the same operating rule: one prompt should never carry the entire safety burden.
Measure finished work and owner interruptions
Do not judge the matrix by how many approvals it creates. A long approval queue can look controlled while consuming the owner's day.
Track:
- eligible items received;
- routine items completed without staff work;
- items correctly routed to
Prepare,Ask, orBlocked; - false approvals, where the owner had to approve a normal action;
- missed escalations;
- items reopened because the final state was wrong;
- owner minutes spent per real exception;
- time from authorized decision to completed downstream action;
- records with missing completion proof;
- permission or tool failures;
- recurring exception types that need a clearer policy.
Use your own baseline. Review the first week case by case, then summarize by exception type. If the owner is approving the same routine action repeatedly, write the policy and move that narrow action to Run. If one rule creates mistakes or ambiguity, narrow it.
Common approval-matrix mistakes
Approving every message
This keeps the owner responsible for routine communication. Approve the purpose, facts, eligible scenarios, template range, and exception triggers instead.
Letting the model classify its own authority without enforcement
The model can help route actions, but tool scopes and execution checks should prevent actions outside the assigned state.
Using a universal dollar threshold
A number copied from another business has no authority in yours. Set thresholds by transaction type, reversibility, policy, cash exposure, and who owns the decision.
Treating Prepare as finished work
A draft is useful only if it removes research and assembly from the decision maker. After a decision, the AI employee should execute the approved next step and close the loop.
Logging the request but not the result
Save what the AI employee proposed, who decided, what changed, whether execution succeeded, and the final customer and system state.
Giving broad access because setup is easier
Convenient credentials create unnecessary exposure. Start with the smallest access that can still finish the job.
Freezing the whole workflow on one exception
Pause the affected action. Continue unrelated work. Show the owner exactly what stopped and what kept moving.
Turn one recurring workflow into a written authority model
A good AI employee approval matrix does not force the owner to supervise software all day. It moves routine work off the owner's desk while preserving the decisions that truly require authority.
Choose one workflow. List the specific actions. Assign Run, Prepare, Ask, or Blocked. Add the boundary and completion proof. Then test the matrix against real cases before widening access.
Request a free business audit and we will map one recurring workflow from trigger to finished result, including routine authority, true exceptions, system permissions, and the evidence the owner should see. You can also review the AI employee service paths by industry to see where this model applies in your business.