It may establish almost anything and decide almost nothing. A reviewing role establishes what is: it reads, compares, reports and prepares. Deciding is something else, namely choosing between several defensible options and carrying the consequences. The first can be handed over, the second cannot be handed over completely. In my own business, a split into three levels has proved useful. Alone the role may establish, meaning everything that is verifiably right or wrong. Under a procedure written down beforehand it may handle standard cases where two people working from the same rule would reach the same result. With the human stays what spends money, goes outside the house, touches contracts or promises, or calls for a judgement that has no rule yet. The reason for the boundary is not that the role would judge worse, but that it cannot answer for its judgement.
I run my business with an AI workforce of my own, each employee with their own role, their own personnel file and their own limits. Who does which job is in the org chart, and how it grew is in My AI Workforce. I collect all role topics on the AI employees page. This post answers the third of three questions. Whether you need an overview role is in Does Your Business Need an AI Lead Role?. How to bring several agents together in practice is in Orchestrating Multiple AI Agents. This one is about what the role is then allowed to do.
Establishing is not deciding
Mixing up the two words produces errors in opposite directions: "AI isn't allowed to judge anything anyway" and "if it's reviewing already, it may as well sign off". Both squeeze capability and authority into a single term.
A reviewing role can absolutely judge on the merits: apply criteria, find errors, name contradictions and explain why something does not hold up. That is capability, and it is the entire point of the role. What it does not get is the authority to turn that judgement into a fact.
Three marks tell you that you are dealing with a decision rather than a finding:
- There is more than one defensible option. If only one result can be correct, you are calculating, not deciding. "The page is reachable" is not a decision. "We take the page offline" is.
- The choice creates facts. The money is gone, the email is out, the text is online, the appointment is promised. A finding creates nothing, it describes.
- Someone answers for it. If it goes wrong, there is a name that carries it. A role cannot fill that position, however good its judgement.
A finding has none of the three marks: it is checkable, it is reversible, and if it was wrong that shows up when it is checked. That is why you can hand it over, and why the review works at all. An instance that reviews and signs off in the same move is reviewing its own work. Why author and reviewer are never the same instance in my setup is in Claude Code Subagents: Best Practices.
Three levels you can write down
Level 1 is more generous than most people cut it. Holding files against a report, searching the project for a sentence, putting two numbers side by side and naming the contradiction, writing a draft: all of it without asking, because none of it creates a fact. The judgement on the merits belongs here too, as long as it is phrased as a finding and not as an instruction.
Level 1 still has an edge, and it is not the write access. A finding that someone builds on unchecked works like a decision: a security assessment, a data protection classification, a statement about creditworthiness or health. So do not only ask whether a task changes a state, but also how far its answer carries and who would notice that it is wrong.
Level 2 is where it goes wrong if you are sloppy. The test is in the table. For "it may put questions to a colleague in the internal mailbox on its own", yes, because nothing leaves the house in the process. For "it may handle the smaller things itself", no, because "smaller" is the judgement you just handed over. My house rule, not a law: what I cannot write down in five sentences is not a level 2 in my setup.
Level 3 stays with you, along four edges:
- Money. Purchase, subscription, budget increase, credit. In my setup without exception, because no level 2 rule makes an invoice disappear again.
- Outside effect. Anything that leaves the house: email, quote, publication, ad, post, and the seemingly harmless question to a client, applicant or supplier as well. Once it is out, it is out.
- Contracts and promises. The small ones too: an appointment given, a deadline named, a price confirmed.
- The one-off judgement. A case with no rule, because it has not come up in this form before. You recognise it by not being able to write the rule without having decided the case first.
The three levels are not a scale from little to much but a ranking: what sits at level 3 stays there, even when a level 2 rule seems to cover it. The same logic sits in the technical side. Claude Code knows exactly these three kinds of rules for tool permissions and, per the documentation, evaluates them in the order deny, ask, allow, first match decides, and a narrower allow rule carves no exception out of a deny (retrieved 20 September 2026, the state of the version current at the time). The prohibition wins against the permission, not the more precise wording.
How to write this down for your business
This needs no software and no rulebook with clauses, only one line per task in the file the role reads its brief from anyway:
<task> | level 1 | 2 | 3 | rule or evidence | who signs off
- List tasks, not roles. Write down what gets done, not who does it. The assignment comes afterwards and changes nothing about the level.
- Put every task on a level, and when in doubt the higher one. Moving from 3 to 2 costs five minutes once you have the rule. The other direction costs you the case.
- Write the procedure next to every level 2 task. Five sentences. If that does not work, it was a level 3.
- Write down how the sign-off is recognised for every level 3 task. In my setup "get it done" is not a sign-off, only a clear "send it" or "put it online", noted with wording and time. Without that mark, level 3 quietly becomes level 2 in daily use.
- Put the list where the role reads it. In the role file, not in the chat window: an agent has no memory across sessions, your file system does. What that role context looks like is in CLAUDE.md: Claude Code's Memory.
- Also enter what your tool enforces technically. Paper describes the intent, the permission rule carries it out.
- Read the list back after four weeks. Which level 3 task did you decide the same way four times? That one is ready for a rule. Which level 2 rule surprised you? Back to 3.
Point 7 keeps the list alive, otherwise it describes the business of back then within a quarter. Cutting the levels is the actual work, and nobody has to do it alone: if you are unsure, put your list next to the others in the community, because the borderline cases repeat from business to business.
When the decision is about people
One area has rules of its own, regardless of how you cut your levels. Article 22(1) of the General Data Protection Regulation gives every data subject the right "not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning him or her or similarly significantly affects him or her". Recital 71 names as examples the automatic refusal of an online credit application and e-recruiting practices without any human intervention.
That sounds sharper than it is for daily work. Both conditions have to hold at once, solely automated and with legal or similarly significant effect, and it is about decisions about people, not about whether your agent may rename a file.
Article 22(2) names exceptions, for instance where the decision is necessary for a contract or is based on explicit consent. That is not a free pass. For exactly those two cases, points (a) and (c), paragraph 3 requires suitable measures, and it names them: the right to obtain human intervention, to express one's point of view and to contest the decision. Paragraph 4 further restricts special categories of data under Article 9.
Alongside the GDPR there is the AI Act: Annex III classifies employment and workforce management (point 4) and the creditworthiness assessment of natural persons (point 5(b)) as high risk, with obligations of their own attached.
The word where the three levels and the regulation meet is "solely". Anyone who reads the role's proposal, checks it and then decides themselves is not making a solely automated decision. Anyone who turns the sign-off into a formality and clicks without looking is betting that their tick counts as human intervention. I would not take that bet, and that is my assessment, not a reading backed by a court ruling. None of this is legal advice in any case: if applications, creditworthiness, terminations or prices for individual people are in play in your business, have it checked by someone who is liable for the answer.
What this looks like in my setup
The position that keeps the overview in my business is called Kai. He reads what the others leave behind anyway: git history, the reports and minutes in the shared folder, the mailbox. After a session he holds an employee's claims against what is in the file system and on the live website, and writes a short report from it with the things I have to decide. That is level 1, entirely. At level 2 there is exactly one thing: he may send questions and notes to an employee through the internal mailbox on his own. Everything else is level 3. He decides nothing on the merits, signs off nothing, commits nothing and changes no other employee's work. Briefs on the merits he issues only on my express instruction.
That this is not a question of his capability is shown by his review of 5 September 2026. He checked the work of two employees and found three things: an incomplete correction run after which a website contradicted itself in five places; one employee sitting permanently on a permission level without prompts; and a changed knowledge file with no record of who changed it. Conversely the same review confirmed a completed rollout across six core routes checked live.
Those are judgements on the merits, and they are better than what I would have seen skimming. Every one of them still ends as a finding, not as an action. He may write down that the site contradicts itself in five places. He may not decide that it therefore goes offline. Where evidence is missing he writes "not verified" instead of smoothing over the gap: three times in this one review. His tools and limits are on Kai, who keeps the overview.
Kai assumes several roles are already running. Anyone heading there starts lower down: Learn Claude: The Path in Five Stages is the way in, and the process is in Hiring an AI Employee: The Process.
Frequently asked questions
Can an AI agent decide anything itself?
Within a procedure written down beforehand, yes, otherwise no. If two people working from the same rule would reach the same result, it is no longer deciding but applying. As soon as the rule contains a judgement you have not written down, you have handed over the judgement rather than the task.
What is the difference between reviewing and deciding?
Reviewing establishes what is, and the result is verifiably right or wrong. Deciding chooses between several defensible options and creates facts in doing so. A reviewing role may judge on the merits and substantiate errors, only the final sign-off is not hers.
Can AI decide on job applications or credit requests?
Article 22 GDPR covers decisions based solely on automated processing with legal or similarly significant effect on a person, and Recital 71 names exactly those two cases as examples. The AI Act additionally lists both in Annex III as high risk. This is not legal advice: if such cases occur in your business, have the arrangement checked professionally.
Do I have to write a rulebook for this?
No, one line per task is enough: task, level, rule or evidence, who signs off. What matters is not the length but the location: the file the role reads its brief from, not a conversation that ends with the session.
What do I do if the role decided something anyway?
You find it in the evidence, that is what the audit trail is for. What follows is not a reprimand but a line: the task moves up a level, and if your tool knows permission rules, it gets blocked there too. Prohibitions win against permissions, in the text as in the configuration.
What to do next
Take the last three things your AI role did and write one line for each: task, level, and for level 3 the name of whoever signs off. If one gives you pause, you have found your first real borderline case. It belongs at level 3 until you can write the procedure in five sentences. That is all it takes to turn "the AI just does that" into an authority you can account for.
The way there, from the first session to your own workforce with clear limits, is what we go through step by step in my community, with the finished role packages as a starting point.