A feature in your mail program is enough when your sorting follows general patterns and the worst case is that you see a message a few hours later. You need a role of its own as soon as the decision requires knowledge about your business that isn't in any message: which client is escalating right now, which deadline is running, which sender is a supplier and which is a prospect. Four criteria separate the two cases: context, rate of change, liability, data. Vendor features sort by patterns from the message and the thread; a role sorts by rules you wrote down and can change at any time. In the default setup neither Copilot nor Gemini nor Apple Mail nor an AI employee sends a reply without you; sorting is a different matter, because moving, archiving, categorizing and deleting are state changes a tool carries out on command. In my setup replies stay drafts; sending and deleting need my approval. Everything below comes from the vendors' own help pages, as of September 2026.
I made this decision for my own office and ended up with the role, because my rules change too often to maintain them in settings dialogs. What an agent actually is is explained on the AI agents hub, the tool comparison between Microsoft's assistant and Claude in Copilot vs Claude: What Works at the Office. This piece is about the inbox only, and about the tool question only; which piles and rules you need in the first place is in Sorting Your Inbox: The Pile System. Without that procedure there is nothing to automate.
What mail programs actually do in September 2026
First the state of play: everything in this table is what the vendors' own help pages say.
Two things stand out. First: what the vendors offer is mostly summarizing and drafting, not deciding; the one rating feature with criteria of your own is Copilot's Prioritize, and it works on three levels. Second: if you work in another language, check first whether a feature is switched on for your language and region instead of inferring it from an announcement.
The four criteria
The choice between feature and role comes down to four questions.
No single criterion decides anything. My rule of thumb: when three out of four land in the same column, the answer is clear enough and you can skip the feature-list comparison.
Criterion 1: How much context the decision needs
This is the real dividing line. A feature sees the message, the thread and your address book; Microsoft names exactly those signals for Prioritize: the people on the thread, their job titles, the content. That reliably spots a message that wants an answer. Not visible from the text: that this sender has been waiting three weeks for a quote, that you agreed something different with another one verbally, or that the invoice attached belongs to a project that hasn't been signed off.
The test is simple. Take the last ten messages you handled wrongly or too late and ask, for each, whether an attentive stranger would have reached the right rating from the text alone. My rule of thumb, not a proven threshold: if they get eight out of ten right, a feature is enough. If half of them need background knowledge, you are sorting by context, and context has to live somewhere.
That is the difference between a settings dialog and a file. Microsoft lets you state criteria as text and recommends whole phrases over single words: more than a rigid filter, but it ends where the dialog ends. A role gets a file instead, with your rules, your exceptions and your reasons, and that file may be as long as your business is complicated.
Criterion 2: How often your rules change
Fixed rules age badly. A filter that catches "invoice" in the subject is right until a client starts writing "statement"; a criterion in Copilot's settings until your project name changes. What costs you isn't the ageing, it's the upkeep. With a feature you maintain rules where the vendor put them: per program, per device, per account. Use Outlook on the desktop and Apple Mail on the phone and you maintain two systems with two vocabularies. With a role you maintain one text file, in the same language you think the rule in.
The practical difference shows up in the doubtful case. A settings dialog has no place for "I don't know", it has high, normal, low. A written rule, by contrast, is allowed to say: if this and that apply, put it in front of me and don't decide yourself. That fifth pile is why written rules hold up in daily use and filter lists don't; how it works is in the pile system.
Criterion 3: Who answers for a message filed in the wrong pile
The short answer: you do. A message doesn't count as undelivered because a piece of software put it in the wrong pile. Under section 130 (1) of the German Civil Code a declaration of intent made to an absent party takes effect at the moment "it reaches them" (statute); what matters is the recipient's sphere of control and the possibility of taking notice under ordinary circumstances. For business dealings the German Federal Court of Justice has held that an email is generally received once it is available on the recipient's mail server during business hours (BGH, 6 October 2022, VII ZR 895/21). So when exactly a message is received depends on the circumstances, retrieval and business hours included. Wrong automatic sorting usually doesn't shift it, because the message sits in your sphere of control either way. That is not legal advice and doesn't replace it in an individual case, it is the reason liability belongs in the tool decision at all.
The vendors see it the same way. Microsoft writes in its privacy and security documentation that the responses generative AI produces "aren't guaranteed to be 100% factual" and that users "should still use their judgment", and describes its own features explicitly as drafts and summaries "rather than fully automating these tasks" (Microsoft Learn). Which means, for the setup: sorting may run automatically, deleting and sending may not.
So check what happens at worst when a message sits in the wrong pile for a week: for a newsletter nothing, for a dunning letter or a notice period quite a lot. The more often the second case comes up, the less you want a rating you can't read back. Written rules have the advantage here, because every decision traces back to a line you wrote yourself.
Criterion 4: What happens to the data
The differences here are smaller than many expect, and the footnotes matter more. Microsoft writes that prompts, responses and data accessed through Microsoft Graph aren't used to train the language models, and names the EU Data Boundary for EU users; models provided by Anthropic as a subprocessor are, per the same documentation, currently excluded from it. Google puts it similarly for Gemini in Workspace: no training on customer data without prior permission, no human review outside your own domain, interactions stay within your organization (Google Workspace Privacy Hub).
The more practical question is: how much gets processed, and where? A feature that summarizes every incoming message processes the full text of each one, including those in the job applications folder and the one with a medical note attached. Where that text goes differs from product to product: Apple writes that its models run entirely on device "in many cases" and that more complex requests go to Private Cloud Compute, where the data is, per Apple, neither stored nor made accessible to Apple (Apple); which path applies to a single mail feature isn't stated there. So compare documented data flows, not labels: "role" doesn't by itself mean less gets read.
What a role has going for it: you draw that line yourself and write it down. Frieda is limited exactly that way in my setup: for sorting she sees only sender, subject and preview, the full text only where a reply is going to be written. That staged reading isn't a property of the category, it's a line in her file. Where your line sits is nothing you have to settle alone: the calls in my community are full of people who have made the same trade-off, and a question in a post beats two evenings of pondering.
And where do n8n and the workflow builders fit?
They are the third answer, and they pay off where triage doesn't end at the mailbox: when a message should reliably trigger a ticket, a row in a sheet or a notification in another system. If the real work is judgment, you end up building a node that asks a language model anyway, and your rule now lives in a flowchart instead of a text. Which tasks belong in a builder is worked through on a concrete example in n8n, Make, or Claude: What For?.
What the role looks like day to day
This is how Frieda works for me. She sorts the inbox into four piles plus a fifth for the cases my rules don't clearly cover, and puts finished drafts in front of me for the messages that need an answer today. Nothing is sent unread, nothing is deleted at all. Sensitive areas such as private matters, job applications or health she doesn't touch, not even to sort them; receipts are a transit stop rather than an archive, and she doesn't answer tax questions.
The difference from a feature isn't the technology, it's the source of the decision: Frieda sorts by the rules in her file, and that file is mine. What such a file looks like word for word is in CLAUDE.md: Claude Code's Memory, including tasks, limits and approvals. It's a set of working instructions, the kind you would give a new hire, only written down instead of passed on verbally.
Frequently asked questions
What does email triage mean in the first place?
Triage means sorting incoming messages by urgency and responsibility before working on them, instead of processing them in order. The term comes from emergency medicine and means the same in a mailbox: classify first, then act. The classifying can be handed off, the handling usually stays with you.
Can Copilot sort my inbox automatically?
Copilot assigns a priority on arrival per Microsoft and can, through its triage feature, pin, flag, move and categorize messages and create rules. Not rated: messages outside the inbox, meeting invitations, encrypted mail and anything that arrived before the feature was switched on. As of September 2026.
Do I need AI at all, or do filter rules do the job?
For anything with a hard feature, filter rules do the job, and they are faster, cheaper and easier to follow; Thunderbird, Outlook and Gmail have done this for years. AI only pays off where the feature is soft: where only the content reveals whether a message is urgent. Start with rules and add AI where rules fail.
What is the difference between an agent and a feature in my mail program?
A feature works within the criteria and actions the vendor provides, in the vocabulary of its settings dialog. An agent works by rules you write in your own words, can treat doubtful cases as doubtful cases, and can be extended in the same file that holds its limits. The price is that you have to maintain that file.
May an AI answer messages in my name?
Technically possible, rarely sensible. Microsoft describes its own features explicitly as drafts and summaries, not as full automation. Do the same: drafts yes, sending no. Responsibility for the content of a sent message stays with you, regardless of who phrased it.
How to go on from here
Run the ten-message test from criterion 1 before you switch anything on: note, for each case, whether the right rating was in the text of the message or only in your head. If it leans toward the text, switch on your mail program's feature and save yourself everything else. If it leans toward your head, write that knowledge down before you pick a tool: it's the part no vendor ships. If you get stuck, put the question in a post.
If you take the second route: Frieda is available as a finished package with exactly this structure in my community Claude Practitioners, together with the course for it and the other AI employees. The rule file is already written there, and you only fill in your own context.