If you already have a transcript, you don't need another tool, you need a better instruction. You hand the text to ChatGPT, Claude or Microsoft Copilot and you do not ask for a summary, because that single word is what turns minutes into a retelling. The instruction that works asks for four separate blocks: a header with date and participants, decisions in the present tense, tasks with an owner and a deadline, and open items. Add three rules that make the difference between useful and dangerous: use only what is in the transcript, give the timestamp of the supporting passage for every entry, and put anything uncertain into a fifth block called "unclear" instead of guessing. Then comes a second pass in which the model has to quote, word for word, the sentence behind every decision and every task, and delete anything it cannot support. Supported is not the same as correctly filed, though, so you look at ownership, deadlines and the question of decision or idea yourself. And before the text goes anywhere: names, companies, amounts and private details that don't belong in minutes, you take out first.
I stopped typing a fresh prompt for jobs like this and use a saved instruction. That is the whole trick: the difference between useless and useful output isn't the model, it's five lines of text. How I build prompts is in Claude prompts that work in daily practice, and which tasks are suited to being handed over at all is covered in Which tasks you can hand to AI.
Why "summarise this" gives you the wrong thing
A summary condenses the course of the conversation, minutes condense the outcomes. Ask for the first and you get the first: not a model failure but a correctly executed instruction. In my own practice the three variants produce three different kinds of text.
The third one works because it gives the model a filing system. Without it the model decides for itself what mattered, and that decision is then not yours.
Handing over the transcript: file or pasted text
Both work, length decides: short calls you paste, long ones you attach. Claude's chat interface accepts up to 20 files per conversation at up to 500 MB each (help article, as of 23 July 2026). With a pure text transcript it is usually the context window that gets tight first, meaning how much text the model can hold in view at once, not the file limit. Both limits remain, though: scans, images or large PDFs reach the file limit too. If the second half of your minutes comes out noticeably thinner than the first, split the transcript by time segments and merge the partial minutes. Long inputs belong at the top, above the instruction: Anthropic explicitly recommends placing lengthy documents above your query and instructions.
Two things not to throw away first. The speaker labels, because without them no task can be assigned to a person and the minutes end up saying "someone will take care of it". And the timestamps, because they are the supporting passage you check every entry against later. Both come out of automatic recognition and can be wrong, in the speaker attribution as much as in the time.
If your transcript comes out of Microsoft Teams, fetch it there as a file first. Where the feature sits, what your administrator has to enable and how long Teams keeps the files is in Transcribing calls in Microsoft Teams. What Teams, Zoom and Meet produce on their own is compared in Creating meeting minutes automatically.
What you take out beforehand
Putting a transcript into a general language model means handing conversation content to a provider. That is no reason to avoid it, but it is a reason to spend two minutes first. I replace names with the role, company names with the industry, amounts with the order of magnitude and places with the region. Passages with no bearing on the matter, so health, family, private finances, I delete rather than replace. The list of which role belongs to which person stays local with me. The minutes lose nothing by it: "managing director" carries a task line as well as a surname, as long as only one was in the call. That does not make the text anonymous: as long as the list exists and role, industry and amount range can together point to a person, it stays pseudonymised and therefore personal data (Art. 4(5) GDPR). Replacing lowers the risk, it does not take the obligations off you.
What the provider does with the text afterwards depends less on the logo than on the plan:
- Claude: Anthropic states for the consumer plans that conversations are used to improve the model only with your consent and that incognito chats are excluded even when model improvement is switched on (help article dated 16 March 2026).
- Microsoft Copilot in a business setting: "Prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs", and EU users additionally get the EU Data Boundary. In the same article Microsoft notes that Anthropic models used as a subprocessor are currently excluded from it (article dated 9 July 2026).
- OpenAI API: according to the provider's documentation, data sent to the API is not used for training, abuse monitoring logs are kept for up to 30 days by default, and eligible customers can get zero data retention. That applies to the API, not automatically to the same company's chat interface.
In practice: look in your account's privacy settings to see whether model improvement is on or off, before the first client call goes in. And then there is the step before that, which no setting fixes: recording the non-publicly spoken word without authorisation is a criminal offence under § 201 of the German Criminal Code, not an administrative fine, and what authorises the recording is prior consent from everyone else involved. What a workable consent looks like and what the GDPR demands for the later processing is in Transcribing conversations: the legal basics. I'm a developer, not a lawyer, this is a working basis and not legal advice.
The instruction you can copy
This is the template I have saved. It works word for word in ChatGPT, Claude and Copilot, because there is nothing product-specific in it.
You are the minute taker. Below is the transcript of a meeting.
<transcript>
[paste transcript here, with speaker labels and timestamps]
</transcript>
Write minutes of record from it, not a summary.
Follow this structure exactly:
1. Header: date, occasion, participants, duration. Only what is in
the transcript.
2. Decisions: one sentence each, present tense, describing what applies
from now on. In brackets behind it, the timestamp of the passage where
it was decided.
3. Tasks: a table with the columns Task, Owner, Deadline, Timestamp. With
no named person the row does not go in this table but under point 4.
4. Open items: everything that was not decided, with the question left
open and, if stated, when and by whom it will be decided.
5. Unclear: anything where you are not sure who something belongs to,
what was meant, or whether it was a decision. Put it here rather
than guessing.
Rules:
- Use only what is in the transcript. Do not use general knowledge and do
not add anything that merely sounds plausible.
- Invent no names, no numbers, no dates. If the transcript says "next
week", write "next week", not a date.
- Do not evaluate and do not recommend.
- If a section stays empty, write "none".
- Give the timestamp for every entry under 2, 3 and 4. If you find no
supporting passage, the entry belongs under point 5.
Two details in there carry more than they look like. The "unclear" block gives the model a permitted way out. Anthropic lists "Allow Claude to say 'I don't know'" as the first basic technique in its guide on reducing hallucinations and says it can drastically reduce false information. The timestamp requirement forces a supporting passage per statement, which the same guide recommends under "Ground responses in quotes". One phrasing is deliberately missing: any request for nice language. Minutes are not supposed to sound good, they are supposed to be right.
The second pass against invented content
The finished minutes are a draft, not a result. Microsoft writes about its own Copilot output that responses produced by generative AI aren't guaranteed to be 100% factual and that users should apply their own judgement before sending them on. So I append a second prompt, in the same chat:
Now check your own minutes against the transcript.
For every decision and every task, quote word for word the sentence from
the transcript you are relying on.
If you find no literal supporting quote, delete the entry and write in its
place: [deleted, no supporting quote].
At the end, list the entries whose quote only partly supports them.
That is the technique Anthropic describes as "Verify with citations": have every claim backed by a quote, and retracted when no quote can be found. It checks the wording, though, not the filing: a sentence can be quoted word for word and still be attributed to the wrong person or put in the wrong category. After that, three places remain that you look at yourself, because no model can derive them from the text:
- Ownership. "Can you do that?" produces no name. If a name appears in the minutes that was never spoken in the call, it was guessed.
- Deadlines. "Next week" is not a date. A model that turns it into the 27th has calculated rather than read.
- Decision or idea. Whether "we should look into that sometime" was an instruction is decided by the context in the room, not by the wording.
When one set of minutes becomes fifty
So far this is about one conversation, and for most meetings that is the end of it. If you run client calls, something else begins: after fifty of them you have fifty correct sets of minutes and still no answer to which question your clients keep asking and which of your own answers worked best.
That is exactly what Gustav, my AI employee for evaluating call transcripts is for. His first working step is not an analysis but the same replacement pass as above, hard-wired rather than optional, with the counter-check of whether role, industry and a rare piece of content together point back to a person after all. If an entry passes it, it is anonymous within the meaning of recital 26 GDPR, if not it stays pseudonymised and gets a purpose, a deadline and access control. Speech recognition runs through a local Whisper model, and processing happens in stages of ten to twenty calls with a spot check after each. His product is not the individual set of minutes but a pattern register: the recurring questions together with your own best answer, which becomes a playbook, an FAQ, an objection library and raw material for content. What he does not deliver is a new opinion. How the procedure runs in detail is in Conversation analysis with AI: the pattern register, which data goes towards AI here is in AI and data protection: what the AI may see, and the other roles I collect under AI employees.
Frequently asked questions
Can ChatGPT write minutes from a transcript?
Yes, and so can Claude or Copilot, not just ChatGPT. What matters is the instruction. Ask for a summary and you get prose about the course of the conversation. Specify a structure of decisions, tasks with owner and deadline, and open items, and demand supporting passages, and you get minutes you can forward.
What is the difference between a summary and minutes?
A summary shows what was discussed, minutes show what applies now. The summary looks back, the minutes set the terms going forward. That is why minutes consist of decisions, tasks and deadlines, and why a task without a name and a date isn't a task but a wish.
Is Copilot better for this than ChatGPT or Claude?
For the minute-taking itself the instruction makes the difference, not the product. Copilot has the practical advantage of already sitting on the Microsoft data the transcript came from. On the data question the providers differ more clearly than on text quality, and there it pays to read your plan.
Can I simply put a client call into a language model?
Two steps that often get mixed up. Recording the non-publicly spoken word without authorisation is a criminal offence under § 201 of the German Criminal Code, and what authorises the recording is prior consent from everyone else involved. The later processing by a provider needs its own legal basis under the GDPR. In practice: ask first, replace the names, set a retention period.
How do I tell whether the model invented something?
Through the second pass: demand a literal quote from the transcript for every decision and every task, and have everything deleted that has none. That catches invented entries, not misfilings. Then check three places by hand where guessing usually happens: names that were never spoken, dates calculated out of vague statements, and ideas promoted to decisions.
How to take this further
Take your last transcript and run it twice: once with "summarise this", once with the instruction above including the checking pass. The gap between the two results is exactly the part of the work you used to do by hand or not at all. Whether a task like this belongs with AI at all is something I decide beforehand by asking which problem it is meant to solve; how I go about that is in my AI strategy. You don't build your own version of the instruction alone: others share theirs, and in the calls we go through yours together.
And if after that you want to evaluate not one conversation but your whole stack of them, Gustav is available as a finished package in my community, with the anonymisation pass, consent templates and an invented demo transcript to practise on. The templates are samples without warranty and do not replace a lawyer's review of your case. Everything about it is at Community.