15 AI prompts for support QA and tone: rewrite passes that fix drafts before they send, pre-send checks, a scoring rubric, and three that QA the real queue.
Key takeaways
- QA prompts are the reverse of drafting prompts: instead of asking the AI to write a reply, you ask it to find what is wrong with the one you wrote. The second job catches what the first cannot.
- The most valuable check happens before sending: did the draft answer everything asked, what did it just promise, and how could it be read wrong. Thirty seconds of prompt beats a day of cleanup.
- Team tone is a QA problem, not a talent problem. A written voice rubric plus a scoring prompt turns five people answering five ways into one voice with five typists.
- Prompts 1 to 12 work by copy-paste in ChatGPT, Claude, or Gemini. The last three run QA on a connected inbox: real sent replies, scored against your rubric, with the receipts.
Table of contents
The short version: these fifteen prompts do the opposite of every drafting prompt on this blog. Instead of asking the AI to write a reply, they ask it to find what is wrong with the one you already wrote: the question it dodged, the promise it accidentally made, the tone that will read colder than you meant it. The first twelve work by copy-paste in ChatGPT, Claude, or Gemini. The last three run against a connected inbox and QA the replies your team actually sent.
One boundary, so you land on the right page: this is the quality page, not the hard-tickets page. If the problem is writing the reply in the first place (the angry customer, the refund refusal), that is the customer service prompts library. This page is for when a draft exists and the question is whether it should be sent.
Rewrite prompts: fix the draft, keep the message (1-4)
Rewriting is the highest-frequency QA there is, and the failure mode is always the same: the fix erases the message along with the problem. Every prompt here names what must survive the rewrite.
1. The tone shift. Blunt to warm without dissolving into mush.
2. The de-jargon pass. For replies written by the person who fixed the problem.
3. The shrink. Long replies read as effort; they are usually avoidance.
4. The bad-news clean. Clear delivery without the corporate chill.
Pre-send checks: the thirty-second review (5-8)
These four are the checks that catch problems while they are still free to fix. They work best as a habit on anything going to an unhappy reader.
5. The completeness check. The most common reply failure is answering the easy half of the email.
6. The promise audit. You are accountable for what the draft commits to, including the commitments you did not notice making.
7. The misread test. Every reply has one sentence that can be taken the wrong way. Find it before the customer does.
8. The thread-context check. A perfectly good reply can still be wrong for this customer.
Team QA prompts: one voice, honest scores (9-12)
Solo senders can stop at prompt 8. Team inboxes have a second-order problem: five people answering in five voices at five quality levels, and nobody keen to police colleagues. A rubric and a scoring prompt make the standard explicit and the feedback impersonal.
9. The rubric builder. You cannot score replies against a standard that lives in one person's head.
10. The reply scorer. The rubric, applied without office politics.
11. The coaching note. Feedback a teammate can use, without rewriting their work for them.
12. The voice-consistency check. The customer should not be able to tell who was on shift.
Drag AI
The inbox your team and your AI work in together
Shared inbox, live chat, and AI in Gmail, with an MCP server your AI tools can drive.
The prompts that QA the real queue (13-15)
Everything above works on what you paste. Connected to the inbox through an MCP server, the same review runs on what your team actually sent, and the scores come with receipts. The connection guide covers the setup in about ten minutes.
13. The weekly QA sample. The review that never happens by hand, because sampling is tedious.
What happens: the assistant pulls closed conversations from the connected inbox, reads the real replies, and scores them with quotes as evidence. The tedious part of QA, the sampling and reading, stops being the reason it never happens.
14. The tone-drift report. Voice rules decay quietly. This catches the drift while it is small.
What happens: search_threads surfaces the week's replies, and the comparison cites real sentences rather than impressions. The report lands as evidence a team lead can actually use.
15. The knowledge-gap miner. Wrong answers are a QA problem before they are a documentation problem.
What happens: search_knowledge checks the connected knowledge base while the assistant reads the replies, so contradictions surface with both sides quoted. Each run either fixes an answer or writes the missing article's brief.
Where the prompts stop and the system starts
A QA prompt improves one reply; a QA habit improves one person. The team-level fix is structural: shared drafts reviewed before sending, a knowledge base the answers come from, and a record of who said what. For disclosure, that is the product we build: Drag turns Gmail into a shared inbox with Drag AI drafting in the team's voice from $18 a seat, and its MCP server (Pro plan) is what makes prompts 13 to 15 run. Everything else on this page needs no Drag account, just a clipboard and any assistant.
FAQ
Can AI review my email before I send it?
Yes, and review is where AI is most reliable, because judging a draft is easier than writing one. The strongest pre-send checks are completeness (did it answer everything asked), commitments (what did it promise), and misreading (how could it land wrong). Paste both the customer's email and your draft; the thread is the context that makes the review accurate.
How do I keep customer service replies consistent across a team?
Write the voice down. Turn your three best replies into a short rubric and a handful of voice rules, then score against them rather than debating taste. Consistency comes from the standard being explicit; the scoring prompt just applies it without anyone having to critique a colleague face to face.
What should a support reply QA rubric measure?
Five criteria cover most teams: correctness of the answer, completeness against what was asked, tone matching your voice rules, ownership (a named next step and date), and economy. Score each 1 to 3 with written descriptions per level, and derive the rubric from your own best replies so it rewards your standard, not a generic one.
The rest of the prompt library
This page is one of seven in the Drag prompt library. The support reporting prompts read the numbers these reviews produce. For writing replies from scratch, the customer service prompts take the hard tickets and the ChatGPT and Claude pages cover everyday drafting. For the queue itself, the triage prompts sort it and the shared inbox prompts run it as a team.
Co-founder
Building Drag for nearly ten years: shared inboxes, boards, and now the AI and agent layer, all on Gmail, plus HeyHelp for the personal inbox. Writes the honest versions of the comparisons.
