AI Prompts for Support QA and Tone: 15 That Fix Replies Before They Send (2026)

Nick Timms
Nick Timms, Co-founder
September 2, 2026·7 min read·verifiedReviewed by Duda Bardavid

15 AI prompts for support QA and tone: rewrite passes that fix drafts before they send, pre-send checks, a scoring rubric, and three that QA the real queue.

Key takeaways

  • QA prompts are the reverse of drafting prompts: instead of asking the AI to write a reply, you ask it to find what is wrong with the one you wrote. The second job catches what the first cannot.
  • The most valuable check happens before sending: did the draft answer everything asked, what did it just promise, and how could it be read wrong. Thirty seconds of prompt beats a day of cleanup.
  • Team tone is a QA problem, not a talent problem. A written voice rubric plus a scoring prompt turns five people answering five ways into one voice with five typists.
  • Prompts 1 to 12 work by copy-paste in ChatGPT, Claude, or Gemini. The last three run QA on a connected inbox: real sent replies, scored against your rubric, with the receipts.
Table of contents

The short version: these fifteen prompts do the opposite of every drafting prompt on this blog. Instead of asking the AI to write a reply, they ask it to find what is wrong with the one you already wrote: the question it dodged, the promise it accidentally made, the tone that will read colder than you meant it. The first twelve work by copy-paste in ChatGPT, Claude, or Gemini. The last three run against a connected inbox and QA the replies your team actually sent.

One boundary, so you land on the right page: this is the quality page, not the hard-tickets page. If the problem is writing the reply in the first place (the angry customer, the refund refusal), that is the customer service prompts library. This page is for when a draft exists and the question is whether it should be sent.

Rewrite prompts: fix the draft, keep the message (1-4)

Rewriting is the highest-frequency QA there is, and the failure mode is always the same: the fix erases the message along with the problem. Every prompt here names what must survive the rewrite.

1. The tone shift. Blunt to warm without dissolving into mush.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Rewrite this reply to sound warmer without weakening it: [paste draft]. Keep every commitment, date, and refusal exactly as strong as it is. Change only the temperature, not the content. Show me what you changed and why.

2. The de-jargon pass. For replies written by the person who fixed the problem.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Rewrite this so a customer with no technical background understands it on first read: [paste draft]. Keep the accurate cause, drop the internal vocabulary, and do not make it longer. If a technical term has to stay, make the sentence around it carry the meaning.

3. The shrink. Long replies read as effort; they are usually avoidance.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Cut this reply to under 100 words: [paste draft]. Everything the customer needs to know or do survives; everything that is context for us rather than for them goes. List what you cut so I can veto.

4. The bad-news clean. Clear delivery without the corporate chill.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
This reply delivers news the customer does not want: [paste draft]. Rewrite it so the news arrives in the first two sentences instead of the last, the reason is honest and brief, and what we can do appears immediately after what we cannot. No "unfortunately" pileups, no policy-speak.

Pre-send checks: the thirty-second review (5-8)

These four are the checks that catch problems while they are still free to fix. They work best as a habit on anything going to an unhappy reader.

5. The completeness check. The most common reply failure is answering the easy half of the email.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Here is the customer's email and my draft reply: [paste both]. List every question and request in their email, and mark which ones my draft actually answers. If any are dodged or half-answered, quote the dodge back to me.

6. The promise audit. You are accountable for what the draft commits to, including the commitments you did not notice making.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Read my draft: [paste]. List every commitment it makes: dates, refunds, features, follow-ups, exceptions to policy. For each, tell me whether it is stated firmly or vaguely, because vague promises get remembered as firm ones.

7. The misread test. Every reply has one sentence that can be taken the wrong way. Find it before the customer does.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Read my draft as an already-frustrated customer would: [paste draft, plus their last message]. What is the least charitable reading of each sentence? Flag the one most likely to make things worse, and offer a rewording that closes the misreading without changing the meaning.

8. The thread-context check. A perfectly good reply can still be wrong for this customer.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Here is the full thread and my draft reply: [paste]. Does my draft contradict anything we told this customer earlier, repeat anything they were already told, or ignore something they have said twice? Quote the clash if you find one.

Team QA prompts: one voice, honest scores (9-12)

Solo senders can stop at prompt 8. Team inboxes have a second-order problem: five people answering in five voices at five quality levels, and nobody keen to police colleagues. A rubric and a scoring prompt make the standard explicit and the feedback impersonal.

9. The rubric builder. You cannot score replies against a standard that lives in one person's head.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Draft a scoring rubric for our support replies, five criteria, each scored 1 to 3 with a one-line description per score. Base it on these three replies we consider excellent: [paste]. The rubric should reward what these do, not generic advice.

10. The reply scorer. The rubric, applied without office politics.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Score this reply against our rubric: [paste rubric and reply]. Give a score and one sentence of evidence per criterion, then the single change that would most improve the score. Do not rewrite the reply.

11. The coaching note. Feedback a teammate can use, without rewriting their work for them.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Here is a teammate's reply and our rubric: [paste]. Write two sentences of feedback I can send them: one naming specifically what worked, one naming the highest-impact improvement. Their voice stays theirs; do not include a corrected version.

12. The voice-consistency check. The customer should not be able to tell who was on shift.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Here are recent replies from three teammates to similar questions: [paste]. Describe each one's voice in a line, then tell me where they diverge most (greeting, directness, sign-off, handling of apology). Propose the two or three voice rules that would close the gap with the least change to anyone's writing.

Drag AI

The inbox your team and your AI work in together

Shared inbox, live chat, and AI in Gmail, with an MCP server your AI tools can drive.

Start free trial
ChromeWebDesktopMobileAPIMCP

The prompts that QA the real queue (13-15)

Everything above works on what you paste. Connected to the inbox through an MCP server, the same review runs on what your team actually sent, and the scores come with receipts. The connection guide covers the setup in about ten minutes.

13. The weekly QA sample. The review that never happens by hand, because sampling is tedious.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Pull ten customer conversations we closed this week, spread across teammates. Score each final reply against our rubric [paste rubric], with one line of evidence per score. End with the two patterns most worth fixing team-wide.
Drag logoReply, using Drag's MCP tools
Scored ten closed conversations across three teammates. Average 12.1 of 15. Strongest criterion: correctness, ten of ten across the sample. The two team-wide patterns: five replies answered the question but named no next step or date (ownership averaged 1.8 of 3), and sign-offs range from "Cheers" to "Kind regards" with the same customer. Per-reply scores with quoted evidence follow.

What happens: the assistant pulls closed conversations from the connected inbox, reads the real replies, and scores them with quotes as evidence. The tedious part of QA, the sampling and reading, stops being the reason it never happens.

14. The tone-drift report. Voice rules decay quietly. This catches the drift while it is small.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Read this week's sent replies across the team and compare them against our voice rules [paste rules]. Who is drifting, on which rule, with one quoted example each? Keep it factual; this feeds coaching, not scorekeeping.

What happens: search_threads surfaces the week's replies, and the comparison cites real sentences rather than impressions. The report lands as evidence a team lead can actually use.

15. The knowledge-gap miner. Wrong answers are a QA problem before they are a documentation problem.

Claude logoChatGPT logoGemini logoWorks in Claude, ChatGPT, or Gemini
Compare this week's replies against our knowledge base. Find any reply that contradicts a documented answer, and any question we answered from memory that has no documented answer at all. List both: the first is a correction, the second is an article we owe ourselves.

What happens: search_knowledge checks the connected knowledge base while the assistant reads the replies, so contradictions surface with both sides quoted. Each run either fixes an answer or writes the missing article's brief.

Where the prompts stop and the system starts

A QA prompt improves one reply; a QA habit improves one person. The team-level fix is structural: shared drafts reviewed before sending, a knowledge base the answers come from, and a record of who said what. For disclosure, that is the product we build: Drag turns Gmail into a shared inbox with Drag AI drafting in the team's voice from $18 a seat, and its MCP server (Pro plan) is what makes prompts 13 to 15 run. Everything else on this page needs no Drag account, just a clipboard and any assistant.

FAQ

Can AI review my email before I send it?

Yes, and review is where AI is most reliable, because judging a draft is easier than writing one. The strongest pre-send checks are completeness (did it answer everything asked), commitments (what did it promise), and misreading (how could it land wrong). Paste both the customer's email and your draft; the thread is the context that makes the review accurate.

How do I keep customer service replies consistent across a team?

Write the voice down. Turn your three best replies into a short rubric and a handful of voice rules, then score against them rather than debating taste. Consistency comes from the standard being explicit; the scoring prompt just applies it without anyone having to critique a colleague face to face.

What should a support reply QA rubric measure?

Five criteria cover most teams: correctness of the answer, completeness against what was asked, tone matching your voice rules, ownership (a named next step and date), and economy. Score each 1 to 3 with written descriptions per level, and derive the rubric from your own best replies so it rewards your standard, not a generic one.

The rest of the prompt library

This page is one of seven in the Drag prompt library. The support reporting prompts read the numbers these reviews produce. For writing replies from scratch, the customer service prompts take the hard tickets and the ChatGPT and Claude pages cover everyday drafting. For the queue itself, the triage prompts sort it and the shared inbox prompts run it as a team.

Nick Timms

Nick Timms

Co-founder

Building Drag for nearly ten years: shared inboxes, boards, and now the AI and agent layer, all on Gmail, plus HeyHelp for the personal inbox. Writes the honest versions of the comparisons.

Drag AI

The inbox your team and your AI work in together

Shared inbox, boards, live chat, and WhatsApp with AI included, in Gmail and beyond, plus an MCP server your AI tools can drive.

7-day trial, no card required4.7 · 1,200+ reviews
Gmail extension
Web app
Desktop
iOS + Android
API
MCP server: works with Claude, ChatGPT, Gemini, Copilot, Cursor