AI Quality Control Checklist for Virtual Assistants (Philippines, 2026)

AI quality control checklist for virtual assistants working with Filipino clients in 2026

An AI quality control checklist for virtual assistants is the short set of checks you run on every AI-assisted deliverable before your client sees it: verify each fact against a source you opened yourself, recompute every number, confirm every name, date, link, and amount, read it once in your client’s voice, and make sure it answers what was actually asked. Ten to fifteen minutes. That is the difference between AI making you faster and AI making you the VA who sent a client a quote that does not exist.

I started as a VA in 2020, before ChatGPT was something anyone used for work. When I was managing Shopify stores, I wrote every product description by hand, so checking my own output was not a separate step, it was just the work. AI changed the speed. It did not change who is responsible when the output is wrong, and the tool does not sign your contract.

Key takeaways

  • Five categories catch almost everything: facts, numbers, names, dates, and links.
  • Low hallucination rates are not zero. A 2026 benchmark across five frontier models and 5,000 prompts reported rates between 3.1% and 19.1%.
  • The costly failures are invented specifics. Deloitte Australia repaid about $63,000 after a government report it delivered was found to contain a fabricated court quote and references to papers that do not exist.
  • Philippine work needs its own pass. Peso amounts, BIR forms, local dates, and holiday calendars are what a US-trained model gets confidently wrong.
  • Upwork’s ethics guidance tells freelancers to always disclose to clients whether they use AI-generated content or tools on a project.

What goes on an AI quality control checklist for virtual assistants?

Ten checks, in this order, and most take a minute or less. The order matters: invented facts, wrong numbers, and wrong names are cheap to catch early and humiliating to catch after your client has forwarded your work.

CheckWhat you are looking forTime
1. Facts and claimsEvery stated fact has a source you opened yourself3 to 5 min
2. Numbers and mathTotals, percentages, and currency recomputed outside the chat2 min
3. Names and titlesPeople, companies, and products spelled right and real1 min
4. Links and quotesEvery URL clicked, every quote found in the actual document2 min
5. Dates and deadlinesDay of week, time zone, and year all correct1 min
6. The instructionThe output answers what your client actually asked for1 min
7. Voice and toneReads like your client, not like a chatbot2 min
8. FormattingNo leftover placeholders, stray markdown, or broken tables1 min
9. Sensitive contentNothing confidential that should not be in the output at all1 min
10. Final readOne slow read out loud, top to bottom2 min

Two carry more weight than the rest. The fact check is the one that ends contracts. The instruction check is the one that quietly loses you the client, because AI is very good at answering a slightly different question than the one you were asked, and proofreading will never catch that. Read your client’s original message again after the draft is done, not before.

The mental model I keep coming back to came from an accounting professor commenting on the Deloitte case in CFO Dive in October 2025: AI output must be reviewed as if an intern or a new hire prepared it. You would not forward an intern’s first draft unread, or assume the intern verified the statistic. If you are still building your stack, my guide to the AI tools every Filipino VA should learn in 2026 covers what to use. This post is what happens after the tool hands you an answer.

Where does AI actually break in client work?

It breaks where it has to be specific: citations, names, numbers, policies, and anything recent. General explanations are usually fine. The moment the model needs a precise detail it does not have, it invents one that looks right, and never tells you it guessed.

The clearest recent example is not a freelancer, it is Deloitte. Australia’s Department of Employment and Workplace Relations commissioned a 237-page review under a contract worth roughly $290,000. After publication, a Sydney University researcher found fabricated references in it, including a made-up quote from a federal court judgment and citations to academic papers that do not exist. A revised version was published disclosing that Microsoft’s Azure OpenAI had been used in drafting, and Deloitte repaid over $63,000, confirmed by the government on October 21, 2025. A firm with a real review process still shipped hallucinated citations to a government client. That tells you what happens with no review.

Support work has its own version. In April 2025, users of the coding tool Cursor were getting logged out across devices, and the company’s AI support bot told them their subscription was limited to one active session. No such policy existed. The bot invented it, users cancelled, and the CEO apologized publicly and refunded those affected. If you handle inbox or tickets, that is your risk: an AI answering a policy question your client never set.

Transcription is the quiet one. In a study presented at the ACM FAccT conference in 2024, researchers examining OpenAI’s Whisper found roughly 1% of audio transcriptions contained entire hallucinated phrases that appeared nowhere in the audio, and 38% of those hallucinations carried explicit harms. I use an AI note-taker on client calls myself, because I am an introvert and would rather listen than type frantically, but I read the summary against my own notes first, kasi a recap that invents one sentence is worse than no recap. A 2026 benchmark of five frontier models and 5,000 prompts put hallucination rates between 3.1% and 19.1%, and one wrong answer in thirty is not a rate you can ship unchecked.

How do you verify facts, numbers, and names in under ten minutes?

Check the claim at its source, never inside the same chat window that produced it. Asking an AI to fact-check itself is asking the system that invented a detail to notice it invented one, and it will often defend the answer or invent a source. OpenAI keeps a line under the ChatGPT input box telling you it can make mistakes. Take that literally.

Here is the sequence I would use if I were starting out today. Ask the model to list its factual claims and figures separately from the prose, so you get a checklist instead of a wall of text. Open a source for each one yourself: the company’s page, the government agency, the platform’s help center. Recompute every number in a calculator or a spreadsheet, including totals, percentages, and currency conversions. Search every proper name in quotes to confirm it exists and is spelled right. Then click every link, because a dead URL is a five-second catch and a public embarrassment.

One rule has no exceptions: never pass along a quote you have not read in the original document. That habit alone would have caught the Deloitte problem, and it is the check people skip most, because a fake quote usually sounds plausible.

As a VA since 2020 who has worked with multiple clients and one stable client for 5 years, I would advise you to record what you checked, not just that you checked. When I hit a problem in client work, I research it first, then write down what I found and what I did about it. Two lines under each deliverable naming the figures you verified and where takes thirty seconds, and becomes your defence the day a client questions something.

What extra checks does Philippine client work need?

Anything local, dated, or in pesos. Most of the training data behind these tools is American, so the model is weakest exactly where our work is most specific: BIR forms, holiday calendars, peso amounts, local addresses, and Filipino names.

Holidays are the cleanest example, and they matter if you manage a calendar or a payroll schedule. Proclamation No. 1006 declared the regular holidays and special non-working days for 2026, but deliberately left out the Islamic holidays, because those dates are only fixed later, once the National Commission on Muslim Filipinos recommends them. Eid’l Adha was declared a regular holiday for May 27, 2026 by a separate proclamation issued in May of that year. So a 2026 Philippine holiday calendar generated by AI early in the year is not just possibly wrong, it is structurally incomplete. Check the Official Gazette, not the chat.

Money is second. Never let a model convert dollars to pesos for an invoice or a rate quote, because rates move daily and the model works from whatever it absorbed months ago. Pull the rate from the platform paying you, and write amounts the way your client’s records expect them: ₱50,000 with the peso sign, $1,200 with the dollar sign. If the work touches tax, treat every AI answer as a starting question. My posts on BIR registration for freelancers and the 1701Q filing deadlines exist because those details change and the penalty lands on a real person.

Then the small stuff that signals sloppiness: Filipino names autocorrected into something else, barangay and city fields collapsing into one mangled line, and dates flipping between the day-first format we write and the month-first format US clients expect. None are dramatic alone. Together they make a client wonder what else you skipped.

How do you keep AI output sounding like your client, not like a chatbot?

Give the model real samples of your client’s writing and edit against them, instead of prompting for a tone in the abstract. Save five to ten things your client actually wrote, their emails, captions, past newsletters, into one file, and paste the relevant ones as reference every time. “Write in a friendly professional tone” gives everyone on earth the same output. “Match the voice in these three emails” gives you something your client recognizes.

Then strip the tells on the read-through. Bullet lists where one sentence would do. Words nobody in your client’s business actually says. Long dashes dropped into sentences your client would never write. And the biggest tell, hedging: AI writes “it is important to consider” where your client would just say the thing. I occasionally handle hiring and interview candidates, and I can tell when someone bulk-sent a template without reading the job post. Unedited AI output reads the same way, technically fine but obviously not written for this person.

Now the part that matters more for us than for a client in Sydney or Chicago. AI detectors are unreliable in a way specifically stacked against non-native English writers. A study published in the journal Patterns in 2023 tested seven widely used GPT detectors and found they wrongly labeled an average of 61.3% of human-written TOEFL essays by non-native English writers as AI-generated, while classifying essays by US students correctly. A flag is not evidence, and you can get flagged for writing that is entirely your own. Keep your drafts, your Google Docs version history, and your notes. That record answers better than any argument.

How much time should checking take, and what if a mistake still gets through?

Budget roughly 15% to 20% of the task time, and never zero. On a one-hour deliverable that is ten to twelve minutes. Tier it by stakes: rough drafts get a light pass, but anything client-facing, anything with a number in it, and anything touching health, money, or a legal matter gets the full checklist. If you do bookkeeping or medical VA support, assume nearly everything you touch sits in that top tier.

Price it in your head as work, not a favor. If you are on a ₱35,000 monthly retainer, the checking is not an unpaid extra, it is part of what your client is buying and the reason they pay a person instead of a $20 subscription. It also stops you cutting the check on the days you are behind, which is when you need it most. High-trust work pays better because your client stops having to verify you, which is the argument in my post on executive VA roles.

When something does get through, tell them the same day, in writing: what was wrong, what you have corrected, where else it might have travelled, and what you changed so it does not repeat. My habit is to research first, fix what I can, then report the outcome with a proposed next step instead of handing the problem back. Deloitte’s report was revised, republished, and given an AI-use disclosure, and the department said the substance was retained. A corrected mistake plus a fixed process survives. A hidden one does not.

One honest note

Every figure above is what the cited source said when I checked on August 1, 2026. Published hallucination rates swing hard depending on who ran the study and what they tested, so treat the 3.1% to 19.1% range as directional. And a checklist lowers your risk, it does not remove it. The goal is that the mistakes slipping through are small ones you catch next week, not a fabricated quote your client forwards to their board.

Frequently asked questions

How long should AI quality control take on a normal VA task?

Budget about 15% to 20% of the task time, roughly ten to twelve minutes on a one-hour deliverable, and never zero. Internal notes get a light pass, while anything client-facing or involving money, health, or a legal matter gets the full checklist. The check that saves you most often is the slow read out loud, because it catches tone problems, leftover placeholders, and sentences that say nothing.

Do I have to tell my client I used AI?

Read your contract first, because many now include clauses covering third-party tools. On Upwork, the platform’s own ethics guidance is explicit: always disclose to clients whether you use AI-generated content or tools during a project. Disclosing plus explaining how you check the output reads as competence, not confession.

What if my client runs my work through an AI detector and it flags it?

A flag is not proof. A 2023 study in the journal Patterns tested seven widely used GPT detectors and found they wrongly labeled an average of 61.3% of human-written TOEFL essays by non-native English writers as AI-generated, while classifying US student essays correctly. That bias hits us directly. Keep your drafts and version history, and offer those instead of arguing about the tool.

Can I just ask ChatGPT to fact-check its own answer?

No, not as your only check. You are asking the same system that produced a detail to notice it invented one, and it will often defend the answer or supply a source that does not exist. Use it to list its claims and figures cleanly, then verify each one yourself against the original document or the agency’s own page.

Checking is the job now

AI made drafting fast and made checking valuable. Everyone has the same tools now, so what separates you is no longer how fast you produce something, it is whether what you send can be trusted without your client re-reading it.

Build the habit on your next task, not someday. List the claims. Open a source for each. Recompute the numbers. Click the links. Read it out loud once. Kaya mo ‘yan. Then tell your client this is how you work, because almost nobody is saying it, and being the VA whose work does not need double-checking beats being the fastest one.

Sources

Jean Aguilar

Jean Aguilar

I’m a Filipina VA based in Cavite. I started in 2020 as a data-entry VA and worked my way up to Shopify manager and operations roles. I started PinoyRemote to share what actually worked, so you can skip the guesswork na pinagdaanan ko the hard way. Connect on LinkedIn →

Keep reading