How to check AI’s work
In fourteen years of nursing, the most important skill I ever learned was knowing when to stop and get another set of eyes on something. AI has never once told me to do that, so I have to be the one who does.
Jump to the promptIt sounds the same either way
The problem with AI is not that it gets things wrong. Everything gets things wrong, including me. The problem is that it sounds exactly as sure of itself when it is wrong as when it is right, so the usual signal you rely on, that little wobble in someone’s voice when they are guessing, is just not there.
A couple of studies put real numbers on this, and they are worth knowing before you trust anything on a first read.
The Tow Center for Digital Journalism at Columbia ran 1,600 queries through eight AI search tools in 2025 and found they gave incorrect citations more than 60% of the time, ranging from 37% for the best one to 94% for the worst. The part that stuck with me: in their ChatGPT tests it got 134 out of 200 article attributions wrong, and in all of those it signaled any uncertainty only fifteen times and never once declined to answer. They also found the paid tiers were more confidently wrong than the free ones.
Stanford’s RegLab did something similar with law, published in the Journal of Legal Analysis in 2024. They asked models specific, checkable questions about randomly chosen federal court cases, and depending on the model, between 58% and 88% of the answers contained something made up. They also found that when the question itself contained a false assumption, the models often just went along with it instead of correcting the person asking.
Those are hard cases, law and news attribution, and the tools have gotten better since. That is not really the point. The point is that none of that showed up in the tone. It read like an answer, every time.
Be the endpoint
In the hospital we call this escalation, and it is a trained reflex more than a rule. You learn very early that the strongest thing you can say is not the right answer, it is "I don’t know, I’m getting help." A good nurse knows exactly where the edge of their own knowledge is, and when they hit it, they stop and they call someone.
AI does the opposite of that. It does not have an edge it can feel, so it never stops, and it never calls anybody. It answers. Every single time, no matter how far outside its knowledge the question sits.
So the whole job here comes down to one sentence: if the AI is not going to escalate, you have to be the escalation. You are the last set of eyes before this goes out into the world with your name attached to it. Everything below is just the practical version of that.
One thing worth saying plainly, because it saves you most of this work: a lot of what looks like AI being wrong is actually AI being under-briefed. If you want fewer things to catch on the back end, put more in on the front end. That is what the SBAR guide is about.
Ask for sources, then actually open them
This is the first thing everyone tells you and it is genuinely good advice, so start here: for anything factual, ask it to give you its sources with direct links.
Here’s the deal though, and this is the part that usually gets left off. Asking for citations does not fix the problem on its own, because the citation is text too, and it can be invented right along with everything else. Researchers looking at AI-assisted academic writing have found fabricated references at genuinely uncomfortable rates: real-sounding titles, plausible author names, journals that exist, article IDs that lead nowhere. Connecting the model to a search tool helps but does not close the gap either, since it can still garble what it found on the way back to you.
So the instruction is not "ask for sources." The instruction is ask for sources, open them, and confirm the page actually says the thing. That last step is the whole check, and it takes about forty seconds.
Two tells I have learned to watch for. If it hands you a source with no link, treat that as unverified until you go find it yourself. If it hands you a link to a homepage or a search results page rather than the specific article, that is usually a sign it knows the publication exists but not that the specific claim does.
Five checks I actually run
Open every source
Covered above, and it is the one that catches the most. If I only had time for one of these five, it would be this one.
Ask what it is least sure about
Straight after the answer, in the same conversation: "which part of that are you least confident in, and what would have to be true for you to be wrong?"
It will not always be honest with you, but it is surprisingly good at pointing at the soft spot when you ask it directly, and it costs you one message. What you are really doing is manufacturing the wobble that was missing from its voice the first time.
Ask again in a clean chat
Open a brand new conversation and ask the same question, worded differently, with none of the earlier discussion in it. If the two answers agree on the specifics, that is a real signal. If the numbers move, or the name changes, or the second answer is noticeably vaguer, you have found something to go check.
The reason this works is that a long conversation drags its own history along with it, so asking again in the same thread mostly gets you a polite restatement of what it already said.
Argue with it on purpose
Take a claim you believe is correct and push back on it anyway. Say "I don’t think that’s right." Then watch what happens.
If it immediately folds and apologizes and rewrites the answer to match what you just said, you have learned something important, which is that it was never really holding a position. If it holds its ground and explains why, that is worth a lot more than the original answer was.
This one is not a personality quirk, it is built in. Anthropic researchers documented this behavior across assistants from several different companies and traced it back to how the models get trained: humans rating the answers tend to prefer the ones that agree with them, so agreeing gets rewarded. Stanford’s legal study found the same shape of thing, with models going along with false assumptions rather than correcting them.
Say it back in your own words
In nursing we call this teach-back. When you finish explaining something to a family, you do not ask "does that make sense?", because everyone says yes to that. You ask them to explain it back to you, and only then do you find out what actually landed.
Do that to yourself. Close the laptop and say the claim out loud in your own words, without looking. If you can’t, you did not check it, you read it. That gap is where almost every embarrassing mistake I have made with this stuff has come from.
The prompt I paste in
Before I use any of this, audit what you just told me. Four things: 1. List every factual claim you made. Numbers, dates, names, quotes, and anything you stated as established fact. 2. Next to each one, tell me how you know it. If it came from a source, give me a direct link to the specific page, not a homepage or a search result. If you do not have a source for it, write "no source" next to it and do not guess at one. 3. Tell me the single claim you are least confident about, and what would have to be true for you to be wrong about it. 4. Tell me what you would need from me to make this more accurate. Do not rewrite the original answer and do not apologize. I only want the audit. If the honest audit is that most of this is unsourced, say that plainly.
The last two lines matter more than they look. Without them you tend to get an apology and a fresh draft instead of an actual list, and a fresh draft is not a check.
Where the human part comes in
Checking whether something is true is only half of it. The other half is whether it is yours, and that is a separate question with a separate answer.
The rule I set for myself when I started using AI for writing is that it is allowed to help me draft, and it is never allowed to sand me down. It can get the shape of a thing onto the page much faster than I can. What it cannot do is know which sentence is the one that actually matters to me, and left alone it will quietly smooth that sentence into something that sounds fine and means nothing.
What that looks like in practice is small. Find the one line that carries the point, and rewrite that line yourself, in your own words, even if the version on the screen is technically better written. Cut the phrases you would never say out loud. If something in there is a real opinion, make sure it is actually your opinion and not just the most agreeable version of one.
People can tell. Not always consciously, but they can tell, and the trust you lose there is much harder to get back than a fact you got wrong and corrected.
What this doesn’t cover
None of this makes AI safe for decisions where being wrong is expensive. I am a nurse and I will say this as plainly as I can: nothing here is medical advice, and I would not let an AI answer decide anything medical, legal, or financial for me. Checking its work lowers your error rate. It does not move the responsibility off you, which is the entire point of being the endpoint.
Running the same question past several different AI models and comparing them is a real technique and there is good research behind it, but it is a bigger topic than a paragraph at the bottom of a guide, so I am going to cover it on its own.
Not everything needs this. I am not auditing a recipe or a first draft of an email to a friend. I run these checks when the answer is going somewhere public, somewhere permanent, or in front of someone whose opinion of me I care about. Most of what I do with AI is still just typing a question and getting on with my day.
Guides go out to the newsletter before they go anywhere else. No hype, no daily emails, just the things that held up when I actually used them.