What AI Can and Cannot Do
The promise: Leave with an honest capability map — and a six-question test you can run on any task before letting AI near it.
Outcome: After this article, the learner can name where today's AI is genuinely strong and where it is unreliable, correct the six most common misconceptions, and sort their own tasks into Draft, Delegate, or Don't.
Who this is for and when to use it
For anyone about to make a decision that touches AI: buying a tool, starting a pilot, reassuring a nervous team, or pushing back on an overexcited one. It pairs with What AI Actually Is, which covers vocabulary; this one covers judgment.
Use it before your first project, before any purchase, and whenever your feed serves you either "AI does everything now" or "AI is a bubble." The truth is more useful than either.
The honest capability map
Modern AI — especially generative AI — is strongest at work that is language-shaped or pattern-shaped, where a decent first draft is valuable and a human polishes the result. It is weakest where a wrong answer is expensive to detect.
Genuinely strong:
- First drafts of almost anything written. Emails, summaries, job postings, outlines, meeting recaps. Fast, competent, tireless.
- Summarizing and extracting. Turning a long document into key points, pulling action items from notes, converting formats.
- Rewriting. Same content, different tone, length, or audience.
- Brainstorming breadth. Twenty options in thirty seconds; humans pick the three that fit.
- Sorting and classifying at volume. Routing messages, tagging feedback, triaging requests — with spot checks.
- Availability. A capable first pass at 2 a.m., every day, without fatigue.
Unreliable without human help:
- Facts. Generative systems can produce plausible, confident, wrong statements — including invented citations, names, and figures. This is called hallucination, and it is a byproduct of how the technology works, not a rare glitch.
- Recent events. Unless a tool is connected to live search, it answers from training data with a cutoff — and even with search, sources need checking.
- Precise math and counting. Without a connected calculator or spreadsheet, multi-step arithmetic is a gamble.
- Your context. It does not know your customers, pricing, history, or politics unless you provide them — and it will not tell you what you forgot to mention.
- Consistency. The same question can produce different answers on different runs.
- Judgment and accountability. It can lay out options and tradeoffs; it cannot own a decision, weigh your values, or be responsible for an outcome. That work stays human.
- The physical world, and true edge cases. Anything outside the patterns it learned gets shaky fast.
| Task area | Genuinely strong | Unreliable without help | The human's job |
|---|---|---|---|
| Writing and drafting | Fast, competent first drafts | Facts, names, and figures inside the draft | Verify specifics; add voice and judgment |
| Summarizing | Long document to key points | Subtle emphasis; may miss what matters to you | Check against the source for critical uses |
| Brainstorming | Breadth of options, quickly | Knowing which option fits your constraints | Choose and refine |
| Sorting and classifying | Volume, speed, consistency | Novel or ambiguous cases | Define categories; sample-check outputs |
| Facts and research | Orientation, vocabulary, leads | Citations and specifics can be invented | Verify with primary sources |
| Math and data | Explaining approaches | Multi-step arithmetic without tools | Use calculators and spreadsheets; check results |
| Current events | — | Anything after its training data, unless search-connected | Confirm dates and sources |
| Judgment calls | Framing options and tradeoffs | Owning the decision | Decide, and be accountable |
Hype versus reality
| You may have heard | Closer to the truth |
|---|---|
| "AI understands what it says" | It predicts likely words. Fluency is not comprehension; judge the substance. |
| "AI will replace entire jobs overnight" | It changes tasks first — usually the repetitive ones. Measured redesign beats panic in either direction. |
| "It's always learning from your chats" | Most assistants do not improve from your usage by default; some do use inputs for training. Check the settings (verify current options). |
| "It searched the web for that answer" | Only some tools search; many answer from training memory. Know which mode you are in. |
| "It got one thing wrong, so it's useless" | Wrong answers prove it needs review — like every first draft you have ever received. |
| "The newest model can do everything" | Capabilities shift quickly and unevenly. Test on your own tasks before trusting claims (verify current capabilities). |
A realistic example: Bright Smile Dental
Bright Smile Dental is a fictional nine-person dental office. Over one quarter, here is what worked and what did not.
Worked well: drafting patient recall and reactivation emails (front desk reviewed each one — drafting time dropped from about 15 minutes to about 4, per their own log); summarizing a 30-page supplier contract into a list of questions for their attorney (a better meeting, not a replaced lawyer); brainstorming twenty ideas to reduce no-shows, of which three were worth piloting.
Caught by review: the assistant drafted a reply to a patient's medication question — the office manager stopped it, because clinical answers come from clinicians, full stop. It also produced an insurance coverage figure that was plausible and wrong; now every number a patient hears comes from the insurer's portal, never from a generated draft.
Their standing rule, one sentence long: AI drafts words; verified systems provide numbers; licensed humans give clinical answers. (All details fictional; the pattern is the takeaway.)
The Draft / Delegate / Don't test
Run these six questions on any task you are considering:
- Stakes. What happens if the output is wrong and slips through?
- Verifiability. Can a knowledgeable person check the output in well under half the time it would take to do the task from scratch?
- Data. Does the task require sensitive or confidential data? Is there an approved tool for that — or can the data be removed?
- Shape. Is the task language- or pattern-shaped: drafting, summarizing, sorting, rewriting, transforming?
- Review. Is there a named person who reviews the output before it counts?
- Measurement. Can you compare before and after — minutes, error rate, response time?
Then sort:
- Draft (the default). Pattern-shaped task plus a named reviewer: AI produces the first pass, a human finishes and approves. Most of the real value lives here.
- Delegate (earn it). Low stakes, easy verification, highly repetitive: let it run with periodic spot checks — for example, review one output in ten, forever.
- Don't (or redesign). High stakes plus hard-to-verify output, or sensitive data with no approved tool: keep it human. Or redesign the task — strip identifying data, add a checkpoint, lower the stakes — and run the test again.
Common mistakes and how to avoid them
- The demo-to-production leap. A great demo is the tool's best 30 seconds. Pilot on 20 real, reviewed runs before you trust anything.
- Generalizing from one result. One failure does not mean useless; one success does not mean magic. Look at rates across many runs.
- Trusting confidence as competence. The tone never wavers, even when the content is wrong. Review the substance.
- Skipping measurement. Without before-and-after numbers, your AI debate is opinions. A simple wins log settles it.
- Assuming today's map is permanent. Capabilities and limits both move. Recheck quarterly — this page carries a review date for the same reason.
Safety, privacy, and human review
- Plan for hallucination; do not hope around it. Build verification into every workflow that produces facts, figures, or citations.
- Watch for bias. These systems learned from human-generated data, bias included. Sample-check outputs, and treat anything affecting people — hiring, lending, housing, health — as high-stakes work requiring human decision-makers and, often, legal review.
- Protect data. Sensitive information only goes into tools approved for it. When in doubt, leave it out or use fictional stand-ins.
- Accountability does not transfer. AI can inform judgment; it cannot replace it, and it cannot be accountable for it. A named person owns every consequential output.
Quick practice (10 minutes)
- List five tasks you or your team repeat every week.
- Run each through the six questions.
- Sort them into Draft, Delegate, and Don't. (Expect mostly Draft — that is normal and good.)
- Tomorrow, run your top Draft task with AI plus your review, and log the minutes saved. That single log entry is worth more than any headline you read this week.
Next recommended resource
You know what AI is and what it can do — now meet the tool landscape: Practical AI Tool Categories Explained.
Business readers: take your Draft list into the AI Opportunity Audit Worksheet — and if you want an experienced second opinion on that list, book a call at /contact. We will tell you honestly which items are worth automating and which are not.