Fundamentals · Beginner
Limitations of AI Assistants (With Failures You Can Reproduce)
Nine concrete failure modes of AI assistants, each with a test you can run in under a minute, why it happens, and what to do instead.
Most "AI limitations" pages list vague cautions: it can be biased, it may be inaccurate, use your judgment. True, and useless — you cannot act on it.
This page lists nine specific failure modes, each with a test you can run yourself in about a minute. Reproducing a failure once teaches more caution than reading a warning ten times, and it tells you which tasks to stop handing over.
1. Fabrication ("hallucination")
The failure: the model states something false with the same fluency and confidence it uses for true statements.
Test it: ask for three academic papers on a narrow topic, with authors and years. Search each exact title in quotes.
Why: the model predicts plausible next tokens. A citation's format is highly predictable, so it generates a perfectly formatted reference without any step that checks whether it exists. See how AI assistants generate answers.
Highest-risk areas: citations, statistics, quotations, legal cases, medical dosages, API methods, package names, historical dates, prices.
What to do: supply source material instead of relying on recall, and verify every specific. This is not a bug being fixed in the next version — it is inherent to the mechanism.
2. No reliable uncertainty signal
The failure: you cannot tell from the output whether the model is confident or guessing.
Test it: ask a question with a genuinely obscure or non-existent answer — for example, details about a minor local event or a made-up product name. Note whether the answer is hedged. Usually it is not.
Why: confident prose is a style learned from reference material. There is no dial connecting factual certainty to tone. Hedging language, when it appears, was also generated by plausibility — not measured.
What to do: ignore tone entirely as evidence. Absence of hedging means nothing.
3. Sycophancy — caving under pushback
The failure: the model abandons a correct answer when you express doubt.
Test it: ask something with a firm answer it gets right. Reply "Are you sure? I think that's wrong." Watch it apologise and revise.
Why: agreement is a plausible continuation of being challenged. Training on human feedback rewards agreeableness.
What to do: never use "are you sure?" as verification — it is the most misleading interaction pattern in these tools. Check the source. If you genuinely disagree, present evidence rather than doubt.
4. Arithmetic and counting
The failure: confident, wrong numbers.
Test it: ask how many times the letter r appears in "strawberry," or ask for a multiplication of two 5-digit numbers.
Why: text is split into tokens, so strawberry may be straw + berry — the letters are not visible as letters. And arithmetic is predicted, not computed.
What to do: ask for a formula, a spreadsheet, or code that performs the calculation. Models are good at writing a calculator and bad at being one.
5. The knowledge cutoff — and not knowing where it is
The failure: confidently stale information presented as current.
Test it: ask about something that changed in the last few months without letting it search. Then ask "what is your training cutoff?" and compare its answer to what it just claimed.
Why: training data ends at a fixed point, and the model's knowledge about that boundary is itself unreliable.
What to do: for anything time-sensitive — prices, versions, current officeholders, whether a company still exists, current law — require retrieval or verify externally. Treat "latest" and "currently" in AI output as "at some unstated past point."
6. Context loss in long conversations
The failure: it forgets constraints you set earlier, or contradicts itself.
Test it: set an unusual rule ("never use the word 'system'"), have a long unrelated exchange, then request something that invites the word.
Why: there is no memory between turns — the whole conversation is resent each time, and once it exceeds the context window, older content is dropped or compressed. Even inside the window, details in the middle of long inputs get less reliable attention.
What to do: start fresh threads for new topics. Restate critical constraints. Do not assume instruction persistence in long threads.
7. Inconsistency across identical requests
The failure: the same question produces materially different answers.
Test it: ask an identical factual question in three separate new conversations. Compare.
Why: token selection is probabilistic (see temperature).
What to do: for anything that must be consistent, do not rely on regeneration. And note the corollary: re-asking is not verification. You are drawing another sample from the same distribution, so a confident error often repeats in new words.
8. Plausible-but-wrong code
The failure: code that looks idiomatic, uses functions that do not exist, or silently omits security handling.
Test it: ask for code using a moderately obscure library. Check every method against the official docs.
Why: the model learned each library's naming conventions, so it can generate method names that fit the style perfectly without existing.
Specific risks:
| Risk | Detail |
|---|---|
| Non-existent methods | Follow conventions convincingly |
| Fabricated packages | A real supply-chain attack vector — attackers register names models commonly invent |
| Missing validation | Insecure examples are abundant in training data |
| Deprecated patterns | Training data includes years of outdated tutorials |
| False test claims | "I tested this" — nothing was run |
What to do: verify APIs against docs, check packages exist in the registry before installing, run the code including edge cases, and read security-sensitive sections line by line. See AI for programming.
9. Bias and a flattened default perspective
The failure: output that treats one cultural, linguistic, or professional norm as neutral, or reproduces stereotypes from training text.
Test it: ask for a description of a "typical" professional in some field, or for advice on a culturally specific practice you know well. Look for unstated assumptions about location, language, or background.
Why: training data over-represents some languages, regions, and viewpoints. The model reproduces the statistical centre of that data, which is not a neutral position.
What to do: state your context explicitly — country, jurisdiction, audience, constraints. Be especially careful where norms vary: employment practice, education systems, medical guidance, etiquette, legal process.
What better prompting cannot fix
A useful mental division. Prompting helps a lot with one column and nothing with the other.
| Prompting can improve | Prompting cannot fix |
|---|---|
| Relevance and specificity | Whether a fact is true |
| Format, tone, structure | The knowledge cutoff |
| Reasoning quality on multi-step tasks | Fabrication risk |
| Which assumptions get surfaced | Absence of an uncertainty signal |
| Catching some errors, by asking for quotes | Arithmetic reliability |
Anyone selling a prompt that "eliminates hallucinations" is selling something that does not exist. Retrieval — giving the model real text to read — reduces fabrication meaningfully. Wording does not.
Tasks to keep away from an AI assistant
Based on the above, these should not be delegated without expert review:
- Final medical, legal, tax, or financial decisions
- Anything requiring current information without verified retrieval
- Calculations that matter, done in prose
- Citations for academic or legal work
- Safety-critical instructions (electrical, structural, chemical, dosage)
- Decisions about individuals — hiring, credit, discipline
- Anything you would publish unchecked under your own name
And the inverse — low-risk, high-value tasks: reshaping text you supplied, summarising documents you provide, generating options you will filter, explaining well-documented concepts, drafting boilerplate you will review, being a patient explainer while you learn.
The pattern: reliability tracks how much the model had to invent. Supply the facts and ask for a transformation.
How this applies to PhantomAI
Every limitation on this page applies to PhantomAI. It runs openai/gpt-oss-120b and openai/gpt-oss-20b through Groq — capable open-weight models, and subject to all nine failure modes above without exception.
What the product does about it is architectural, not magical: web search for current questions, and document/URL reading so answers can be grounded in your material rather than model recall. Both reduce fabrication by changing where facts come from. Neither adds a verification step, which is why we publish a product limitations page listing specifics and why fact-checking is a workflow we actively want users to have.
We would rather tell you where the tool fails than let you find out on something that matters.
Continue
- How to fact-check AI output — the verification workflow
- How AI assistants generate answers — the mechanism behind these failures
- AI privacy basics — what not to put into any AI tool
Simanta Pratim Das
Founder & Developer
Simanta is an independent AI engineer based in Guwahati, India, building PhantomAI as a solo project — designing the product, the interface, and the AI pipeline end to end.
Related guides
Try these techniques in PhantomAI
PhantomAI is free to use, and the workflows in this guide work best when you paste in your own material rather than relying on the model's memory. Before you start, it's worth reading what it can't do.