Skip to main content
Building with LLMs introduces a class of bugs that don’t exist in deterministic software — here are the gotchas that catch most engineers at least once. Some of these cost me hours to debug. A few cost real money. All of them are avoidable once you know to look for them.

The Gotcha Catalog

The problem: User input contains instructions that override or subvert your system prompt. The model treats user-supplied text as authoritative instructions rather than data to be processed.Classic example: A user sends:
Adversarial Input
The model, trained to be helpful and follow instructions, may comply — even with a well-crafted system prompt.Mitigations:
  • Separate untrusted content from instructions — never interpolate raw user input directly into instruction-level positions in your prompt
  • Use structured outputs — JSON mode makes it harder for injected instructions to affect response format
  • System-prompt hardening — explicitly instruct the model to ignore user requests to change its behavior or ignore instructions
  • Input sanitization — filter or flag known injection patterns before the prompt is constructed
injection_defense.py
Detection: Add adversarial injection attempts to your eval suite. Test regularly — model updates can change susceptibility.
The problem: Even with retrieved context in the prompt, LLMs generate plausible-sounding facts that aren’t present in the retrieved documents. The model fills gaps with its parametric knowledge — which may be wrong, outdated, or simply fabricated.Why it happens: Models are trained to produce fluent, confident-sounding text. They don’t have a built-in “I don’t know” signal — they interpolate from patterns in training data when the answer isn’t in context.Mitigations:
  • Explicit grounding instructions in every RAG prompt
  • Self-consistency checking (ask the model to verify its answer against the context)
  • Faithfulness scoring with RAGAs in your eval pipeline
  • Citation-based answers that can be spot-checked
Here’s the grounding prompt pattern I use in production:
grounding_prompt.py
The grounding prompt reduces hallucination but doesn’t eliminate it. Always run faithfulness scoring on a sample of production outputs, especially after model upgrades — newer models aren’t always more grounded.
The problem: LLMs perform significantly worse at recalling information from the middle of long context windows. Research has consistently shown that models attend most strongly to content at the beginning and end of their context — information buried in the middle gets lost.This is particularly damaging in RAG: if you retrieve 20 chunks and concatenate them, the most relevant chunks may end up in the middle where the model effectively ignores them.Mitigations:
  • Rerank before inserting into context — use a cross-encoder reranker (Cohere Rerank, BGE reranker) to find the truly most relevant chunks
  • Limit context to top-3 to top-5 chunks, not top-20 — less context, better retrieval, lower cost
  • Position matters — put your most important context at the beginning or end of the prompt, not the middle
reranking.py
The problem: Assuming character count approximates token count leads to silent context truncation. Tokens are not characters — a single word can be 1–4 tokens depending on its frequency in the training corpus. Code, URLs, and non-English text tokenize very differently from plain English prose.What actually happens: You construct a prompt, it exceeds the context limit, and the API either returns an error (if you’re lucky) or silently truncates the input (if you’re not). Your RAG context — carefully retrieved — gets chopped without warning.The fix: use tiktoken for exact token counting before every API call.
token_counting.py
Different models use different tokenizers. gpt-4o and gpt-4-turbo use cl100k_base. Always pass the actual model name to tiktoken.encoding_for_model() rather than hardcoding the encoding.
The problem: A common assumption is that temperature=0 gives identical outputs on repeated calls with identical inputs. It doesn’t — not reliably.Why: Floating-point arithmetic in GPU operations isn’t perfectly reproducible across batches, distributed inference nodes, or model versions. The differences are usually small, but they’re real.Why it matters: If you’re building parsing logic that relies on the model always responding in exactly the same format — e.g., always starting with “Answer:” or always returning valid JSON without JSON mode — you’ll get silent breakage when the format shifts by even one character.The fix: don’t rely on assumed exact output format. Use:
  • response_format={"type": "json_object"} for JSON outputs (OpenAI)
  • Pydantic output parsers with .with_structured_output() in LangChain
  • Robust parsing with fallbacks, never brittle string matching
structured_output.py
The problem: In LangChain’s ToolNode, when a tool raises an exception, the error is caught and returned as a message to the agent — but the agent may not recognize it as an error state and will continue confidently, often hallucinating a successful result.I’ve had agents proceed to “summarize the search results” after their search tool returned a stack trace. The agent treated the error message as if it were real data.The fix: test tools individually before wiring them into agents, implement retry logic with exponential backoff, and log every tool failure explicitly.
robust_tool.py
Add instructions to your system prompt telling the agent how to handle TOOL_ERROR responses. Without explicit guidance, the model will often treat error strings as valid data.
The problem: Hitting OpenAI or Anthropic rate limits in production causes cascading failures. A burst of requests triggers rate limiting, retries pile up, the queue grows, and suddenly your entire application is waiting on API responses.The layered fix:
  • Exponential backoff with jitter on every API call — the tenacity library makes this easy
  • A request queue with concurrency limits — cap parallel inflight requests to stay under TPM limits
  • Track rate limit headers — OpenAI returns x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens on every response
  • LiteLLM proxy for multi-provider setups — it handles rate limit routing and failover automatically
rate_limit_handling.py
The wait_random_exponential function adds jitter to prevent the “thundering herd” problem — all retrying clients hitting the API again at exactly the same moment.
The problem: In long agent conversations, early mistakes compound. A bad tool call in turn 3 produces an incorrect observation. The model incorporates that incorrect observation into its working understanding. By turn 8, it’s confidently reasoning from false premises that it introduced itself six turns ago.This is especially insidious because the conversation looks coherent — the model isn’t obviously confused, it’s just wrong in a self-consistent way.The fixes:
  • Detect and flag agent errors — when a tool returns an error, treat it as a recoverable failure and consider resetting agent state
  • LangGraph checkpointing — save state at known-good points and roll back when something goes wrong
checkpoint_recovery.py
  • Conversation reset for agent errors — sometimes the cleanest fix is to start fresh and re-summarize the goal
  • Scratchpad vs conversation history — separate the agent’s internal reasoning from the user-visible conversation so errors in reasoning don’t pollute the visible context