The Gotcha Catalog
Gotcha 1: Prompt Injection
Gotcha 1: Prompt Injection
The problem: User input contains instructions that override or subvert your system prompt. The model treats user-supplied text as authoritative instructions rather than data to be processed.Classic example: A user sends:The model, trained to be helpful and follow instructions, may comply — even with a well-crafted system prompt.Mitigations:Detection: Add adversarial injection attempts to your eval suite. Test regularly — model updates can change susceptibility.
Adversarial Input
- Separate untrusted content from instructions — never interpolate raw user input directly into instruction-level positions in your prompt
- Use structured outputs — JSON mode makes it harder for injected instructions to affect response format
- System-prompt hardening — explicitly instruct the model to ignore user requests to change its behavior or ignore instructions
- Input sanitization — filter or flag known injection patterns before the prompt is constructed
injection_defense.py
Gotcha 2: Hallucination in RAG
Gotcha 2: Hallucination in RAG
The problem: Even with retrieved context in the prompt, LLMs generate plausible-sounding facts that aren’t present in the retrieved documents. The model fills gaps with its parametric knowledge — which may be wrong, outdated, or simply fabricated.Why it happens: Models are trained to produce fluent, confident-sounding text. They don’t have a built-in “I don’t know” signal — they interpolate from patterns in training data when the answer isn’t in context.Mitigations:
- Explicit grounding instructions in every RAG prompt
- Self-consistency checking (ask the model to verify its answer against the context)
- Faithfulness scoring with RAGAs in your eval pipeline
- Citation-based answers that can be spot-checked
grounding_prompt.py
Gotcha 3: The "Lost in the Middle" Problem
Gotcha 3: The "Lost in the Middle" Problem
The problem: LLMs perform significantly worse at recalling information from the middle of long context windows. Research has consistently shown that models attend most strongly to content at the beginning and end of their context — information buried in the middle gets lost.This is particularly damaging in RAG: if you retrieve 20 chunks and concatenate them, the most relevant chunks may end up in the middle where the model effectively ignores them.Mitigations:
- Rerank before inserting into context — use a cross-encoder reranker (Cohere Rerank, BGE reranker) to find the truly most relevant chunks
- Limit context to top-3 to top-5 chunks, not top-20 — less context, better retrieval, lower cost
- Position matters — put your most important context at the beginning or end of the prompt, not the middle
reranking.py
Gotcha 4: Token Counting Surprises
Gotcha 4: Token Counting Surprises
The problem: Assuming character count approximates token count leads to silent context truncation. Tokens are not characters — a single word can be 1–4 tokens depending on its frequency in the training corpus. Code, URLs, and non-English text tokenize very differently from plain English prose.What actually happens: You construct a prompt, it exceeds the context limit, and the API either returns an error (if you’re lucky) or silently truncates the input (if you’re not). Your RAG context — carefully retrieved — gets chopped without warning.The fix: use
tiktoken for exact token counting before every API call.token_counting.py
Gotcha 5: Temperature=0 Is Not Deterministic
Gotcha 5: Temperature=0 Is Not Deterministic
The problem: A common assumption is that
temperature=0 gives identical outputs on repeated calls with identical inputs. It doesn’t — not reliably.Why: Floating-point arithmetic in GPU operations isn’t perfectly reproducible across batches, distributed inference nodes, or model versions. The differences are usually small, but they’re real.Why it matters: If you’re building parsing logic that relies on the model always responding in exactly the same format — e.g., always starting with “Answer:” or always returning valid JSON without JSON mode — you’ll get silent breakage when the format shifts by even one character.The fix: don’t rely on assumed exact output format. Use:response_format={"type": "json_object"}for JSON outputs (OpenAI)- Pydantic output parsers with
.with_structured_output()in LangChain - Robust parsing with fallbacks, never brittle string matching
structured_output.py
Gotcha 6: Tool Call Errors Are Silent by Default
Gotcha 6: Tool Call Errors Are Silent by Default
The problem: In LangChain’s
ToolNode, when a tool raises an exception, the error is caught and returned as a message to the agent — but the agent may not recognize it as an error state and will continue confidently, often hallucinating a successful result.I’ve had agents proceed to “summarize the search results” after their search tool returned a stack trace. The agent treated the error message as if it were real data.The fix: test tools individually before wiring them into agents, implement retry logic with exponential backoff, and log every tool failure explicitly.robust_tool.py
Gotcha 7: Rate Limits and Exponential Backoff
Gotcha 7: Rate Limits and Exponential Backoff
The problem: Hitting OpenAI or Anthropic rate limits in production causes cascading failures. A burst of requests triggers rate limiting, retries pile up, the queue grows, and suddenly your entire application is waiting on API responses.The layered fix:
- Exponential backoff with jitter on every API call — the
tenacitylibrary makes this easy - A request queue with concurrency limits — cap parallel inflight requests to stay under TPM limits
- Track rate limit headers — OpenAI returns
x-ratelimit-remaining-requestsandx-ratelimit-remaining-tokenson every response - LiteLLM proxy for multi-provider setups — it handles rate limit routing and failover automatically
rate_limit_handling.py
The
wait_random_exponential function adds jitter to prevent the
“thundering herd” problem — all retrying clients hitting the API again at
exactly the same moment.Gotcha 8: Context Pollution in Multi-Turn Conversations
Gotcha 8: Context Pollution in Multi-Turn Conversations
The problem: In long agent conversations, early mistakes compound. A bad tool call in turn 3 produces an incorrect observation. The model incorporates that incorrect observation into its working understanding. By turn 8, it’s confidently reasoning from false premises that it introduced itself six turns ago.This is especially insidious because the conversation looks coherent — the model isn’t obviously confused, it’s just wrong in a self-consistent way.The fixes:
- Detect and flag agent errors — when a tool returns an error, treat it as a recoverable failure and consider resetting agent state
- LangGraph checkpointing — save state at known-good points and roll back when something goes wrong
checkpoint_recovery.py
- Conversation reset for agent errors — sometimes the cleanest fix is to start fresh and re-summarize the goal
- Scratchpad vs conversation history — separate the agent’s internal reasoning from the user-visible conversation so errors in reasoning don’t pollute the visible context