Skip to main content
LLM applications have unique performance characteristics — the bottleneck is almost always LLM inference latency, not compute or I/O. That means standard optimization approaches (faster database queries, better indexing, CDN caching) barely move the needle. You need a different mental model: minimize the number of LLM calls, make the ones you do make as fast as possible, and hide latency from users through streaming.
The optimization hierarchy that consistently pays off: async > caching > model routing > prompt compression > inference optimization. Work in that order. Async and caching are high-impact and low-risk; inference tuning is low-impact unless you’re running at scale.

Measure First

Profile your pipeline before optimizing — every system has a different bottleneck, and guessing wrong wastes time. The key metrics to track for LLM applications are different from traditional services: Use LangSmith tracing or OpenTelemetry to get per-step timing automatically. For quick profiling during development, a timing decorator works well:
timing_decorator.py

Async Everything

Python async is essential for LLM applications — LLM calls are pure I/O wait. The Python process is doing nothing while waiting for the API response. With synchronous code, that idle time blocks every other operation. With async, you can run independent operations concurrently. Use ainvoke / astream instead of invoke / stream throughout your LangChain/LangGraph code. The async versions are almost always available and the migration is usually a one-line change. The real win comes from parallelizing independent operations:
parallel_async.py
asyncio.gather is your most powerful latency tool for retrieval-augmented pipelines. If you’re fetching conversation history, user profile data, and retrieved chunks, run all three in parallel. It turns a 600ms sequential operation into a 200ms parallel one.

Semantic Caching

Exact-match caching (caching on the verbatim query string) has very low hit rates for LLM applications because users phrase questions differently even when they mean the same thing. Semantic caching solves this by caching on meaning rather than exact text. The idea: embed each incoming query, find the nearest cached query by cosine similarity, and if the similarity exceeds a threshold (typically 0.95), return the cached response.
semantic_cache.py
Semantic caching is especially effective for:
  • FAQ-style queries — “What are your pricing plans?” phrased 20 different ways
  • High-traffic, low-variance queries — product search, support queries with common patterns
  • Expensive chains — multi-step chains where a cache hit saves several LLM calls
Cache hit rates of 30–50% are realistic for support/FAQ use cases. At that rate, you’re cutting LLM costs by a third and slashing average latency dramatically — cached responses return in <50ms vs 1–3 seconds for live inference.

Prompt Compression

Long prompts mean higher latency and higher cost. Every token in your prompt adds to the prefill time and the cost per call. Retrieved context is often the biggest contributor — chunking strategies frequently produce more text than the model actually needs. LLMLingua is the most effective tool I’ve found for this. It compresses retrieved context by up to 20x by removing tokens that are redundant or low-information, with minimal quality loss on downstream tasks.
prompt_compression.py
Manual compression strategies that work without a compression model:
  • Extract only relevant sentences from chunks using a small, fast model
  • Reranking reduces context size by keeping only top-3 to top-5 chunks rather than top-20
  • Summarize document-level context rather than chunking full documents

Model Routing by Complexity

Not all queries need GPT-4o. Using your most capable (and most expensive) model for every query is like using a sledgehammer to crack a nut. Simple factual lookups, yes/no questions, and classification tasks run equally well on smaller, faster, cheaper models. Here’s a lightweight routing implementation using a classifier prompt:
model_router.py

Streaming for Perceived Performance

Perceived latency is often more important than actual latency. A user waiting for a 3-second response with no feedback feels much longer than a 3-second response where text starts appearing in 300ms. Stream LLM output to the user immediately — don’t wait for the full response to arrive before displaying anything.
streaming_endpoint.py
Key principles for streaming UX:
  • First token latency matters more than total generation time — optimize your pipeline to minimize time from request to first token
  • Combine with a skeleton UI — render placeholder shapes before content arrives; users perceive this as faster
  • Stream processing steps too — if your pipeline has multiple steps, stream status updates (“Searching knowledge base…”, “Generating response…”) rather than showing a blank loading state

vLLM Throughput Tuning

For self-hosted inference at scale, vLLM is the standard. These are the configuration knobs that actually matter:
Controls how many requests vLLM processes in parallel. Higher values increase throughput but also increase memory pressure and per-request latency.Start with the default, then increase incrementally while monitoring GPU memory utilization and p99 latency. A good rule of thumb: increase max-num-seqs until GPU memory hits ~85%, then stop.
vllm_startup.sh
The KV cache is what enables efficient batching — it stores intermediate attention states so vLLM doesn’t recompute them on every token. Larger KV cache = more sequences can be batched = higher throughput.Set --gpu-memory-utilization 0.90 to give vLLM 90% of GPU memory for model weights and KV cache. The remaining 10% is headroom for other processes and to avoid OOM on unexpected memory spikes.
vllm_memory.sh
AWQ (Activation-aware Weight Quantization) reduces model memory by roughly 50% with minimal quality degradation — typically less than 1–2% on standard benchmarks. This lets you fit a larger model on the same GPU, or fit the same model while leaving more room for KV cache.Use pre-quantized AWQ models from HuggingFace rather than quantizing yourself:
vllm_awq.sh
AWQ is generally preferable to GPTQ for inference — similar compression ratio with better throughput characteristics.
For models that don’t fit on a single GPU (70B+ parameter models, or smaller models where you need very high KV cache), tensor parallelism splits the model across multiple GPUs.
vllm_tensor_parallel.sh
Tensor parallelism adds inter-GPU communication overhead. For smaller models that fit on one GPU, running multiple single-GPU instances behind a load balancer often gives better throughput than tensor parallelism across multiple GPUs.