The optimization hierarchy that consistently pays off: async > caching > model routing > prompt compression > inference optimization. Work in that order. Async and caching are high-impact and low-risk; inference tuning is low-impact unless you’re running at scale.
Measure First
Profile your pipeline before optimizing — every system has a different bottleneck, and guessing wrong wastes time. The key metrics to track for LLM applications are different from traditional services:
Use LangSmith tracing or OpenTelemetry to get per-step timing automatically. For quick profiling during development, a timing decorator works well:
timing_decorator.py
Async Everything
Python async is essential for LLM applications — LLM calls are pure I/O wait. The Python process is doing nothing while waiting for the API response. With synchronous code, that idle time blocks every other operation. With async, you can run independent operations concurrently. Useainvoke / astream instead of invoke / stream throughout your LangChain/LangGraph code. The async versions are almost always available and the migration is usually a one-line change.
The real win comes from parallelizing independent operations:
parallel_async.py
Semantic Caching
Exact-match caching (caching on the verbatim query string) has very low hit rates for LLM applications because users phrase questions differently even when they mean the same thing. Semantic caching solves this by caching on meaning rather than exact text. The idea: embed each incoming query, find the nearest cached query by cosine similarity, and if the similarity exceeds a threshold (typically 0.95), return the cached response.semantic_cache.py
- FAQ-style queries — “What are your pricing plans?” phrased 20 different ways
- High-traffic, low-variance queries — product search, support queries with common patterns
- Expensive chains — multi-step chains where a cache hit saves several LLM calls
Cache hit rates of 30–50% are realistic for support/FAQ use cases. At that
rate, you’re cutting LLM costs by a third and slashing average latency
dramatically — cached responses return in <50ms vs 1–3 seconds for live
inference.
Prompt Compression
Long prompts mean higher latency and higher cost. Every token in your prompt adds to the prefill time and the cost per call. Retrieved context is often the biggest contributor — chunking strategies frequently produce more text than the model actually needs. LLMLingua is the most effective tool I’ve found for this. It compresses retrieved context by up to 20x by removing tokens that are redundant or low-information, with minimal quality loss on downstream tasks.prompt_compression.py
- Extract only relevant sentences from chunks using a small, fast model
- Reranking reduces context size by keeping only top-3 to top-5 chunks rather than top-20
- Summarize document-level context rather than chunking full documents
Model Routing by Complexity
Not all queries need GPT-4o. Using your most capable (and most expensive) model for every query is like using a sledgehammer to crack a nut. Simple factual lookups, yes/no questions, and classification tasks run equally well on smaller, faster, cheaper models.
Here’s a lightweight routing implementation using a classifier prompt:
model_router.py
Streaming for Perceived Performance
Perceived latency is often more important than actual latency. A user waiting for a 3-second response with no feedback feels much longer than a 3-second response where text starts appearing in 300ms. Stream LLM output to the user immediately — don’t wait for the full response to arrive before displaying anything.streaming_endpoint.py
- First token latency matters more than total generation time — optimize your pipeline to minimize time from request to first token
- Combine with a skeleton UI — render placeholder shapes before content arrives; users perceive this as faster
- Stream processing steps too — if your pipeline has multiple steps, stream status updates (“Searching knowledge base…”, “Generating response…”) rather than showing a blank loading state
vLLM Throughput Tuning
For self-hosted inference at scale, vLLM is the standard. These are the configuration knobs that actually matter:--max-num-seqs: Maximum Concurrent Sequences
--max-num-seqs: Maximum Concurrent Sequences
Controls how many requests vLLM processes in parallel. Higher values increase
throughput but also increase memory pressure and per-request latency.Start with the default, then increase incrementally while monitoring GPU
memory utilization and p99 latency. A good rule of thumb: increase
max-num-seqs until GPU memory hits ~85%, then stop.vllm_startup.sh
--gpu-memory-utilization: Maximize KV Cache
--gpu-memory-utilization: Maximize KV Cache
The KV cache is what enables efficient batching — it stores intermediate
attention states so vLLM doesn’t recompute them on every token. Larger KV
cache = more sequences can be batched = higher throughput.Set
--gpu-memory-utilization 0.90 to give vLLM 90% of GPU memory for
model weights and KV cache. The remaining 10% is headroom for other processes
and to avoid OOM on unexpected memory spikes.vllm_memory.sh
Quantization: AWQ for Memory Efficiency
Quantization: AWQ for Memory Efficiency
AWQ (Activation-aware Weight Quantization) reduces model memory by roughly
50% with minimal quality degradation — typically less than 1–2% on standard
benchmarks. This lets you fit a larger model on the same GPU, or fit the same
model while leaving more room for KV cache.Use pre-quantized AWQ models from HuggingFace rather than quantizing yourself:AWQ is generally preferable to GPTQ for inference — similar compression ratio
with better throughput characteristics.
vllm_awq.sh
Tensor Parallelism: Multi-GPU Scaling
Tensor Parallelism: Multi-GPU Scaling
For models that don’t fit on a single GPU (70B+ parameter models, or smaller
models where you need very high KV cache), tensor parallelism splits the model
across multiple GPUs.Tensor parallelism adds inter-GPU communication overhead. For smaller models
that fit on one GPU, running multiple single-GPU instances behind a load
balancer often gives better throughput than tensor parallelism across multiple
GPUs.
vllm_tensor_parallel.sh