Building a Retrieval-Augmented Generation (RAG) prototype in a Jupyter notebook is straightforward. Scaling it to handle thousands of concurrent production users is a completely different engineering challenge.
In previous posts, we focused heavily on optimizing the retrieval layer—from vector indexing strategies to unified Graph-RAG architectures in Postgres. However, once your data pipeline is optimized, the operational bottleneck inevitably shifts to the inference layer.
Relying on external LLM APIs introduces vendor lock-in, unpredictable latencies, and escalating costs. Conversely, wrapping an open-source embedding model or LLM in a standard FastAPI service with Hugging Face Transformers creates severe performance bottlenecks: single-threaded GIL constraints, memory fragmentation, and poor GPU kernel saturation.
To achieve enterprise-grade throughput and sub-hundred-millisecond p99 latencies, you need a dedicated, C++-backed inference engine. NVIDIA Triton Inference Server is the standard industry solution for this problem.
The Architectural Bottleneck: Why Native Python Serving Fails
When multiple asynchronous requests hit a standard Python web server hosting an AI model, several issues emerge:
- Unbatched GPU Execution: Each incoming HTTP request triggers its own forward pass on the GPU. Modern Tensor Cores require large, aligned matrices to reach peak TFLOPS; processing requests one-by-one leaves over 80% of GPU compute capacity idle.
- Memory Fragmentation: Allocating dynamic tensors per request leads to VRAM fragmentation and frequent, costly Out-Of-Memory (OOM) crashes under peak load.
- Application-Level Queuing: Queueing requests in Python application code adds latency spikes due to event-loop blocking and thread context switching.
Triton addresses these issues by moving the serving layer into C++, decoupling model execution from the API web worker, and managing GPU hardware scheduling natively.
Target Architecture: Unified Inference Hub
In a production Agentic RAG system, Triton acts as a centralized inference microservice serving both your vector embedding models (e.g., bge-large-en-v1.5 compiled to ONNX) and your generative models (e.g., Llama-3 via TensorRT-LLM or vLLM backends).
Communication between your application backends (e.g., FastAPI, Pydantic AI agents) and Triton occurs over gRPC using HTTP/2 protocol buffers, eliminating HTTP/1.1 JSON serialization overhead.

Dynamic Batching Configuration
Triton’s primary throughput driver is Dynamic Batching. Instead of executing model inference immediately, Triton holds incoming requests in a lock-free queue for a configurable microsecond window (max_queue_delay_microseconds), concatenates them into a single matrix, and executes one highly efficient GPU kernel pass.
Here is an abbreviated production config.pbtxt for an ONNX embedding model repository:
name: “bge_large_onnx”
platform: “onnxruntime_onnx”
max_batch_size: 64input [
{ name: “input_ids”, data_type: TYPE_INT64, dims: [ -1 ] },
{ name: “attention_mask”, data_type: TYPE_INT64, dims: [ -1 ] }
]
output [
{ name: “last_hidden_state”, data_type: TYPE_FP32, dims: [ -1, 1024 ] }
]dynamic_batching {
max_queue_delay_microseconds: 5000
preferred_batch_size: [ 8, 16, 32, 64 ]
}
Production Integration: Async gRPC Client
The following Python implementation demonstrates how to integrate Triton into an asynchronous RAG service layer using gRPC:
import numpy as np
import tritonclient.grpc.aio as grpcclient
from transformers import AutoTokenizer
class TritonEmbeddingClient:
def init(self, triton_url=”localhost:8001″, model=”bge_large_onnx”):
self.url, self.model = triton_url, model
self.tokenizer = AutoTokenizer.from_pretrained(“BAAI/bge-large-en-v1.5”)
async def embed(self, texts: list[str]) -> np.ndarray:
client = grpcclient.InferenceServerClient(url=self.url)
encoded = self.tokenizer(texts, padding=True, truncation=True, return_tensors="np")
inputs = [
grpcclient.InferInput("input_ids", encoded["input_ids"].shape, "INT64"),
grpcclient.InferInput("attention_mask", encoded["attention_mask"].shape, "INT64")
]
inputs[0].set_data_from_numpy(encoded["input_ids"].astype(np.int64))
inputs[1].set_data_from_numpy(encoded["attention_mask"].astype(np.int64))
res = await client.infer(model_name=self.model, inputs=inputs)
embeddings = res.as_numpy("last_hidden_state")[:, 0, :] # CLS token pooling
await client.close()
return embeddings / np.linalg.norm(embeddings, axis=1, keepdims=True)
Load Testing & Benchmarking
To measure performance under load, use this compact async benchmark script (benchmark_triton.py):
import asyncio, time, numpy as np
import tritonclient.grpc.aio as grpcclient
async def send_req(client, dummy_id, dummy_mask):
inputs = [
grpcclient.InferInput(“input_ids”, dummy_id.shape, “INT64”),
grpcclient.InferInput(“attention_mask”, dummy_mask.shape, “INT64”)
]
inputs[0].set_data_from_numpy(dummy_id)
inputs[1].set_data_from_numpy(dummy_mask)
start = time.perf_counter()
await client.infer(“bge_large_onnx”, inputs=inputs)
return time.perf_counter() – start
async def benchmark(total_reqs=1000, concurrency=32):
client = grpcclient.InferenceServerClient(“localhost:8001”)
ids, mask = np.ones((1, 128), dtype=np.int64), np.ones((1, 128), dtype=np.int64)
sem = asyncio.Semaphore(concurrency)
async def worker():
async with sem:
return await send_req(client, ids, mask)
t0 = time.perf_counter()
latencies = await asyncio.gather(*[worker() for _ in range(total_reqs)])
total_time = time.perf_counter() - t0
print(f"Throughput: {total_reqs / total_time:.2f} req/s")
print(f"p95 Latency: {np.percentile(latencies, 95) * 1000:.2f} ms")
await client.close()
if name == “main“:
asyncio.run(benchmark())
Impact: What Improvements Can You Expect?
When migrating from a naive Python application (such as FastAPI wrapped around PyTorch/Hugging Face) to Triton with ONNX and Dynamic Batching under heavy concurrent traffic:
- 10x+ Increase in Throughput: Dynamic batching merges individual async requests into packed tensor operations, drastically boosting requests per second without adding extra hardware.
- Over 90% Reduction in p99 Tail Latency: Decoupling web workers from model execution eliminates application-level queueing, keeping tail latency tightly bounded under traffic spikes.
- Predictable VRAM Memory Footprint: Allocating memory during server warm-up eliminates GPU memory fragmentation and runtime Out-Of-Memory (OOM) failures.
Operational Best Practices
- Decouple Pre-Processing: Tokenization and text cleaning should occur in your application layer (or a dedicated CPU container) to avoid burning valuable GPU cycles on string manipulations.
- Use gRPC Health Checks: In Kubernetes deployments, leverage Triton’s native health endpoints (
/v2/health/readyand/v2/health/live) for Liveness/Readiness probes. - Model Versioning: Maintain strict model artifact management in Object Storage (S3/GCS). Triton can dynamically load or unload new model versions without restarting the server instance.