Performance benchmarking inference serving of vLLM.
Keywords:
LLM Inference Serving, vLLM, TensorRT-LLM, Triton Inference Server, PagedAttention, Continuous Batching, Throughput Benchmarking, Production Concurrency, Tail Latency.Abstract
The rapid adoption of large language models (LLMs) in production systems hasshifted the central engineering bottleneck from model training to inference serving, wherethroughput, latency, and cost-efficiency under concurrent load determine the economic viability of a deployment. This paper presents a comparative benchmarking study of threewidely deployed LLM inference serving stacks
References
1: Mayank Atreya, Navin Chhibber, Harvendra Singh, Explainable Machine Learning For Dynamic Pricing In Fast-Changing Retail Environments, 2022/4/9, Journal ,Available at SSRN 6011354, https://scholar.google.com/citations?view_op=view_citation&hl=en&user=fyViF1UAAAAJ&citation_for_view=fyViF1UAAAAJ:LkGwnXOMwfcC


