Performance benchmarking inference serving of vLLM.

Authors

  • Gunjan Shegade

Keywords:

LLM Inference Serving, vLLM, TensorRT-LLM, Triton Inference Server, PagedAttention, Continuous Batching, Throughput Benchmarking, Production Concurrency, Tail Latency.

Abstract

The rapid adoption of large language models (LLMs) in production systems hasshifted the central engineering bottleneck from model training to inference serving, wherethroughput, latency, and cost-efficiency under concurrent load determine the economic viability of a deployment. This paper presents a comparative benchmarking study of threewidely deployed LLM inference serving stacks

References

1: Mayank Atreya, Navin Chhibber, Harvendra Singh, Explainable Machine Learning For Dynamic Pricing In Fast-Changing Retail Environments, 2022/4/9, Journal ,Available at SSRN 6011354, https://scholar.google.com/citations?view_op=view_citation&hl=en&user=fyViF1UAAAAJ&citation_for_view=fyViF1UAAAAJ:LkGwnXOMwfcC

Downloads

Published

2024-06-20

How to Cite

Gunjan Shegade. (2024). Performance benchmarking inference serving of vLLM. Journal of Computational Analysis and Applications (JoCAAA), 33(06), 4184–4198. Retrieved from https://eudoxuspress.com/index.php/pub/article/view/5723

Issue

Section

Articles