Abstract
Inference Perf is a generative AI (GenAI) inference performance benchmarking tool aimed at benchmarking and analyzing the performance of inference deployments. It is designed to be model-server agnostic, allowing for apples-to-apples comparisons across different model servers and serving stacks. As part of the inference benchmarking and metrics standardization effort in the Kubernetes wg-serving, it seeks to standardize tooling and metrics for measuring inference performance across the Kubernetes and model server communities.