{inference-perf} (JOSS) - A production-scale GenAI inference performance benchmarking tool that allows you to benchmark and analyze the performance of inference deployments. It is agnostic of model servers and can be used to measure performance and compare different systems.
Cost
Reasoning queries (∼5,000 output tokens) raise energy use ∼13× vs. standard queries (source)