Packages
{ inference-perf } (JOSS ) - A production-scale GenAI inference performance benchmarking tool that allows you to benchmark and analyze the performance of inference deployments. It is agnostic of model servers and can be used to measure performance and compare different systems.
Resources
Cost
LLMs are priced in terms of input and output tokens, with output generally charged at significantly higher rates
Reasoning queries (∼5,000 output tokens) raise energy use ∼13× vs. standard queries (source )
Hosted Model: input tokens × input-token price + output tokens × output-token price
Reasoning causes the model to generate many more internal/output tokens (e.g. 1000s). Then, the API provider can charge substantially more.
Track per-step accuracy separately from overall task completion.
When per-step accuracy drops from 90% to 87% on a 10-step task, overall success rate drops from 35% to 24%.
You want to catch that degradation in monitoring, not in a post-incident review
The Docker Model Runner (DMR) - A new Docker Desktop feature that enables the run of Large Language Models natively with Docker Desktop.
Follows the common Docker workflow:
Pull models from registries (e.g., Docker Hub)
Run models locally with GPU acceleration
Integrate models into the development workflows
Resources