Tag
Cost-Optimization
2 posts filed under Cost-Optimization.
Cost Control for LLM Apps: Caching, Batching, Model Tiers
Cut your Azure OpenAI bill with model-tier routing, the four caches, output discipline, and batching. The levers that move the bill, in order of impact.
Streamlining AI Development with LiteLLM Proxy: A Comprehensive Guide
Running LiteLLM proxy with Open WebUI, Postgres, and Redis in Docker Compose — one OpenAI-compatible endpoint in front of OpenAI, Anthropic, and Ollama.

