Self-hosted LLM inference platform
Internal
Designed and optimised a self-hosted ~27B-parameter LLM inference platform supporting multi-user workloads.
Model serving, quantization and precision trade-offs, GPU memory, request queues and concurrency, with monitoring and load testing under simulated agent workloads.
Stack: Python, Docker, GPU inference servers, Linux
