A dedicated server for local language models: GPU acceleration (AMD MI25), substantial storage, and the model-serving toolchain. Tracked in part by its subprojects - llm-integration (client tooling and routers), amd-mi25-fan-service (GPU cooling), and model-router (inference routing).