AI Runtime & Serving

Engines and apps for running open-weight models on hardware you control, from a laptop to a GPU cluster, usually exposed through an OpenAI-compatible API.