AI Runtime & Serving

Inference servers, model hubs, and API gateways for running or serving LLMs — local inference engines, hosted model catalogues, and aggregators that route across multiple providers.