Developers deploying open models on local or on-premise machines need to get models running and expose callable inference endpoints.
Developers hand-roll inference services with scripts or containers, or call cloud model APIs directly.
Manually configuring inference environments, model downloads and memory management is tedious, and local versus cloud usage is hard to view in one place.