A developer building locally or running agents needs a GPU that can host open large models, and wants the cost fixed before starting.
Renting cloud GPU instances directly, or using token-metered hosted APIs.
Metered cloud GPU costs are unpredictable, while local machines cannot host large models, making experimentation expensive.