also quotas, 429
Capping how many requests a user, team or app can make in a period, to protect capacity and spend.
You have done this if
You set per-team token limits on the model gateway.
Say it in a review
Every team has a token budget per minute at the gateway, and we return 429 with a retry-after when they hit it.
On the AI Application map LLM Gateway, API Gateway
Read Your AI feature has unit economics whether you measured them or not