1 / 9 · Purpose
Sharing a spending rate
Allocations control how quickly API usage can spend money, in USD/hour. They are not prepaid balances.
- Each request consumes its key's allowance and every applicable ancestor allowance.
- Idle allocations do not reserve bandwidth away from other users.
- A separate monthly ceiling can limit total spending.
Local proof of concept. Example prices and rates are illustrative.
2 / 9 · Hierarchy
A tree of shared limits
Company · $20/hour
└── Engineering · $10/hour
├── Morgan · $2/hour (personal)
│ └── Individual key
└── Platform · $5/hour
├── Alice · $2/hour (personal)
│ ├── Editor key · $2/hour
│ └── CLI key · $2/hour
├── Bob · $2/hour (personal)
│ └── Individual key
└── Agent key · $3/hour
The agent key consumes Platform, Engineering and Company capacity—not Morgan's personal allowance.
3 / 9 · Token bucket
Refill steadily; spend when needed
Think of a bucket containing a dollar allowance. It refills continuously at the configured rate and stops refilling when full.
- A request reserves estimated input plus bounded output cost before dispatch.
- If a required bucket cannot cover that reservation, the gateway returns a rate-limit error.
- Actual usage reconciles the estimate afterward; an underestimate can leave debt.
Here “token bucket” is the limiter algorithm. Its tokens represent dollars, not model text tokens.
4 / 9 · Bursts
Short bursts without a higher sustained rate
With a ratio of 0.25 hours, a $2/hour allocation has a $0.50 bucket. Fifteen idle minutes can refill an empty bucket.
- Saved allowance lets a short interactive request start immediately.
- A burst spends saved capacity; it does not increase the long-term rate.
- The administrator sets one ratio for everyone. Users do not size bursts.
Reservations larger than a bucket's full capacity cannot fit, even after waiting. Reduce the request size or ask the responsible manager about the allocation.
5 / 9 · Overallocation
Child rates may add up to more than their parent
Alice has $2/hour. Her Editor and CLI keys may each be set to $2/hour.
- Editor idle: CLI can use Alice's full available rate.
- Both active: each key's cap applies, then Alice's shared cap.
- The parent is charged for usage, not for the sum of configured child rates.
- Platform and the company still cap their combined descendants.
Admission is dynamic and request-based, not a fixed split. The PoC does not guarantee equal shares, weighted scheduling or Linux HTB borrowing priorities under sustained contention.
6 / 9 · Unlimited and model caps
Unlimited means no cap at this level
- An Unlimited key still consumes its user's or team's finite allocation.
- An Unlimited personal allocation still consumes organizational capacity.
- Zero blocks positive-cost requests. Zero is not Unlimited.
- All-model and per-model limits both apply. Switching keys or models cannot bypass shared caps.
- Monthly ceilings and upstream provider limits remain independent constraints.
The API represents Unlimited as JSON null, not a very large number.
7 / 9 · Responsibilities
Personal and organizational allocations are separate
| Allocation or key | Who controls it? |
|---|---|
| Personal allocation | A responsible manager; never the user themselves, including managers. |
| Organizational allocation | An ancestor manager, not the manager of that same organization. |
| Company root allocation | The explicitly designated administrator. |
| Individual keys | Their owner, within the inherited personal allocation. |
| Agent keys | Responsible organizational managers, within organizational allocations. |
Managers allocate rates to subordinate organizations and people. They also have their own, separate personal allocation.
8 / 9 · Personnel changes
Personal access ends; automation continues
- Deprovisioning disables the user's portal access and permanently revokes their individual keys.
- Agent keys belonging to that manager move up one organizational level; their secrets remain valid.
- A surviving ancestor manager takes responsibility. At the root, administrator stewardship is the terminal case.
- Historical usage remains attributed to the original request hierarchy.
SUSE SCIM notifications will supply provisioning events during integration. Exact API, authentication and hierarchy mappings await the integration documents; the local trusted hierarchy exercises lifecycle behavior now.
9 / 9 · Reading a rate-limit response
Find the scope that is limiting the request
- Check the reported key, personal or organizational scope.
- If allowance is temporarily exhausted, wait for the reported retry interval when present.
- If the request cannot fit at full capacity, reduce its maximum output or discuss the rate with the responsible manager.
- Check monthly budgets separately; a monthly ceiling does not refill hourly.
In-flight reservations and uncertain charges remain visible. Ambiguous provider failures retain the estimate rather than silently freeing capacity.
Enforcement is approximate at request boundaries. This presentation does not claim a measured 1% overshoot bound or production SCIM/provider integration.
Back to start