OpenSLA
Most tools tell you after a promise is broken. OpenSLA stops you making promises you cannot keep. Every plan is an agreement the platform enforces: limits come from the agreement itself rather than being configured by hand, capacity is checked before you sell it, one customer cannot slow everyone else, and you are warned as traffic nears its limits — before anything is refused. What cannot be prevented by refusing a request — latency — is checked instead, and every refusal and breach arrives already attributed to the customer and what they were promised.
Almost everywhere, an API’s service level agreement is a document. It lives in a wiki or a contract annexe, the gateway configuration that is supposed to implement it lives somewhere else, and the two begin drifting apart the day they are written. Nobody can tell you when they diverged, because nothing compares them.
The numbers in that document are usually not real either. They come from a performance test run against the component on its own — no ingress, no TLS termination, no token validation, no policy chain, no network hop to the service. A consumer’s request traverses all of that, so the published figure describes a path nobody actually takes. It is optimistic by construction, which is why latency and throughput commitments are treated as soft and quietly dropped the first time they are inconvenient.
OpenSLA is a vendor-neutral specification that removes the gap. The document you publish is the thing the gateway enforces — same artefact, versioned with the API, deployed with it — and it is validated against what the whole chain actually delivers rather than against a benchmark of one part of it.
Service Levels From Evidence
Section titled “Service Levels From Evidence”The numbers in a service level should describe what your API really delivers, not what one part of it managed on a good day.
- Coming — for you, as the producer: Apiway will recommend the numbers for your plans from two kinds of evidence — Assurance runs through the whole path your customers’ calls take, and the traffic your API actually serves. You accept them or adjust them; what you publish is then what you can keep.
- For your customers: before they subscribe, Apiway recommends the smallest plan that covers the throughput and daily usage they expect, so nobody buys a plan that will refuse them, or pays for one far larger than they need.
What an OpenSLA Document Contains
Section titled “What an OpenSLA Document Contains”| Section | What It Specifies |
|---|---|
| Metadata | API name, version, provider details |
| Guarantees | P95/P99 latency targets, availability percentage |
| Protection | The API’s own ceiling — limit, window, and whether it is counted per source, per subnet or globally |
| Quotas | Daily/monthly/quarterly RU limits |
| Rate limits | Requests per minute (RPM), transactions per second (TPS) |
| Cost | RU cost rate per operation, currency |
Two Halves, and They Are Not the Same Thing
Section titled “Two Halves, and They Are Not the Same Thing”An OpenSLA document declares two things that are easy to confuse and must not be:
| Protection | Limits | |
|---|---|---|
| Whose interest | The producer’s | The consumer’s |
| What it is | The ceiling the API can serve | The terms a consumer bought |
| Who it applies to | All traffic | One subscription |
| Partitioned by | ip, subnet, or global | The consumer |
Limits are commercial. They are what a consumer agreed to and paid for, and every consumer’s limits are a bounded subset of the ceiling — a consumer cannot be sold more than the API can serve.
Protection is infrastructure. It is the producer’s ceiling on their own surface, and it exists because rate limits alone do not keep an API standing. A thousand consumers each behaving perfectly inside their own limits can still exhaust it. Only a ceiling that applies to all traffic prevents that.
Protection also covers the traffic that has no consumer yet. A rate limit is per subscription, so it cannot apply to a request arriving before any identity exists — a token request, a discovery document. Those endpoints are reached by everyone, including anyone probing them, and the ceiling is what protects them.
The Ceiling Outranks the Entitlement
Section titled “The Ceiling Outranks the Entitlement”Because they answer to different interests, the two are enforced in a fixed order: the ceiling first, the consumer’s limits second. A consumer comfortably inside their own rate limit is still refused once the API’s quota is exhausted.
This is deliberate and it is not configurable. A consumer’s commercial terms cannot switch off the producer’s protection of their own surface — otherwise the ceiling would be worth only as much as the most generous contract anyone had signed.
Both Sides of the Coin
Section titled “Both Sides of the Coin”Declaring the ceiling and the sold terms in the same machine-readable artefact makes something possible that is simply unavailable to anyone whose SLA is prose: the arithmetic.
The platform knows what the API can serve and the sum of what has been promised to consumers, so it can answer questions that are otherwise guesswork:
- How much headroom remains on this API right now
- Whether the next consumer’s requested tier fits inside it
- Whether an upgrade a consumer is asking for can actually be granted
- How close the current subscription base is to the ceiling
Subscription and upgrade flows check headroom before granting, rather than discovering the answer when the API falls over. Onboarding becomes a check instead of a hope, and the producing team learns that capacity is running out before a customer tells them.
Tiered Pricing
Section titled “Tiered Pricing”Producers package limits into ordered tiers, which is what a consumer actually chooses between. An illustrative shape:
| Tier | Rate Limit | RU Quota | Latency P95 | Price |
|---|---|---|---|---|
| Free | 10 RPM | 1,000 RU/month | Best effort | Free |
| Standard | 100 RPM | 50,000 RU/month | 200ms | Per-unit |
| Premium | 1,000 RPM | 500,000 RU/month | 100ms | Per-unit |
| Enterprise | Custom | Unlimited | 50ms | Custom |
Tiers are ordered, so the platform knows which tier is above the current one and can recommend an upgrade. Every tier remains a subset of the ceiling — tiers are how capacity is sold, not a way of exceeding it.
How Tiers Work
Section titled “How Tiers Work”For Producers
Section titled “For Producers”- Define your tiers when creating or updating an API
- Set per-tier rate limits, RU quotas, latency guarantees, and pricing
- Each tier maps to a specific set of gateway enforcement rules
For Consumers
Section titled “For Consumers”- Browse available tiers in the developer portal or marketplace
- Select a tier manually, or let Apiway recommend one based on estimated usage
- The selected tier drives the subscription’s rate limits, RU quota, and cost model
Upgrade Flow
Section titled “Upgrade Flow”When a consumer consistently hits their tier’s limits:
- After 3+ budget/rate limit breaches in 30 days →
ConsumerSlaUpgradeRecommendedgovernance event - The event includes the recommended next tier
- Consumer can one-click upgrade → starts a governance flow
- Capacity planning validates the new tier has headroom against the ceiling
- New limits take effect after governance approval
Gateway Enforcement
Section titled “Gateway Enforcement”The specification is not documentation that something else implements. It is what runs:
| SLA Element | Gateway Enforcement |
|---|---|
| Protection (quota) | Applied first, before any customer’s own limits — 429 with Retry-After on breach |
| Rate limit (RPM) | Per customer, per minute — 429 on breach |
| RU quota | Tracked per customer — 402 when exhausted |
| TPS (spike arrest) | Per-second limit on bursts |
| Latency | Measured through the deployed path; Assurance validates against targets |
Response headers expose the current state to consumers:
RateLimit-Limit: 100RateLimit-Remaining: 73X-RU-Limit: 50000X-RU-Remaining: 41200X-RU-Cost: 1Latency is the one guarantee that cannot be enforced by refusing a request, so it is checked instead of capped. Assurance runs real cases against the deployed API and compares observed P95 and P99 against the committed targets — through the ingress, the authentication, the policy chain and the network, because that is the path a consumer’s request takes. A target that the real deployment cannot meet is visible as a failing check rather than as a complaint.
Keeping every customer’s promise visible
Section titled “Keeping every customer’s promise visible”Why it matters. A promise to a customer is only worth what you can show you kept — and the customer should never be the one who tells you that you did not.
How. Throughput promises are enforced on every call, so they cannot be broken quietly. Latency and availability cannot be enforced by refusing a request, so they are checked instead.
What you have today, and what is coming — stated plainly, so you can decide whether the gap matters for you:
- Per-customer latency and availability on live traffic are not measured yet. Today limits and quotas are enforced on every call, attributed to the customer, and latency and availability are checked by Assurance after every deployment and on a schedule. A slowdown that affects one customer’s live traffic between those checks is not yet raised against that customer’s commitment.
- Until it is, cover it with the monitoring you already run. Every refusal and breach arrives already attributed to a named customer through the live event stream, and capacity and tier events by webhook (Events to your on-call). Point your existing latency monitoring at the same customer identities and the two line up without mapping anything.
Consumer-Side SLA Snapshot
Section titled “Consumer-Side SLA Snapshot”When a consumer subscribes across tenants, the producer’s OpenSLA is not shared live — it is snapshotted on the consumer tenant at the moment the producer approves the subscription. The snapshot is the frozen contract; subsequent producer tier edits do not retroactively alter a consumer’s agreed terms.
| Property on the snapshot | Value |
|---|---|
ApiId / SubscriptionId | Stamped to the consumer-side API and subscription |
FunctionallyValid | true — producer approval stands in for validation |
Distribute | true — readers filtering on “currently applied SLA” find it |
| Tier variant | Filtered to the single variant the consumer was granted |
Readers that need the currently-applied SLA for a subscription query
functionallyValid=true && distribute=true and sort by most recent — the same pattern used for
producer-side catalogue promotion. Superseded snapshots keep Active=false so historical versions
are preserved for audit without polluting current-state queries.
Wealth Engine Integration
Section titled “Wealth Engine Integration”OpenSLA tiers are the foundation of API economics:
- Each tier has a defined RU cost rate
- Revenue is tracked per tier — producers see which tiers generate the most revenue
- The tier structure IS the monetisation plan
- The Wealth Engine reports revenue by tier across the entire API portfolio
Warnings reach your on-call rather than waiting on a dashboard: capacity and tier events by webhook (Events to your on-call), limit breaches through the risk stream (Risk Management).
See also: Assurance · Metering · Cost Control · Subscribing