Skip to content

OpenSLA

Most tools tell you after a promise is broken. OpenSLA stops you making promises you cannot keep. Every plan is an agreement the platform enforces: limits come from the agreement itself rather than being configured by hand, capacity is checked before you sell it, one customer cannot slow everyone else, and you are warned as traffic nears its limits — before anything is refused. What cannot be prevented by refusing a request — latency — is checked instead, and every refusal and breach arrives already attributed to the customer and what they were promised.

Almost everywhere, an API’s service level agreement is a document. It lives in a wiki or a contract annexe, the gateway configuration that is supposed to implement it lives somewhere else, and the two begin drifting apart the day they are written. Nobody can tell you when they diverged, because nothing compares them.

The numbers in that document are usually not real either. They come from a performance test run against the component on its own — no ingress, no TLS termination, no token validation, no policy chain, no network hop to the service. A consumer’s request traverses all of that, so the published figure describes a path nobody actually takes. It is optimistic by construction, which is why latency and throughput commitments are treated as soft and quietly dropped the first time they are inconvenient.

OpenSLA is a vendor-neutral specification that removes the gap. The document you publish is the thing the gateway enforces — same artefact, versioned with the API, deployed with it — and it is validated against what the whole chain actually delivers rather than against a benchmark of one part of it.

The numbers in a service level should describe what your API really delivers, not what one part of it managed on a good day.

  • Coming — for you, as the producer: Apiway will recommend the numbers for your plans from two kinds of evidence — Assurance runs through the whole path your customers’ calls take, and the traffic your API actually serves. You accept them or adjust them; what you publish is then what you can keep.
  • For your customers: before they subscribe, Apiway recommends the smallest plan that covers the throughput and daily usage they expect, so nobody buys a plan that will refuse them, or pays for one far larger than they need.
SectionWhat It Specifies
MetadataAPI name, version, provider details
GuaranteesP95/P99 latency targets, availability percentage
ProtectionThe API’s own ceiling — limit, window, and whether it is counted per source, per subnet or globally
QuotasDaily/monthly/quarterly RU limits
Rate limitsRequests per minute (RPM), transactions per second (TPS)
CostRU cost rate per operation, currency

Two Halves, and They Are Not the Same Thing

Section titled “Two Halves, and They Are Not the Same Thing”

An OpenSLA document declares two things that are easy to confuse and must not be:

ProtectionLimits
Whose interestThe producer’sThe consumer’s
What it isThe ceiling the API can serveThe terms a consumer bought
Who it applies toAll trafficOne subscription
Partitioned byip, subnet, or globalThe consumer

Limits are commercial. They are what a consumer agreed to and paid for, and every consumer’s limits are a bounded subset of the ceiling — a consumer cannot be sold more than the API can serve.

Protection is infrastructure. It is the producer’s ceiling on their own surface, and it exists because rate limits alone do not keep an API standing. A thousand consumers each behaving perfectly inside their own limits can still exhaust it. Only a ceiling that applies to all traffic prevents that.

Protection also covers the traffic that has no consumer yet. A rate limit is per subscription, so it cannot apply to a request arriving before any identity exists — a token request, a discovery document. Those endpoints are reached by everyone, including anyone probing them, and the ceiling is what protects them.

Because they answer to different interests, the two are enforced in a fixed order: the ceiling first, the consumer’s limits second. A consumer comfortably inside their own rate limit is still refused once the API’s quota is exhausted.

This is deliberate and it is not configurable. A consumer’s commercial terms cannot switch off the producer’s protection of their own surface — otherwise the ceiling would be worth only as much as the most generous contract anyone had signed.

Declaring the ceiling and the sold terms in the same machine-readable artefact makes something possible that is simply unavailable to anyone whose SLA is prose: the arithmetic.

The platform knows what the API can serve and the sum of what has been promised to consumers, so it can answer questions that are otherwise guesswork:

  • How much headroom remains on this API right now
  • Whether the next consumer’s requested tier fits inside it
  • Whether an upgrade a consumer is asking for can actually be granted
  • How close the current subscription base is to the ceiling

Subscription and upgrade flows check headroom before granting, rather than discovering the answer when the API falls over. Onboarding becomes a check instead of a hope, and the producing team learns that capacity is running out before a customer tells them.

Producers package limits into ordered tiers, which is what a consumer actually chooses between. An illustrative shape:

TierRate LimitRU QuotaLatency P95Price
Free10 RPM1,000 RU/monthBest effortFree
Standard100 RPM50,000 RU/month200msPer-unit
Premium1,000 RPM500,000 RU/month100msPer-unit
EnterpriseCustomUnlimited50msCustom

Tiers are ordered, so the platform knows which tier is above the current one and can recommend an upgrade. Every tier remains a subset of the ceiling — tiers are how capacity is sold, not a way of exceeding it.

  1. Define your tiers when creating or updating an API
  2. Set per-tier rate limits, RU quotas, latency guarantees, and pricing
  3. Each tier maps to a specific set of gateway enforcement rules
  1. Browse available tiers in the developer portal or marketplace
  2. Select a tier manually, or let Apiway recommend one based on estimated usage
  3. The selected tier drives the subscription’s rate limits, RU quota, and cost model

When a consumer consistently hits their tier’s limits:

  1. After 3+ budget/rate limit breaches in 30 days → ConsumerSlaUpgradeRecommended governance event
  2. The event includes the recommended next tier
  3. Consumer can one-click upgrade → starts a governance flow
  4. Capacity planning validates the new tier has headroom against the ceiling
  5. New limits take effect after governance approval

The specification is not documentation that something else implements. It is what runs:

SLA ElementGateway Enforcement
Protection (quota)Applied first, before any customer’s own limits — 429 with Retry-After on breach
Rate limit (RPM)Per customer, per minute — 429 on breach
RU quotaTracked per customer — 402 when exhausted
TPS (spike arrest)Per-second limit on bursts
LatencyMeasured through the deployed path; Assurance validates against targets

Response headers expose the current state to consumers:

RateLimit-Limit: 100
RateLimit-Remaining: 73
X-RU-Limit: 50000
X-RU-Remaining: 41200
X-RU-Cost: 1

Latency is the one guarantee that cannot be enforced by refusing a request, so it is checked instead of capped. Assurance runs real cases against the deployed API and compares observed P95 and P99 against the committed targets — through the ingress, the authentication, the policy chain and the network, because that is the path a consumer’s request takes. A target that the real deployment cannot meet is visible as a failing check rather than as a complaint.

Keeping every customer’s promise visible

Section titled “Keeping every customer’s promise visible”

Why it matters. A promise to a customer is only worth what you can show you kept — and the customer should never be the one who tells you that you did not.

How. Throughput promises are enforced on every call, so they cannot be broken quietly. Latency and availability cannot be enforced by refusing a request, so they are checked instead.

What you have today, and what is coming — stated plainly, so you can decide whether the gap matters for you:

  • Per-customer latency and availability on live traffic are not measured yet. Today limits and quotas are enforced on every call, attributed to the customer, and latency and availability are checked by Assurance after every deployment and on a schedule. A slowdown that affects one customer’s live traffic between those checks is not yet raised against that customer’s commitment.
  • Until it is, cover it with the monitoring you already run. Every refusal and breach arrives already attributed to a named customer through the live event stream, and capacity and tier events by webhook (Events to your on-call). Point your existing latency monitoring at the same customer identities and the two line up without mapping anything.

When a consumer subscribes across tenants, the producer’s OpenSLA is not shared live — it is snapshotted on the consumer tenant at the moment the producer approves the subscription. The snapshot is the frozen contract; subsequent producer tier edits do not retroactively alter a consumer’s agreed terms.

Property on the snapshotValue
ApiId / SubscriptionIdStamped to the consumer-side API and subscription
FunctionallyValidtrue — producer approval stands in for validation
Distributetrue — readers filtering on “currently applied SLA” find it
Tier variantFiltered to the single variant the consumer was granted

Readers that need the currently-applied SLA for a subscription query functionallyValid=true && distribute=true and sort by most recent — the same pattern used for producer-side catalogue promotion. Superseded snapshots keep Active=false so historical versions are preserved for audit without polluting current-state queries.

OpenSLA tiers are the foundation of API economics:

  • Each tier has a defined RU cost rate
  • Revenue is tracked per tier — producers see which tiers generate the most revenue
  • The tier structure IS the monetisation plan
  • The Wealth Engine reports revenue by tier across the entire API portfolio

Warnings reach your on-call rather than waiting on a dashboard: capacity and tier events by webhook (Events to your on-call), limit breaches through the risk stream (Risk Management).

See also: Assurance · Metering · Cost Control · Subscribing