Skip to main content

Command Palette

Search for a command to run...

Rate-limiting

Published
•3 min read•View as Markdown

Rate limiting is a strategy for limiting network traffic, Rate limiting makes it harder for malicious actors to overburden the system and cause attacks,

A rate limit is the maximum number of calls you want to allow in a particular time interval.
Types of Rate Limiting:

Rate limiting are of two types: hard (enforced) or soft

i) Hard Rate limit: if the rate limiting call exceeds the limit, then the call is aborted and an error is returned.When a hard rate limit is reached, no more calls are accepted from that customer until the beginning of the next time period.

ii) Soft Rate limit: A soft rate limit allows the call to complete but logs a warning message.

Where can you apply rate limits?

You can configure rate limiting at several levels of an API

i) You can define a rate limit for a specific path and operation of an API

ii) You can define a rate limit on an API in the API definition (the Design page in the user interface).

iii)You can define rate limit on an API assembly, When you apply a rate limit to an API assembly, the limit only affects calls at that point in the assembly

Why rate limiting is used

i) Preventing resource starvation: The most common reason for rate limiting is to improve the availability of API-based services by avoiding resource starvation, Generally, a service applies rate limiting at a step before the constrained resource, with some advanced-warning safety margin. Margin is required because there can be some lag in loads, and the protection of rate limiting needs to be in place before critical contention for a resource happens. For example, a RESTful API might apply rate limiting to protect an underlying database; without rate limiting, a scalable API service could make large numbers of calls to the database concurrently, and the database might not be able to send clear rate-limiting signals.

ii) Managing policies and quotas: When the capacity of a service is shared among many users or consumers, it can apply rate limiting per user to provide fair and reasonable use, without affecting other users.

iii) Controlling flow: In complex linked systems that process large volumes of data and messages, you can use rate limiting to control these flows—whether merging many streams into a single service or distributing a single work stream to many workers.

How does rate limiting work for APIs?

An API, or application programming interface, is a way to request functionality from a program. APIs are invisible to most users, but they're extremely important for applications to function properly. For example, a restaurant's website could rely upon the API of a table reservation service to enable customers to make reservations online. Or, an eCommerce platform could integrate a shipping company's API to provide users with accurate shipping costs.

Every time an API responds to a request, the owner of that API has to pay for compute time: the server resources required for code to run and produce a response to that API request. In the example above, the restaurant's API integration will cause the table reservation service to pay for compute time whenever a restaurant customer makes a reservation.

For this reason, any application or service that offers an API for developers will have limitations on how many API calls can be made per hour or day by each unique user. In this way, third-party developers don't overuse an API.

Rate limiting can also motivate developers to pay more for leveraging the API: often they can only make so many API calls before paying more for the API service.

Rate limiting for APIs helps protect against malicious bot attacks as well. An attacker can use bots to make so many repeated calls to an API that it renders the service unavailable for anyone else, or crashes the service altogether. This is a type of DoS or DDoS attack.