Skip to main content
APIsBackendArchitecture

REST API architecture: 7 practices so your backend doesn't fall over in production

Andres Betancourt7 min read

Most APIs fail in production not for lack of features, but because of architecture decisions nobody made on purpose. These are the seven practices that separate an API that survives your first real traffic spike from one that goes down on Black Friday — and how to prioritize them if you can't implement all of them at once.

1. Version the API from the first endpoint

A prefix like /api/v1/ costs nothing to add today, and it keeps you from breaking every client (web, mobile, integrations) the day you need to change a contract. The common mistake is thinking 'I don't need it yet, I'll add it when it matters' — but by the time it matters, there are already clients in production consuming the unprefixed version, and adding versioning retroactively means either maintaining two routes for the same thing or forcing an awkward client migration. The cost of versioning from day one is practically zero; the cost of adding it later is a full client migration.

2. Validate at the edge, not in the middle

Every external input gets validated before it touches business logic — with a schema library (Zod, Yup), not scattered if-statements across the codebase. This has a benefit that goes beyond security: when validation lives in one place, with an explicit schema, API documentation practically writes itself, and anyone new on the team can see exactly what shape each endpoint expects without reading the full implementation.

const OrderSchema = z.object({
  items: z.array(ItemSchema).min(1),
  customerId: z.string().uuid(),
});

const body = OrderSchema.parse(await req.json());

3. Consistent HTTP status codes

201 on create, 204 on delete with no content, 409 on conflicts, 422 on failed validation. A client consuming your API should be able to make decisions from the status code alone, without parsing the error message. This matters especially for automated integrations (another system consuming your API with no human reading the response): a webhook or a scheduled job needs to decide whether to retry, discard, or alert based purely on the code, and a backend that returns 200 with a `success: false` field buried in the body breaks that automated decision-making.

4. Authentication and authorization, kept separate

Authentication answers 'who are you?'. Authorization answers 'can you do this?'. Mixing them into the same middleware is the most common way to end up with an endpoint that leaks data it shouldn't. The safer pattern is to resolve identity once, early in the middleware chain, and then let each endpoint (or a dedicated per-resource middleware) decide the specific permission it needs — never assume 'authenticated' is the same as 'allowed to see this.'

5. Rate limiting before you need it

An endpoint with no rate limit isn't just vulnerable to abuse — a single misconfigured client (a broken cron job, a frontend stuck in a loop) can take down your backend without anyone attacking anything. Rate limiting doesn't need to be sophisticated from the start: a simple per-IP or per-API-token limit, with a clear 429 response and a `Retry-After` header, already covers most real-world cases. What matters is that it exists before the first incident, not after.

6. Idempotency on critical operations

If a payment is retried after a network timeout, the second call shouldn't charge twice. An idempotency key (Idempotency-Key) in the header solves this without complex logic in every endpoint: the server stores the result of the first execution tied to that key, and if another request arrives with the same key, it returns the stored result instead of running the operation again. This is especially critical for any endpoint that moves money, creates resources with external side effects (sending an email, triggering a notification), or modifies inventory.

const existing = await getIdempotentResult(idempotencyKey);
if (existing) return existing;

const result = await processPayment(body);
await saveIdempotentResult(idempotencyKey, result);
return result;

A common mistake is assuming idempotency only matters for payments. Any operation with an external effect that isn't trivially reversible — sending a push notification, firing a transactional email, creating a ticket in an external system — benefits from the same pattern. The practical rule is to ask: if this call runs twice by accident, does the visible result for the user change noticeably? If the answer is yes, that endpoint needs its own idempotency key, whether it moves money or not.

7. Logging and observability from day one

When something breaks in production at 2am, the question isn't 'what happened' but 'do I have a way to find out?'. Structured logs with a correlation ID per request are the difference between a five-minute diagnosis and an entire night of guessing. The correlation ID is the key piece: generated at the start of each request and propagated through every log line, external service call, and background job that request triggers, it lets you reconstruct the full path of a specific request without manually cross-referencing logs by approximate timestamp.

'Structured' is the word that makes the difference here: a JSON-formatted log with consistent fields (timestamp, level, correlation ID, message, extra context) can be queried, filtered, and aggregated with tooling — a free-text log can only be read, line by line, hoping to spot the relevant one among thousands. The investment of turning `console.log('error processing order')` into a structured log with the order ID and the correlation ID is minimal at write time, and enormous the first time you have to investigate a real incident under pressure.

8. Consistent pagination and filtering

A listing endpoint with no pagination is a time bomb: it works perfectly with ten records in development and falls over with a hundred thousand in production. The simplest, most predictable convention is cursor-based pagination (an identifier for the last item seen, not a page number) for fast-growing collections, because it avoids the problem of duplicated or skipped results when new records get inserted between one page and the next. Filtering and sorting should follow the same query-param convention across every endpoint in the API (`?status=active&sort=-createdAt`), not a different convention per resource — inconsistency between endpoints is one of the most common sources of integration bugs in clients that consume the API.

9. CORS configured on purpose, not copy-pasted from Stack Overflow

It's common to see `Access-Control-Allow-Origin: *` in production because 'that's what made the error go away in development' — and it's exactly the kind of configuration that seems to work fine right up until it becomes a real security problem, especially on endpoints handling session cookies or sensitive data. The correct configuration explicitly declares which origins are allowed (the real frontend domain, not a wildcard), which methods and headers are permitted, and whether credentials are allowed — every one of those values should be a conscious decision, not the first result that made the local error disappear.

How to prioritize if you can't do them all at once

On a new project, the reasonable order is: versioning and validation first (they're nearly free and prevent immediate technical debt), separated authentication/authorization and consistent status codes second (they affect how every endpoint gets designed from the start), pagination and CORS when defining the first listing endpoints (retrofitting them later is more expensive than designing them right from the first resource), and rate limiting, idempotency, and observability as the next wave, before the first launch with real traffic — not after the first incident that would have been prevented by them.

A pattern that repeats in projects we inherit from another vendor is that none of these practices are completely missing — what's missing is consistency. One endpoint versions and another doesn't, one validates with a schema and another with scattered if-statements, one paginates with a cursor and another with an offset. That inconsistency is, in practice, almost as costly as not having the practice at all, because it forces every API client to handle special cases per endpoint instead of assuming a uniform contract. The discipline of applying these seven — or nine — practices evenly across the entire backend matters as much as each individual practice.

The most effective way to sustain that consistency over time, especially as the team grows, is documenting it in a single place anyone new can read before writing their first endpoint — a brief internal API style guide, with real examples from the project itself. It doesn't replace code review, but it drastically reduces how often that review ends up debating basic conventions instead of the actual business logic of the change. An API documented with OpenAPI/Swagger generated from the same validation schemas (instead of maintained by hand separately) serves that same purpose automatically, without depending on someone remembering to update a separate document.

None of these practices are exotic — they're documented architecture decisions that cost the same to implement well from the start as to get wrong, and they're exactly what we review whenever we audit or hand an API off to a client.

Back to blog