Section650

Writing

TopicArchitecture
Reading8 min

API Architecture: What Building and Operating Real Systems Has Taught Me

Endpoints are the visible part of an API. Ownership, contracts, authorization, failure behavior and operations decide whether it’s still useful two years later.

By Ka Lun Chan · Software architecture · Architecture / Leadership

Start with the consumer and the workflow

An API is a promise other people build on. The endpoints are the part everyone sees. Ownership, the contract, how it behaves when something fails and how it’s operated decide whether that promise is still worth anything two years later.

I’ve built and integrated APIs for more than 23 years across VoIP, publishing, SaaS and government applications: in Django and Django REST Framework, Flask, including Flask APIs at Brainome, FastAPI, Rails and Node.js, consumed by React and Next.js front ends and by other services. The frameworks changed often. The questions in this post didn’t. Many of them first came up operating a VoIP platform, where the contract between systems was a phone call that had to connect.

API architecture is how systems agree to cooperate: the contracts between them, who owns each one, how access is controlled, and how requests behave when something fails. Good API architecture lets consumers depend on a stable contract while the team changes the implementation safely.

To keep it concrete, I’ll use one hypothetical example throughout: a government application platform. Residents submit information, reviewers assess it, and external agency systems receive updates. That’s four consumers with different needs: a browser form, a reviewer dashboard, possibly a mobile app, and partner systems that integrate over months, not days.

Those consumers shouldn’t see the database. If the API mirrors the tables, every schema change becomes a breaking change for someone else, and partners end up depending on column names you never meant to promise. The contract should describe the workflow: submit an application, assign a reviewer, record a decision.

Choose an API style for a concrete reason

REST is an architectural style, usually over HTTP. GraphQL is a query language and runtime for an API. gRPC is an RPC framework built on HTTP/2 and Protocol Buffers. Django, FastAPI or Rails can implement more than one of them. The choice should come from the consumers, the tooling and what the team already knows how to run.

Fig. 669-1 REST, GraphQL and gRPC in practice
StyleGood fitStrengthsTradeoffs
RESTPublic and partner APIs, browser apps, resource-shaped dataWidely understood; HTTP caching, status codes and tooling work out of the boxScreens that need many resources can turn into many requests; versioning needs discipline
GraphQLFront ends that combine data from many sources and change oftenClients ask for exactly the fields they need; one typed schemaHTTP caching is harder; query cost and field-level authorization need deliberate work
gRPCInternal service-to-service calls with high volume or streamingCompact binary messages, generated clients, strict contractsBrowsers need a proxy such as gRPC-Web; payloads are harder to inspect by hand

For the hypothetical platform, REST would suit the partner integrations, because agencies can call it with standard tools. A reviewer dashboard pulling from many places might justify GraphQL or a dedicated endpoint for that screen. Neither choice says one style is better in general.

Design boundaries around responsibilities

Some things are resources: an application, a document, a reviewer. Others are business operations: submit, withdraw, approve. Modeling “approve” as an operation with its own rules is clearer than letting clients set a status field and hoping they set it correctly.

Each boundary needs an owner, the team that decides how it changes and gets paged when it breaks. An API boundary doesn’t require a separate service or deployment. A well-structured application can expose clean boundaries from one codebase, which I argue for in your team size should influence your architecture.

Treat the contract as something others depend on

Validate input at the edge and return errors in one consistent shape, with a machine-readable code and a message a person can act on. Paginate every list that can grow. Document filtering and sorting instead of letting consumers discover them. Keep the documentation generated from the same schema the code uses, so they can’t drift apart.

Then treat compatibility as a feature. Adding an optional field is usually safe. Renaming a field, changing a type or tightening validation is not. On the hypothetical platform, renaming one field in the application response could create work for the web team, the mobile team and three partner agencies, each on its own release schedule. Version deliberately, announce deprecations, and watch which consumers still use the old version before you remove it.

Security has to follow the data and the action

Authentication answers who is calling. Authorization answers whether they may do this specific thing to this specific record. A valid JWT proves the first. It says nothing about the second.

If a resident can change the ID in /applications/1042 and read someone else’s application, the token was valid and the API still failed. OWASP lists this, broken object level authorization, first in its API Security Top 10. Every request needs an object-level check: is this record in the caller’s tenant, is it theirs or assigned to them, and does their role allow this action? Privileged operations like approving or exporting deserve their own checks and an audit trail. Role-based access and SSO help, but neither replaces the check on the record itself.

Design for uncertain outcomes

A resident submits an application. The server saves it, then the response is lost on a slow mobile connection, and the client times out. From the client’s side, the outcome is unknown. A timeout doesn’t prove the operation failed. It only proves the client didn’t hear back.

If the client retries a plain POST, the resident may now have two applications. The fix is idempotency: the client sends a unique key with the request, and the server stores the result under that key. A retry with the same key returns the original result instead of creating a second record. Reads and well-designed PUTs are naturally safe to retry. Creates and anything that moves money or sends messages need the key.

Decide what belongs in the request path

The resident needs to know the application was received. They don’t need to wait while documents are scanned, partner systems are notified and reports are updated. The API can save the application, return 202 Accepted with a link to a status endpoint, and let background workers do the rest.

A job queue handles that well. If several systems need to react to the same submission, an event stream like Kafka may earn its place, with tradeoffs I cover in event-driven architecture with Kafka. Kafka isn’t required to make an API asynchronous.

Performance and operations cover the whole path

Slow APIs are rarely slow because of the framework. The usual causes are an N+1 query, a response carrying fields nobody reads, a screen that makes twenty requests where one would do, missing caching, a connection pool that runs dry, or a downstream call with no timeout. Rewriting the service in a faster framework keeps all of those.

Run the API as a production service. Track request rate, error rate and latency per endpoint. Write structured logs with a request ID that travels into downstream calls and traces. Make health checks reflect whether the service can do its job. Set rate limits per consumer, and deploy so old and new versions can run side by side during a rollout. When a partner calls at 4 p.m. on a Friday, that’s what lets support answer in minutes.

Cost, and matching the architecture to the team

API design shows up on the bill. Chatty clients multiply requests and gateway charges. Oversized responses add network transfer. Repeated queries add database load. Every extra service boundary adds hosting, observability and another hop. Third-party APIs you call can charge per request. Each integration also has to be supported long after launch.

The organizational costs are larger and harder to see: breaking changes that pull several teams off their roadmaps, endpoints nobody owns, and the same integration built three times because nobody knew it existed.

For many teams, one well-structured API is enough. A backend-for-frontend helps when a web app and a mobile app need different shapes of the same data. Independent services make sense when different teams own them, scale them differently or need a security boundary, and when the team can pay for that overhead.

A decision checklist

  1. Who consumes this API, and what do they need?
  2. Who owns its contract and production behavior?
  3. What happens when a request is retried or a dependency fails?
  4. How do we enforce access to each object and operation?
  5. How will consumers survive changes?
  6. Can the team operate this design at an acceptable cost?

Short answers

Should I use REST, GraphQL or gRPC?

Choose from the consumers’ needs. REST suits public, partner and browser APIs and works well with HTTP caching. GraphQL suits front ends that combine many data sources. gRPC suits high-volume internal service calls. Team familiarity and operational tooling matter as much as the style itself.

Does a valid JWT mean a user can access a record?

No. A valid JWT authenticates the caller. The API still has to check, on every request, that this caller may perform this action on this specific record, including tenant and ownership checks. Missing those checks is broken object level authorization, first on the OWASP API Security Top 10.

How do you make API retries safe?

Use idempotency. The client sends a unique key with a create or payment request, and the server stores the result under that key, so a retry after a timeout returns the original result instead of creating a duplicate.

A dependable contract, room to change

Consumers rarely care which framework an API runs on. They care whether it keeps doing what it promised.

A good API gives consumers a dependable contract and gives the team room to evolve the implementation safely.

Get the ownership, the authorization, the retries and the compatibility right, and the choice of framework becomes one of the easier decisions.

Talk through an API design