GraphQL Federation: Building a Unified API Gateway for Microservices

Microservices solve organizational scaling problems. Each team owns a service, deploys independently, and chooses its own technology stack. But microservices create an API fragmentation problem — clients need data from multiple services to render a single page, and orchestrating those calls from the client is painful. GraphQL Federation solves this by composing multiple service schemas into a single, unified API that clients query as if it were one server.

Federation is not schema stitching. Stitching merges schemas at runtime, delegating field resolution across services in a pattern that is difficult to debug and impossible to validate statically. Federation composes schemas at build time using directives that declare type ownership and cross-service references. The composed supergraph is a static artifact that can be validated, tested, and versioned before deployment.

Federation Architecture

A federated GraphQL architecture has three components: subgraphs, the gateway (also called the router), and the schema registry. Each subgraph is an independent GraphQL server that defines the types and fields it owns. The gateway composes these subgraphs into a supergraph and routes incoming queries to the appropriate subgraphs. The schema registry stores subgraph schemas and validates composition before deployment.

When a client sends a query, the gateway creates a query plan — a directed acyclic graph of subgraph requests that resolves the query with the minimum number of network calls. The query planner considers type ownership, key fields, and cross-service references to determine the optimal execution order. Fields from the same subgraph are batched into a single request. Fields that depend on results from another subgraph are fetched sequentially.

Designing Subgraphs

The primary design decision in federation is type ownership. Each type should have exactly one owning subgraph that defines its key fields and core data. Other subgraphs can extend the type with additional fields relevant to their domain.

# Users subgraph — owns the User type
type User @key(fields: "id") {
  id: ID!
  email: String!
  displayName: String!
  createdAt: DateTime!
}

type Query {
  user(id: ID!): User
  me: User
}
# Orders subgraph — extends User with order data
type User @key(fields: "id") {
  id: ID! @external
  orders(limit: Int = 10): [Order!]!
  totalSpent: Float!
}

type Order @key(fields: "id") {
  id: ID!
  items: [OrderItem!]!
  total: Float!
  status: OrderStatus!
  placedAt: DateTime!
}

The @external directive on id in the Orders subgraph tells the gateway that the Orders service does not own the id field — it receives the id from the gateway when resolving orders or totalSpent for a specific user. This reference resolution is the core mechanism that enables cross-service type composition.

Entity Resolution

Entity resolution is how the gateway fetches data across subgraph boundaries. When the gateway needs to resolve fields from a subgraph that extends a type owned by another subgraph, it calls the extending subgraph's _entities resolver with the key fields. The extending subgraph looks up the entity by its key and returns the requested fields.

// Orders subgraph entity resolver (Node.js)
const resolvers = {
  User: {
    __resolveReference(user) {
      // user.id is provided by the gateway
      return {
        id: user.id,
        orders: () => orderService.getOrdersByUserId(user.id),
        totalSpent: () => orderService.getTotalSpentByUserId(user.id)
      };
    }
  }
};

The __resolveReference function receives an object containing the key fields (id in this case) and returns the extended fields. This pattern keeps cross-service data resolution inside the gateway's query execution — subgraphs never call each other directly, which eliminates circular dependencies and simplifies the service topology.

Query Planning and Performance

The query planner's job is to minimize the number of subgraph requests needed to resolve a query. Consider a query that fetches a user's profile, their recent orders, and product details for each order. The planner executes this as a cascade: first it fetches the user from the Users subgraph, then it fetches orders from the Orders subgraph using the user's id, then it fetches products from the Products subgraph using the product ids from the order items.

This cascade introduces latency proportional to the depth of cross-service references. Three sequential subgraph requests, each taking 50ms, add 150ms to the query. Two optimization techniques reduce this cost.

First, batch entity resolution. When the gateway needs to resolve 20 orders, it sends a single _entities query with all 20 keys rather than 20 individual queries. The subgraph's __resolveReference receives all keys at once and can use a single database query to fetch them. This is critical for the performance metrics that matter in production.

// Batched entity resolution with DataLoader
const resolvers = {
  Order: {
    __resolveReference: async (orders) => {
      const ids = orders.map(o => o.id);
      const results = await orderLoader.loadMany(ids);
      return results;
    }
  }
};

Second, use the @requires directive to fetch dependent fields from the owning subgraph alongside the key. If the Orders subgraph needs the user's email to calculate shipping costs, it can declare email as a required field. The gateway fetches the email from the Users subgraph and includes it in the entity resolution request, avoiding a separate roundtrip.

Schema Evolution and Versioning

Federated schemas evolve independently — each team can add fields, types, and queries to their subgraph without coordinating with other teams. The schema registry validates composition after each change, ensuring the supergraph is valid before the change reaches production.

Breaking changes — removing fields, changing types, or altering key fields — require the same expand-contract approach used for database migrations. Deprecate the field first, give clients time to migrate, then remove it. The schema registry tracks field usage from the gateway's query logs, showing which clients still reference deprecated fields.

# Deprecating a field before removal
type User @key(fields: "id") {
  id: ID!
  email: String!
  displayName: String!
  name: String @deprecated(reason: "Use displayName instead. Will be removed 2027-01-01.")
}

Error Handling and Partial Failures

Federation's partial failure model is both its greatest strength and its most subtle complexity. When the Orders subgraph is down, the gateway can still resolve user.email and user.displayName from the Users subgraph. The user.orders field returns null with an error extension. The client receives a partial response and can render the user profile without order data.

This behavior requires clients to handle nullable fields defensively. Fields that are String! in a subgraph become effectively nullable from the client's perspective because the gateway may return null if the owning subgraph fails. Design client code to handle nulls for any federated field, regardless of the schema's nullability declarations.

Circuit breakers at the gateway level prevent cascading failures. When a subgraph fails repeatedly, the circuit breaker opens and returns cached responses or null values immediately instead of waiting for timeouts. Configure separate circuit breakers per subgraph so that a failure in the Orders service does not affect the Users service's circuit state.

Monitoring and Observability

Federated architectures require distributed tracing to diagnose performance issues. A single client query may fan out to five subgraph requests, each with its own latency, error rate, and resource consumption. Without tracing, it is impossible to determine which subgraph caused a slow response.

The gateway should propagate trace context headers (like traceparent from W3C Trace Context) to every subgraph request. Each subgraph adds its own span to the trace, creating a complete picture of the query execution across all services. Track these metrics per subgraph: p50/p95/p99 latency, error rate, entity resolution batch size, and query plan complexity.

# Apollo Router configuration for tracing
telemetry:
  exporters:
    tracing:
      otlp:
        endpoint: http://otel-collector:4317
  instrumentation:
    spans:
      mode: spec_compliant
      subgraph:
        attributes:
          subgraph.name: true
          subgraph.graphql.operation.name: true

Field-level usage tracking is equally important. The gateway logs which fields are queried, how often, and by which clients. This data drives deprecation decisions, capacity planning, and performance optimization. A field that is queried once a month does not justify the same caching investment as a field queried ten thousand times per second.

When Federation Is the Wrong Choice

Federation adds operational complexity: a schema registry, a gateway process, composition validation in CI, distributed tracing, and subgraph-level monitoring. For teams with fewer than three backend services, this complexity is not justified. A single monolithic GraphQL server with clear module boundaries provides the same developer experience for clients without the distributed systems overhead.

Federation also struggles with highly interconnected domains where every type references every other type. The query planner's cascade model works well when the type graph is shallow (two to three levels of cross-service references) but degrades with deep graphs. If most queries require data from every subgraph, the latency overhead of entity resolution may exceed the latency of a monolithic server that accesses all databases directly.

Start with a monolith and extract subgraphs when team boundaries solidify. The build tooling for federation is mature enough that extraction is a mechanical process once the domain boundaries are clear. Premature federation — splitting a schema before teams have stable ownership — creates coordination costs that outweigh the benefits of independent deployment.