Gateway API and the mesh decision
Ingress to Gateway API route by route, and an honest list of what a service mesh costs you.
Ingress did one job and did it with annotations. Gateway API replaces it with real typed resources and, more importantly, a split of responsibility that matches how teams actually work.
Why the split matters more than the schema
- GatewayClass — the implementation. Installed once, by whoever runs the platform.
- Gateway — the listener: ports, protocols, TLS. Owned by the platform team.
- HTTPRoute — hostnames, paths, backends, weights. Owned by the application team, in their own namespace.
With Ingress, a developer needing a header-based rule had to edit an annotation blob that only the controller understood, usually in a shared object. With Gateway API they write an HTTPRoute in their namespace and attach it to a Gateway they are permitted to use. The RBAC boundary finally lines up with the ownership boundary.
A route, and a canary
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: api
namespace: team-a
spec:
parentRefs:
- name: public-gateway
namespace: infra
hostnames: ["api.example.com"]
rules:
- matches:
- path: { type: PathPrefix, value: /v2 }
backendRefs:
- { name: api-v2, port: 8080, weight: 90 }
- { name: api-v2-canary, port: 8080, weight: 10 }
Weighted backends are in the core specification, not an annotation and not a separate product. That single feature removes the most common reason teams installed a mesh.
Migrating without a flag day
- Install a Gateway API implementation alongside your existing ingress controller. They coexist happily.
- Create the Gateway with a different hostname —
new.api.example.com— and port it one route at a time. - Shift traffic at DNS, with a low TTL, once the new path is proven.
- Delete the Ingress objects last, and only after a week of quiet.
ReferenceGrant in mind: a route in one namespace pointing at a Service in another is denied by default. The target namespace has to publish a ReferenceGrant permitting it. This is a feature — it stops a team routing traffic to someone else’s service — and it is the most common reason a migrated route silently returns 404.When a service mesh is the wrong answer
A mesh gives you mutual TLS everywhere, uniform retries and timeouts, L7 authorisation, and per-call telemetry you did not have to instrument. Those are real. So is the bill:
- A sidecar per pod, or an eBPF agent per node. Either way, memory and CPU across your whole fleet — often 10–15% before you have shipped anything.
- Two more hops in every request path, and two more places a request can be dropped.
- A second control plane to upgrade, certificate-rotate and debug, with its own failure modes that look like application failures.
- Startup ordering — the classic sidecar race where your app starts before the proxy is ready and its first calls fail.
A reasonable sequence: Gateway API first, because you need an ingress anyway. Then NetworkPolicy, which covers most of what people want mTLS for internally. Reach for a mesh when you have a concrete requirement — mTLS for compliance, retries you cannot put in clients, L7 authorisation between services — not because the architecture diagram looks incomplete without one.
If you do want mTLS and nothing else, an eBPF CNI can give you transparent node-to-node encryption with no sidecars at all. That is a much smaller commitment than a full mesh and it covers the single most common compliance ask.
What to actually do with this
- Install a Gateway API implementation next to your ingress controller and port one route.
- Write down the specific requirement that would justify a mesh. If you cannot name one, you have your answer.
- Check whether your CNI can do transparent encryption before pricing a mesh.
Something wrong or out of date? Open an issue — corrections are welcome and get credited.