Go Microservices in Production: Lessons From Building a Distributed Platform From Scratch

Go microservices hold up in production when each service owns its data, every outbound call has a timeout, and you can follow one request through all of them. I’ve built distributed systems in Go at LodgeLink and 90POE and worked on a global ad network at Aylo, and most of what went wrong had little to do with Go. This post is the list I’d hand a team about to split its first monolith, including the parts I’d skip.

This is the first post in the Go in production series.

Table of Contents

Build the modular monolith first

If the team is under ten people and the product is still changing weekly, I’d start with one Go binary, one database, and strict package boundaries inside it. Microservices make the boundaries physical, and physical boundaries are expensive to move. A package boundary moves in one pull request.

The layout I use for a service is the same whether it ends up alone or as one of thirty: domain logic in the middle, ports as interfaces, adapters for HTTP, Postgres and the broker on the outside. I wrote up how that looks in a real codebase in the pREST v2 hexagonal refactor. A module that only talks to the rest of the binary through an interface can be lifted into its own process later without touching the domain code.

When I do split, it’s for one of these reasons:

  • Two parts of the system need to scale very differently (a public read path versus a nightly batch job).
  • Two teams keep blocking each other in the same deploy pipeline.
  • One part has a different failure or security profile and I want a blast radius around it.
  • One part needs a runtime the rest doesn’t (another language, a GPU, an old dependency).

“It would be cleaner” is not on the list. LodgeLink 3.0 was a net-new platform built as distributed Go microservices from the start, and that was the right call there because several pods were building in parallel from day one. With two other engineers I would have shipped a monolith and split it a year later.

Service boundaries that survive

The boundaries that held up for me were drawn around data that one team writes and everyone else reads. The ones that failed were drawn around nouns from a whiteboard, or around technical layers.

Boundary drawn around Tends to Why
A business capability with its own data (bookings, invoicing, identity) Survive One owner, one write path, a stable contract outward
A team that already exists Survive The service follows the people who are on call for it
A technical layer (the “database service”, the “validation service”) Fail Every feature touches every layer, so every change touches every service
A noun every flow needs (the “user service” called on every request) Become a bottleneck It collects everyone’s fields and everyone’s latency budget
A single CRUD table Fail Too small to own anything; it becomes a remote DAO with network cost

Two tests before I agree to a new service. Can the owning team describe what it means for this service to be down, in product terms, without naming another service? If the answer is “then nothing works”, the boundary is wrong. And can they deploy it without coordinating with anyone? Two services that always ship together are one service with extra steps.

Team shape matters more than the diagram. A service with no owning team drifts, and a team with five services it doesn’t understand treats them as one. When I led the Go chapter at 90POE, the useful work was less about which boxes existed and more about making sure every box had a team that could explain it.

Sync or async between services

My default is synchronous for queries and for anything the caller must know the result of right now, asynchronous for everything that can be told later. “Create the booking” is sync. “Send the confirmation, update the search index, notify finance” is async.

HTTP or gRPC call Message on a broker
Caller needs the answer now Yes No
Callee down Caller fails or degrades immediately Message waits; caller unaffected
Ordering Per request Per partition or key, if the broker supports it
Debugging One trace, one span chain Trace has to carry through message headers
Coupling Caller knows the callee’s address and contract Caller knows a topic and an event schema
Failure mode to design for Timeouts and partial failure Duplicates and out-of-order delivery

Within the sync column I use plain HTTP with JSON unless there’s a reason for gRPC. gRPC gives you generated clients, streaming and a schema in the repo. It costs you readable logs, an easy curl story, and a protobuf toolchain in every service. On a team already fluent in protobuf it’s pleasant; on one that isn’t, HTTP plus a shared OpenAPI document gets most of the benefit.

The async column has one rule: every consumer must be idempotent. Brokers redeliver. Deploys restart consumers mid-batch. Your producer will retry after a timeout on a publish that succeeded. In practice that means a processed-events table or a unique constraint on the event ID, written in the same transaction as the side effect, so handling event 4821 a second time does nothing.

Data ownership and the eventual-consistency tax

One service, one database, and nobody else gets a connection string. A shared database is the fastest route to a distributed monolith, because the schema becomes the API and nobody owns it.

The price of that rule is that any view spanning two services is eventually consistent. The booking service knows a booking exists some milliseconds before the reporting service does. Most of the time nobody notices. Some of the time a user creates something, refreshes, and it isn’t there, and that becomes a support ticket.

Ways I’ve paid that tax that were worth it:

  • Return the created object in the write response and render the UI from that, so the user never reads back from the slow path.
  • Show the read model’s lag in the UI (“updated 2 seconds ago”) instead of pretending it’s live.
  • Keep user-visible reads inside the owning service and let the eventually consistent copies serve analytics and search.
  • Use an outbox table so the event and the row commit together. A service that writes to Postgres and then publishes to the broker as a second step will, eventually, do one without the other.

Two that weren’t worth it: distributed transactions across services, and a shared read replica that every service queries directly. Both fix today’s ticket and cost you the next two years.

Observability is a prerequisite

The first incident on a distributed system teaches you that logs from eight services, each with its own timestamp format and no shared ID, are eight separate mysteries. I’d now refuse to split a monolith until three things exist: structured logs, a request ID that crosses every hop, and distributed tracing.

Structured logs mean log/slog with JSON output and the same field names everywhere, kept in a shared package. The request ID is the cheapest of the three. Accept it from the inbound header if a trusted caller set it, mint one if not, store it in the context, and send it on every outbound call.

package reqid

import (
	"context"
	"crypto/rand"
	"net/http"
)

const Header = "X-Request-ID"

type ctxKey struct{}

// Middleware reads or mints a request ID and stores it in the context.
func Middleware(next http.Handler) http.Handler {
	return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
		id := r.Header.Get(Header)
		if id == "" {
			id = rand.Text() // Go 1.24+; crypto/rand, base32
		}
		w.Header().Set(Header, id)
		ctx := context.WithValue(r.Context(), ctxKey{}, id)
		next.ServeHTTP(w, r.WithContext(ctx))
	})
}

// From returns the request ID stored in ctx, or "".
func From(ctx context.Context) string {
	id, _ := ctx.Value(ctxKey{}).(string)
	return id
}

On the outbound side every client does req.Header.Set(reqid.Header, reqid.From(ctx)). A request ID on an error page then turns into one query across all logs.

Tracing is the next step. OpenTelemetry’s HTTP instrumentation propagates a traceparent header for you and gives you the span tree across services, so you can see that the slow checkout was mostly waiting on the pricing service. I added OpenTelemetry to pREST in v2.4.0 with one [otel] config block, HTTP spans named by route template, and Postgres spans at the driver seam; the v2.4.0 post walks through the config and a local SigNoz stack to view it, and the same shape works for any Go service.

One detail from that work I’d repeat everywhere: name spans and metric labels by the matched route template (GET /bookings/{id}), never the raw URL. Otherwise every ID becomes its own time series and your metrics bill becomes a conversation.

The Kubernetes basics that matter

Most Go services I’ve run ended up on Kubernetes, and most of the Kubernetes problems came from four settings.

Requests and limits. The Kubernetes docs put it plainly: the scheduler uses the request to pick a node, and the kubelet enforces the limit so the container “is not allowed to use more of that resource than the limit you set”. The two resources fail differently. Over the memory limit, the kernel may terminate the container. Over the CPU limit, the kernel waits before letting the cgroup run again, so the process is throttled with no error anywhere. In Go that shows up as latency spikes when GOMAXPROCS is higher than the limit allows. Since Go 1.25 the runtime defaults GOMAXPROCS to the cgroup CPU limit on Linux; on older Go, set it yourself. Take the memory numbers from a load test.

Readiness versus liveness. Per the probes documentation, a failed readiness probe removes the Pod’s IP from the EndpointSlices of its Services, so traffic stops; a failed liveness probe, past the configured tolerance, makes the kubelet restart the container. The same page warns that liveness probes “must be configured carefully to ensure that they truly indicate unrecoverable application failure, for example a deadlock”, because getting it wrong “can lead to cascading failures”. My rule is that readiness checks dependencies (can I reach my database?) and liveness checks only the process (is the HTTP server answering?). A liveness probe that fails when Postgres is slow restarts every replica during the exact minute you need them. For slow warm-ups, a startup probe holds the other two off until the app has started.

PodDisruptionBudget. A PDB “limits the number of Pods of a replicated application that are down simultaneously from voluntary disruptions”, in the words of the disruptions page: node drains, upgrades, scale-downs. Without one, a routine node pool upgrade can take all three replicas of a service down together. Set minAvailable on anything that serves traffic. The page also notes that a PDB can’t prevent involuntary disruptions such as hardware failure, though those still count against the budget.

One process per container. This one is my opinion; Kubernetes doesn’t enforce it. A Go binary is a single static executable and runs fine as PID 1. Add a shell loop or a second server to the same container and the probes stop meaning anything, because they check one process while the other dies silently.

Resilience: timeouts, retries, breakers, bulkheads

Every production outage I’ve been part of on a distributed system had a missing timeout somewhere in the chain. Go’s http.DefaultClient has none; the net/http docs say “A Timeout of zero means no timeout.” One slow downstream can then pin every goroutine in your service until it runs out of memory.

I set two timeouts on every outbound call. The http.Client timeout is the ceiling, and it covers connection time, redirects and reading the body. The context deadline is the per-call budget, and it inherits whatever the inbound request had left.

var client = &http.Client{Timeout: 5 * time.Second}

func fetchPrice(ctx context.Context, id string) (Price, error) {
	ctx, cancel := context.WithTimeout(ctx, 800*time.Millisecond)
	defer cancel()

	req, err := http.NewRequestWithContext(ctx, http.MethodGet, pricingURL+"/prices/"+id, nil)
	if err != nil {
		return Price{}, err
	}
	req.Header.Set(reqid.Header, reqid.From(ctx))

	resp, err := client.Do(req)
	if err != nil {
		return Price{}, err
	}
	defer resp.Body.Close()

	if resp.StatusCode != http.StatusOK {
		return Price{}, fmt.Errorf("pricing: unexpected status %d", resp.StatusCode)
	}
	var p Price
	err = json.NewDecoder(resp.Body).Decode(&p)
	return p, err
}

The body is decoded before the function returns so cancel() doesn’t cut it off mid-read. If the inbound request only has 300 ms left, the child context gets 300 ms, which is what you want; there’s no point finishing a call whose caller already gave up.

Retries come after timeouts, and they need a budget. A retry on every failure from every caller is how a 10% error rate at one service becomes double the load on the service below it. Retry only errors that are safe to retry (timeouts, 503s, connection resets), only on idempotent operations, cap it at two or three attempts, and add jitter so the retries don’t line up. I’d also cap retries as a share of each service’s total traffic, so a bad minute can’t amplify itself.

Circuit breakers stop the calls that are going to fail anyway. If the pricing service failed the last fifty calls, calling it again for the next ten seconds helps nobody; return the cached price or a degraded response and let it recover.

Bulkheads keep one bad dependency from taking the whole process down. In Go that usually means a separate http.Client and transport per downstream, with its own connection pool limits, and a semaphore in front of anything expensive. If search hangs, checkout shouldn’t notice.

The organizational cost

The part I underestimated most is that a Go service is a few thousand lines, and nearly all the cost sits around it: ownership, standards, deploys, on-call.

Pods, or whatever the company calls its teams, need clear ownership of each service and an on-call rotation that matches. A service that pages a team who didn’t write it gets patched and never fixed.

Standards are what made the LodgeLink work manageable across pods: one service template with the same layout, logging fields, health endpoints, Dockerfile and CI pipeline. The template is boring on purpose. Every time a team deviated from it for a good reason, the good reason came back as an incident a few months later. A platform team helps here; at Globo, where I contributed to the Tsuru PaaS, the platform was the template, and services that used it got deploys, logs and routing the same way without each team deciding. I also keep a short list of what a service needs before it takes production traffic: structured logs with a request ID, readiness and liveness endpoints, timeouts on every client, a dashboard with request rate, error rate and latency, and a documented owner.

The smell I watch for is the distributed monolith. You have one when a feature needs four pull requests in four repos deployed in a specific order, when the shared integration environment is the only place anything works, or when services share a database or a library whose version must match everywhere. You’re paying the operational cost of microservices for the coupling of a monolith. The fix is usually to merge a few services back together, which teams are oddly reluctant to do.

This topic comes up every time I talk to Go developers in Calgary; at the first Golang Calgary meetup the hallway questions were as much about team boundaries as about goroutines. If you’re working through a split and want a second opinion, my background is on the bio page and I’m happy to talk.

FAQ

When should you split a monolith into microservices?

Split when parts of the system need to scale, deploy or fail independently, or when separate teams keep blocking each other in one codebase. Until then a modular monolith with strict package boundaries gives most of the benefit at a fraction of the operational cost. If you can’t name the team that will own and be on call for the new service, you’re not ready to create it.

Is Go a good language for microservices?

Yes, mostly for operational reasons. A Go service compiles to one static binary, starts fast, uses little memory at idle, and has an HTTP server, a JSON codec and context-based cancellation in the standard library. Much of the tooling around gRPC, OpenTelemetry and Kubernetes is itself written in Go, so it fits well.

How many microservices is too many?

More than the team can explain and keep on call for. A reasonable ceiling is one to three services per owning team; a team of five running fifteen services spends its time on plumbing. If a feature routinely needs changes in four or more services, some of them should merge.

What is a distributed monolith?

A distributed monolith is a set of services that must be deployed together, share a database or schema, or can only be tested as a whole. It has the network cost and operational overhead of microservices with the coupling of a monolith. The usual causes are boundaries drawn around technical layers and a shared database that became the real API.

gRPC or HTTP between Go services?

Either works; pick based on the team. gRPC gives generated clients, streaming and a schema in the repo, at the cost of a protobuf toolchain and less readable traffic. Plain HTTP with JSON and an OpenAPI document is easier to debug with curl and to onboard people onto. The choice matters far less than having timeouts, request IDs and idempotent consumers on whichever you pick.