easypprof: Production-Safe pprof for Go Services

easypprof is a Go package I wrote so that pprof can stay on in production without becoming an incident. It serves the standard /debug/pprof/ endpoints on a dedicated listener bound to 127.0.0.1:6060, behind named tokens, with profile duration capped at 60 seconds, two concurrent profiles at most, and a slog audit line for every request. It is MIT licensed, needs Go 1.23 or newer, and go get github.com/arxdsilva/easypprof is the whole install.

Table of Contents

The problem with import _ “net/http/pprof”

The usual way to get profiling in a Go service is one import line. net/http/pprof registers its handlers on http.DefaultServeMux at import time, which is convenient right up to the moment that mux is the one serving your public port. From then on anyone who can reach the service can pull a heap profile, a full goroutine dump with ?debug=2, a 30 second CPU profile, or hold a /debug/pprof/trace open for as long as they like.

None of that leaks variable values. Profiles carry function names, file paths and stack shapes. That is still plenty: the goroutine dump shows what your service is doing and to whom, the file paths show your internal layout, and an unbounded CPU profile is a cheap way to make a service slow on purpose. It is also one of the first things a scanner tries, because the path is the same on every Go service in the world.

The standard fixes are all manual. Put pprof on a second mux, remember not to use the default one anywhere else, add auth in front of it, cap the seconds parameter yourself, and hope nobody adds import _ "net/http/pprof" back in a year from now. I have done that dance enough times on services where latency matters, like the trading systems I work on today, to want the safe version to be the easy one.

What easypprof does instead

The package never touches DefaultServeMux. You construct it with a config, start it, and shut it down with the rest of the service:

dbg, err := easypprof.New(easypprof.Config{
    Tokens: map[string]string{"oncall": os.Getenv("PPROF_TOKEN")},
})
if err != nil {
    log.Fatal(err)
}
if err := dbg.Start(); err != nil {
    log.Fatal(err)
}
defer dbg.Shutdown(context.Background())

That gives you the full set of endpoints under /debug/pprof/ (index, heap, allocs, goroutine, block, mutex, threadcreate, profile?seconds=N, trace?seconds=N, and flightrecorder on Go 1.25) on 127.0.0.1:6060, GET only, and a token check on every one. Pulling a profile is one curl:

curl -sH "Authorization: Bearer $PPROF_TOKEN" -o cpu.pb.gz \
  'localhost:6060/debug/pprof/profile?seconds=30'
go tool pprof -http=:7070 cpu.pb.gz

HTTP Basic works too, with the token name as the username, which is handy because go tool pprof can fetch straight from a URL:

go tool pprof -http=:7070 "http://oncall:$PPROF_TOKEN@localhost:6060/debug/pprof/heap"

The defaults are the part I care most about, because defaults are what ships:

Setting Default What it does
Addr 127.0.0.1:6060 Loopback only. A non-loopback address is refused unless you set tokens, an Authorizer, or explicitly AllowUnauthenticated: true
Tokens none Name to token. New rejects any token shorter than 16 characters
AllowedCIDRs none Checked against the TCP peer address, never a header
MaxProfileDuration 60s Caps ?seconds= on profile and trace; over the cap is a 400, not a silent clamp
MaxConcurrent 2 Extra profile requests get a 429 with Retry-After: 5
BlockProfileRate 10000 Nanoseconds, applied with runtime.SetBlockProfileRate on Start
MutexProfileFraction 100 One in a hundred events, applied on Start
EnableCmdline false /cmdline is off unless you ask
DisableFullGoroutineDump false Set it to reject goroutine?debug=2
Logger slog.Default() Audit and lifecycle logs

Two of those are opinions, not just safety. Block and mutex profiling are off in a stock Go binary, so /debug/pprof/block and /debug/pprof/mutex come back empty exactly when you need them during an incident. easypprof turns them on at a sampling rate that costs little, and restores the previous rates on Shutdown.

Three ways to run it

The default is the separate listener above, and for most services that is the right answer: the profiling port is a different socket from the API port, and whether it is reachable is a network question you can answer with ss or a port-forward.

If you would rather not open a second port, mount the handlers on the mux you already have. Every control still applies:

if err := dbg.Mount(mux); err != nil {
    log.Fatal(err)
}

For chi, gorilla or anything else that is not a *http.ServeMux, take the handler and route it yourself:

h, err := dbg.Handler()
if err != nil {
    log.Fatal(err)
}
r.Handle("/debug/pprof/*", h)

Embedding changes four things, and the README spells them out. Authentication becomes mandatory, because the mux is presumably reachable from somewhere other than loopback. The server’s WriteTimeout is extended for the long-running endpoints, since a 30 second CPU profile would otherwise be cut off by a 10 second server deadline. X-Forwarded-For is never trusted for the CIDR check, so behind a proxy you rely on tokens, which is the correct answer anyway. And because routes cannot be removed from a ServeMux, after Shutdown the endpoints answer 503 instead of disappearing.

What a request goes through

The handler is a chain of checks in a fixed order, and knowing the order explains every status code you will see:

  1. Disabled or stopped: 503, before anything else, so a disabled profiler does no work at all.
  2. Method: GET and HEAD only, anything else is a 405 with an Allow header.
  3. Network: if AllowedCIDRs is set, the peer address from RemoteAddr is parsed, IPv4-mapped IPv6 addresses are unmapped, and a miss is a 403.
  4. Auth: Bearer or Basic. Tokens are stored as SHA-256 hashes and compared with subtle.ConstantTimeCompare, iterating every configured token without an early exit so timing does not reveal which name matched.
  5. Duration: ?seconds= above MaxProfileDuration is a 400 that names the limit. This also applies to the implicit 30 second default of profile if you set the cap lower than that.
  6. Concurrency: a semaphore sized by MaxConcurrent, skipped for the index page, 429 with Retry-After when full.
  7. Write deadline: for profile, trace and flightrecorder, the response deadline is pushed to MaxProfileDuration plus 30 seconds.
  8. Audit: one slog record per request with subject, remote, method, path, query, status, bytes, duration, and a reason on denials.

The audit line is the piece I would not want to run without. “Who pulled a goroutine dump from the payments service at 03:12” is a question you want to be able to answer from logs, not from memory.

Kubernetes

With the loopback default, the access path in a cluster is a port-forward, and the access control is whatever RBAC you already have on pods/portforward:

kubectl port-forward pod/ledger-7c9f 6060:6060

If a scraper such as Pyroscope or Parca needs to reach it from inside the cluster, bind to the pod IP, give the scraper its own token, and allowlist the pod network:

easypprof.Config{
    Addr:         os.Getenv("POD_IP") + ":6060",
    Tokens:       map[string]string{"pyroscope": os.Getenv("PPROF_SCRAPER_TOKEN")},
    AllowedCIDRs: []string{"10.20.0.0/16"},
}

Named tokens matter here: the scraper’s token and the on-call engineer’s token are different entries, so the audit log says which one was used and you can rotate one without the other. Nothing about this should ever sit behind a LoadBalancer Service or an Ingress.

Turning it off without a deploy

dbg.Disable() makes every endpoint return 503 before the auth check, and dbg.Enable() brings it back. The listener stays open and nothing is unmounted, so flipping it is instant. The README shows two ways to wire that: an admin endpoint using Go 1.22 method patterns, and a SIGUSR1 handler that toggles the state. I like the signal version for services that already have an admin port they would rather not grow.

Labels, so profiles say which operation was hot

A CPU profile of a service with forty handlers tells you which functions were hot, not which requests. pprof labels fix that, and easypprof wraps the two shapes you need:

easypprof.Do(ctx, func(ctx context.Context) {
    postEntry(ctx, e)
}, "op", "post_entry")

mux.Handle("POST /transfers", easypprof.Label("transfers", transfersHandler))

Then go tool pprof -tags cpu.pb.gz lists the label values with their share, and -tagfocus op=post_entry narrows the graph to one operation. Keep the values low-cardinality. Operation names are fine; account IDs are not, both because the profile becomes unreadable and because a profile is production data that will be copied to laptops.

Flight recorder on Go 1.25

Go 1.25 added a flight recorder to runtime/trace: a ring buffer of recent execution that you can snapshot after something odd happens, instead of having to be tracing before it happened. easypprof runs one when you configure it:

dbg, _ := easypprof.New(easypprof.Config{
    Tokens: tokens,
    FlightRecorder: &easypprof.FlightRecorderConfig{
        MinAge:      10 * time.Second,
        MaxBytes:    16 << 20,
        Dir:         "/var/tmp/traces",
        MinInterval: time.Minute,
    },
})

if elapsed > 250*time.Millisecond {
    go dbg.Snapshot("slow-post-entry")
}

A snapshot writes a trace file named with the timestamp and your label into Dir, created 0600 inside a 0700 directory, and MinInterval throttles how often that can happen so a latency storm does not fill the disk. go tool trace opens the result. You can also SnapshotTo any io.Writer, or pull the current buffer with GET /debug/pprof/flightrecorder. On older Go versions New returns ErrFlightRecorderUnsupported rather than silently doing nothing. If you want the background on what you are looking at in that trace, the Go scheduler post walks through the viewer.

What it deliberately does not do

It is not a continuous profiler. Pyroscope, Parca, Datadog and Cloud Profiler already do that well, and they can scrape easypprof’s endpoints with a token of their own. It does not export runtime metrics; that is Prometheus or OpenTelemetry’s job, and the OpenTelemetry setup in pREST v2.4.0 is how I do it. It does not do delta profiles; pull two and diff them with go tool pprof -diff_base before.pb.gz after.pb.gz. And it has no opinion on profile-guided optimization beyond the obvious one: a representative CPU profile pulled this way is a fine default.pgo.

Observability as a prerequisite rather than a phase is the argument I made in Go microservices in production; this package is the profiling part of that argument turned into a dependency.

Try it on the example ledger

The repo ships example/ledger, a small service with a deliberately hot lock, so you can see the mutex profile light up:

export PPROF_TOKEN=$(openssl rand -hex 24)
go run ./example/ledger &
for i in $(seq 1 2000); do curl -s -XPOST 'localhost:8080/transfers?from=a&to=b&amount=1' >/dev/null & done
curl -sH "Authorization: Bearer $PPROF_TOKEN" -o mutex.pb.gz localhost:6060/debug/pprof/mutex
go tool pprof -top mutex.pb.gz

The README also has a short checklist for profiling during a security test (baseline, active, post-test, with -diff_base on the heap profiles), which is where a capped, audited pprof earns its keep. Issues and pull requests go to github.com/arxdsilva/easypprof.

FAQ

Is it safe to leave pprof enabled in production?

With the standard net/http/pprof import on a public mux, no: anyone who can reach the service can pull profiles and run unbounded CPU profiles or traces. With a loopback listener, token authentication, a duration cap and a concurrency limit, the exposure is the same as any other authenticated admin endpoint, and the audit log tells you who used it.

Does easypprof change my application’s performance?

Serving the endpoints costs nothing until someone requests a profile, and each profile costs what pprof always costs. The one standing change is that block and mutex profiling are switched on at a low sampling rate (10µs block rate, 1 in 100 mutex events) so those profiles are not empty during an incident; both are restored on Shutdown.

How do I pull a profile from a pod?

Keep the loopback default and use kubectl port-forward pod/<name> 6060:6060, then curl localhost:6060 with your token. Access control is your cluster’s RBAC on port-forwarding plus the token.

What is the difference between easypprof and net/http/pprof?

Same profiles, different plumbing. net/http/pprof registers on DefaultServeMux at import time with no auth, no limits and no logging. easypprof uses its own listener or your mux, requires tokens off loopback, caps duration, limits concurrency, writes an audit record per request, hides /cmdline by default and enables block and mutex sampling.

Do I need Go 1.25?

No. The package requires Go 1.23 or newer. Only the flight recorder needs Go 1.25, and configuring it on an older toolchain returns ErrFlightRecorderUnsupported from New so the mistake is visible at startup.