How the Go Scheduler Works: G, M and P Explained With What You See in pprof

The Go scheduler multiplexes goroutines (G) onto OS threads (M) through a fixed number of logical processors (P), each with a local run queue, backed by a global queue, work stealing between Ps, and handoff of the P when a thread blocks in a syscall. This post is the detail behind each clause, checked against the runtime source, and what each piece looks like in go tool trace and pprof. It expands my talk at the first Golang Calgary meetup.

Table of Contents

The three letters

The comment at the top of runtime/proc.go defines the model in three lines: G is a goroutine, M is a worker thread (or machine), and P is a processor, “a resource that is required to execute Go code”. An M needs a P to run Go code, but can block or sit in a syscall without one. The design is Dmitry Vyukov’s go11sched doc.

Letter What it is How many
G A goroutine: a stack (2 KB to start, stackMin = 2048 in stack.go), a program counter and scheduling state As many as you create
M An OS thread managed by the runtime Grows with demand, capped at 10,000 by default (sched.maxmcount)
P A logical processor: a 256-slot run queue, a runnext slot, a timer heap, an allocation cache Exactly GOMAXPROCS

P exists so that scheduler state is, in the comment’s words, “intentionally distributed (in particular, per-P work queues)” and most scheduling decisions take no global lock.

Run queues: local, global and runnext

Each P has a local run queue, a ring of 256 goroutine pointers (runq [256]guintptr in runtime2.go) that only its owner pushes to; other Ps pop from it when they steal. There is also one global queue, sched.runq, behind sched.lock.

Then there is runnext, the part people miss: one extra slot per P for a goroutine made runnable by the goroutine currently running. The comment says it “should be run next instead of what’s in runq if there’s time remaining in the running G’s time slice” and “will inherit the time left in the current time slice”. The reason is locality: a communicate-and-wait pair gets scheduled as a unit. Two things put a goroutine there: a go statement (newproc calls runqput(pp, newg, true)) and waking a parked goroutine, for example by sending on a channel it was receiving from (goready calls ready(gp, traceskip, true)).

When the local queue is full, half of it (plus the new goroutine) moves to the global queue in one batch; when it is empty, findRunnable pulls a batch back. Every 61st scheduling tick checks the global queue first, because otherwise “two goroutines can completely occupy the local runqueue by constantly respawning each other”.

What happens when a goroutine blocks

This drew the most questions at the talk, because blocking means three different things to the scheduler.

A blocking syscall

When a goroutine enters a syscall, the G and its M go into the kernel together and the P stays attached for now. If the syscall returns quickly, exitsyscall takes the fast path: the M still has its P and nothing else happens.

If it takes long, sysmon notices. Its retake loop checks each P, and if one has been in the same syscall for more than one sysmon tick (at least 20 microseconds) and there is other work, it takes the P away and calls handoffp, which gives it to another M when there is work for it. When the syscall returns, the M tries to reacquire its old P, then any idle P; if neither is free, the goroutine goes to the global queue and the M parks. The thread blocks, the processor does not.

A network read

A read on a net.Conn with nothing to read does not block a thread. The netpoller wraps epoll, kqueue and their equivalents (listed at the top of netpoll.go). The goroutine parks on the fd’s poll descriptor, the M and P move on, and when the kernel reports the fd ready, netpoll hands the goroutine back to a run queue. A parked goroutine costs its stack and no thread, which is how a Go server holds thousands of idle connections on a few threads.

A channel operation

A receive on an empty channel, a send on a full one, a held sync.Mutex, or a select with nothing ready all go through gopark. The goroutine is recorded on the channel’s wait queue and the M picks the next goroutine from the P; no thread blocks and no P changes hands. When the other side arrives, goready puts it in runnext, so a tight ping-pong between two goroutines shares one time slice until sysmon preempts them.

Preemption: cooperative to signals

Until Go 1.14 the scheduler could only preempt at a function call, because the check lived in the function prologue, so a loop with no calls could not be preempted. The Go 1.14 release notes changed that: “Goroutines are now asynchronously preemptible. As a result, loops without function calls no longer potentially deadlock the scheduler or significantly delay garbage collection.”

The mechanism is a signal: on Unix the runtime sends SIGURG to the thread running the goroutine (sigPreempt = _SIGURG in signal_unix.go), and the handler switches it out at a safe point. The same notes warn that programs now see more EINTR errors from slow syscalls and must retry them.

The time slice is 10 ms: forcePreemptNS in proc.go is “the time slice given to a G before it is preempted”, and retake calls preemptone on any P that has run past it.

Work stealing and spinning threads

When a P finds nothing locally, in the global queue or in the netpoller, its M becomes a spinning thread and calls stealWork. It makes up to four passes over the other Ps in random order (stealTries = 4), skipping idle ones, and runqsteal takes half of a victim’s local queue. Only on the last pass does it also take timers and the victim’s runnext, which the comment calls “the last resort”.

Spinning is bounded: findRunnable lets an M spin only if 2*nmspinning < gomaxprocs - npidle, limiting spinning Ms “to half the number of busy Ps”.

The design comment explains why spinning exists at all: handing a readied goroutine straight to a woken thread was rejected because that thread “can be out of work the very next moment” and it “would destroy locality of computation”. In a CPU profile, spinning shows up as runtime.findRunnable, runtime.stealWork and runtime.runqsteal; a few percent is normal, a large share means Ps keep starving and then fighting for scraps.

sysmon, the thread without a P

sysmon is a runtime thread that runs without a P and never executes user code. Its loop sleeps 20 microseconds between iterations, starts doubling after 50 idle iterations, and caps at 10 ms.

Each iteration it polls the network if nobody has for over 10 ms, runs retake (handing off Ps stuck in syscalls and preempting goroutines past their slice), forces a GC if none ran in two minutes, and, since Go 1.25, re-checks the GOMAXPROCS default about once a second.

sysmon is also what makes runnext safe: two goroutines sharing a slice would starve everyone else without preemption, so on platforms with no sysmon the runtime does not use runnext at all.

GOMAXPROCS and containers

GOMAXPROCS is the number of Ps. The runtime docs describe it as limiting “the number of operating system threads that can execute user-level Go code simultaneously”, and threads blocked in syscalls “do not count against the GOMAXPROCS limit”.

For years the default was runtime.NumCPU(). In a container with a CPU limit of 2 on a 64-core node, that meant 64 Ps sharing two CPUs’ worth of quota and the kernel throttling the process whenever it ran out. The usual fix was Uber’s automaxprocs, which sets GOMAXPROCS from the container CPU quota.

Go 1.25 changed the default. From the Go 1.25 release notes: “On Linux, the runtime considers the CPU bandwidth limit of the cgroup containing the process, if any. If the CPU bandwidth limit is lower than the number of logical CPUs available, GOMAXPROCS will default to the lower limit.” This maps to the Kubernetes CPU limit; the runtime “does not consider the ‘CPU requests’ option”. It also “periodically updates GOMAXPROCS” when the limit changes. Both switch off if you set GOMAXPROCS yourself, or with GODEBUG=containermaxprocs=0 and GODEBUG=updatemaxprocs=0.

So on Go 1.25 or later, in a container with a CPU limit, you probably do not need automaxprocs. With requests only and no limit, the runtime still sees every core on the node. More on that in Go microservices in production.

What you see in pprof and go tool trace

A CPU profile tells you where the program burns cycles, not why a goroutine waited 40 ms for a P. For that you need the execution tracer, “a tool to detect latency and utilization problems” in the words of the Go diagnostics guide.

This program starts a trace, launches 64 CPU-bound goroutines and collects their results.

package main

import (
	"log"
	"os"
	"runtime/trace"
	"sync"
)

func burn(n int) int {
	s := 0
	for i := 0; i < n; i++ {
		s += i % 7
	}
	return s
}

func main() {
	f, err := os.Create("trace.out")
	if err != nil {
		log.Fatal(err)
	}
	defer f.Close()

	if err := trace.Start(f); err != nil {
		log.Fatal(err)
	}
	defer trace.Stop()

	results := make(chan int, 64)
	var wg sync.WaitGroup
	for i := 0; i < 64; i++ {
		wg.Add(1)
		go func() {
			defer wg.Done()
			results <- burn(20_000_000)
		}()
	}
	wg.Wait()
	close(results)

	total := 0
	for r := range results {
		total += r
	}
	log.Println("total:", total)
}
go run . && go tool trace trace.out

The viewer’s landing page is documented in traceviewer/http.go. The parts I use:

  • “View trace by proc” draws one lane per P, “showing which goroutine (if any) was running on that logical processor at each moment”. With 64 goroutines and 8 Ps you see 8 lanes filled with short bars; gaps are idle Ps. That is your parallelism, read directly.
  • “Goroutine analysis” groups goroutines by start location and splits each one’s time into execution, “Sched wait time” (runnable but waiting for a P), syscall time, and “Block time” by reason. If sched wait dominates, you have more runnable goroutines than Ps.
  • Four profiles (network blocking, synchronization blocking, syscall, scheduler latency) cover “causes of delay that prevent a goroutine from running on a logical processor”. go tool trace -pprof=sched trace.out > sched.pprof writes one for pprof. For the program above it is tiny; swap burn for I/O and the sync and net profiles carry the story.

Tracing is cheap enough for short windows in production: the Go team’s post on execution traces puts the overhead at “1–2% for many applications” since Go 1.21, and net/http/pprof serves live captures at /debug/pprof/trace. Go 1.25 added runtime/trace.FlightRecorder, a ring buffer you snapshot after a slow request.

In a plain CPU profile the scheduler appears as runtime.schedule, runtime.findRunnable, runtime.mcall and runtime.futex (Ms parking and unparking on Linux). The goroutine profile lists every goroutine with its wait reason in brackets, like [chan receive], [IO wait] or [syscall].

Common misreadings

More goroutines do not mean more parallelism, which is capped at GOMAXPROCS. Ten thousand goroutines doing CPU-bound work on 8 Ps gives you 8 running, 9,992 runnable, and the memory for 10,000 stacks; a worker pool the size of GOMAXPROCS is usually right. For I/O-bound work, many goroutines are fine, since nearly all are parked.

Raising GOMAXPROCS above the CPU count does not speed up CPU-bound code; the extra Ps share the same cores and now also contend. It can help with heavy cgo or blocking-syscall work, and handoff usually covers that already.

GOMAXPROCS=1 does not make code free of data races. Goroutines still interleave at every channel or mutex operation and, since 1.14, at almost any instruction.

FAQ

What is the difference between G, M and P?

G is a goroutine, M is an OS thread managed by the runtime, and P is a logical processor that an M must hold to run Go code. There are as many Gs as you create, as many Ms as blocking work demands, and exactly GOMAXPROCS Ps.

How many OS threads does Go use?

At least one per P plus sysmon, and more when goroutines block in syscalls or cgo, because the runtime hands the P to another thread while the blocked one stays in the kernel. The default cap is 10,000; debug.SetMaxThreads changes it.

Is the Go scheduler preemptive?

Yes, since Go 1.14, whose release notes state goroutines are “asynchronously preemptible”, implemented on Unix by sending SIGURG to the thread. Before 1.14 preemption only happened at function calls, so a loop with no calls could hold a P indefinitely.

What does GOMAXPROCS do?

It sets the number of Ps, which is the maximum number of threads running Go code at once; threads blocked in syscalls do not count. Since Go 1.25 the default on Linux respects the cgroup CPU limit; before that it was the logical CPU count, which oversubscribed CPU-limited containers.

Why is my CPU not fully used with many goroutines?

Because most of them are waiting on something other than the CPU: a lock, a channel, the network or a syscall. go tool trace shows this as idle gaps in the per-P lanes and as block time by cause. If they are runnable but not running, check GOMAXPROCS against the container’s CPU limit.