Profiling & Performance Analysis¶
Learning Objectives
Apply the “measure, don’t guess” discipline and target the bottleneck that matters
Inspect a machine’s topology and bind threads to specific cores
Measure achieved memory bandwidth and cache behavior with hardware counters or timing
Find hotspots, cache misses, and memory errors with Valgrind
Place a kernel on the roofline model and decide whether it is compute- or bandwidth-bound
Measure, Don’t Guess¶
The most common performance mistake is optimizing the wrong thing. Amdahl’s Law is blunt: speeding up code that accounts for 5% of runtime can never give more than 5% improvement. Before you change anything, measure where time actually goes.
A disciplined workflow:
Time the whole program and its phases to find where the time is.
Profile the hot region to find why it is slow (counters, cache misses).
Classify the bottleneck with the roofline: compute-bound or bandwidth-bound?
Optimize the dominant cost.
Re-measure to confirm a real gain.
LIKWID — Hardware Performance Counters¶
LIKWID is a lightweight command-line toolkit for performance measurement on CPUs.
Inspect the Machine¶
likwid-topology -g
This prints the socket/core/cache layout. You need it to interpret everything else — to know how big L1/L2/L3 are, how many physical cores exist, and how hyperthreads and NUMA domains are arranged.
Measure Hardware Counters¶
List available performance groups:
likwid-perfctr -a
Then run your program pinned to a core under a group:
likwid-perfctr -C 0 -g MEM1 ./stream_triad
-C 0pins execution to core 0.-g MEM1selects the memory-bandwidth counter group.
The output reports bytes read/written and achieved bandwidth in GB/s. Compare to your CPU’s theoretical peak.
Pin Threads¶
For OpenMP code, where threads run dominates memory performance on NUMA systems:
likwid-pin -c 0-7 ./vecadd # pin 8 threads to cores 0-7
likwid-perfctr -C 0-7 -g MEM ./vecadd
Binding threads close to the memory they touch avoids cross-socket traffic.
Alternative: Linux perf¶
If LIKWID is not installed, use perf:
perf stat -e LLC-loads,LLC-load-misses ./program
Valgrind — Hotspots and Correctness¶
Where LIKWID reads real hardware counters, Valgrind runs your program on a synthetic CPU — slow, but needs no privileges.
Tool |
Purpose |
|---|---|
|
Find hotspots by instruction count |
|
Simulate cache and branch behavior |
|
Catch memory errors (leaks, overruns) |
|
Detect data races |
Basic Usage¶
# Find hotspots
valgrind --tool=callgrind ./program
# Check for memory errors
valgrind --tool=memcheck ./program
# Check for data races
valgrind --tool=helgrind ./program
Note
Valgrind slows execution 10–50×. Use it for debugging, not benchmarking.
The Roofline Model¶
The roofline turns measurements into decisions. It plots attainable performance against arithmetic intensity (FLOPs per byte):
Two ceilings:
A sloped bandwidth ceiling on the left.
A flat compute ceiling on the right.
Kernel |
FLOPs |
Bytes |
Intensity |
Regime |
|---|---|---|---|---|
Stream Triad |
\(2N\) |
\(24N\) |
\(\approx 0.08\) |
bandwidth-bound |
Dense matmul (\(N=1024\)) |
\(2N^3\) |
\(\approx 24N^2\) |
\(\approx N/12\) |
compute-bound |
Running on ARC Clusters¶
For LIKWID:
module load likwid
likwid-perfctr -C 0 -g MEM1 ./stream_triad
For Valgrind (workstation or login nodes):
valgrind --tool=memcheck ./program