npx skills add ...
npx skills add mohitmishra786/low-level-dev-skills --skill compiler-optimizations-deep
Deep compiler optimizations skill for RA, ISel, and PGO. Use when explaining register allocation, instruction selection, LICM, vectorization limits, or profile-guided optimization beyond -O3. Activates on queries about register allocation, instruction selection, LICM, auto-vectorization failure, PGO, or BOLT.
npx skills add mohitmishra786/low-level-dev-skills --skill compiler-optimizations-deep
Explain optimization phases beyond flags: mid-level IR opts, register allocation, instruction selection/scheduling, vectorization boundaries, PGO, and post-link BOLT — bridging skills/compilers/pgo and LLVM/GCC internals.
-O3 did not vectorize a hot loop| Miss reason | Typical fix |
|---|---|
| Unknown trip count | peel loop; assert count |
| Dependence | reorder / separate accumulators |
| Function call in loop | inline or outline |
| Alignment unknown | __builtin_assume_aligned |
When live ranges exceed physical registers, the allocator spills to stack slots — costly loads/stores. Reducing live ranges (splitting variables, rematerialization) helps.
GCC/LLVM both use graph coloring variants (LLVM "greedy regalloc").
Improves branch layout, inlining, and vectorization thresholds.
See skills/compilers/pgo for GCC and BOLT.
Optimizes layout after linker — needs relocations (-Wl,--emit-relocs).
Loop-invariant code motion hoists x * scale out of inner loop when legal — reduces work per iteration.
| Symptom | Cause | Fix |
|---|---|---|
| PGO no gain | Unrepresentative training | Match production input |
| BOLT crash | Stripped binary | Keep symbols + relocs |
| Spills in asm | Register pressure | Simplify live ranges |
-O3 slower | Code bloat / cache | Try -O2 or PGO |
| Different GCC/Clang | Pass ordering differs | Compare IR + asm |
skills/compilers/pgo — PGO and BOLT detailskills/compiler-internals/llvm-ir-and-passes — IR-level optsskills/compiler-internals/code-generation-and-backends — ISel and backendsskills/computer-architecture/cpu-pipelines-and-hazards — scheduling contextskills/low-level-programming/simd-intrinsics — manual vectorizationclang -fprofile-instr-generate -O2 -o app foo.c
./app # training workload
llvm-profdata merge default.profraw -o default.profdata
clang -fprofile-instr-use=default.profdata -O2 -o app_pgo foo.cllvm-bolt -instrument app -o app.inst
./app.inst
llvm-bolt -data=perf.fdata -reorder-blocks=+ -o app.bolt app/compiler-optimizations-deep Why did LLVM fail to vectorize this reduction loop?