How do you adapt scientific algorithms to parallel, multi-core HPC systems? The first step is building an empirical performance model. This post summarizes the required hardware background—from instruction cycles and cache locality to data-level and instruction-level parallelism—based on Anthony Joseph’s 2009 thesis.