P3 Analysis#
ppbcc p2analysis and ppbcc p3analysis turn benchmark CSVs into
application-efficiency and performance-portability figures. p2analysis
renders the charts that need benchmark results only; p3analysis joins a
code-complexity CSV and renders the charts that set portability against
complexity. The chart gallery lives in Available Plots;
this page covers the metrics and the data selection behind them.
The metrics implemented here are those of Pennycook et al. [Pennycook2019] [Pennycook2021], and the Cascade and Navchart layouts follow the P3 Analysis Library that accompanies that work. See References.
ppbcc p2analysis PLOT CSV [CSV ...] [options]
ppbcc p3analysis PLOT COMPLEXITY_CSV CSV [CSV ...] [options]
-n/--name selects the benchmark problem. An exact case-insensitive match
wins; a unique substring is accepted, so Polyhedral resolves to
PolyhedralGravity. It may be omitted when the benchmark CSVs contain exactly
one problem. The one output that spans problems rather than picking one is
rank-correlation: it takes a comma-separated list of at least two problems,
defaults to every problem in the CSVs, and reports how far their paradigm
orderings agree (see Available Plots).
Metrics#
Application efficiency#
For every hardware platform and workload, an implementation’s runtime is compared against the best runtime observed for that platform and workload:
Duplicate measurements are reduced to their median runtime first, and the per-workload efficiencies are then averaged arithmetically.
Performance portability \(\Phi\)#
\(\Phi\) is the harmonic mean of the application efficiencies over the set of platforms \(H\) [Pennycook2019]:
By default, a platform an implementation does not support contributes zero and
therefore drives \(\Phi\) to zero — the strict reading of the original
definition, which rewards implementations that run everywhere.
--non-zero-pp restricts the harmonic mean to the platforms that actually
produced a result; the application-efficiency output still retains the zeros,
so the heatmap and the Cascade platform ranking stay honest about the gaps.
Note
Reporting \(\Phi\) without naming the platform set \(H\) is
meaningless — the same implementation scores very differently over
“all NVIDIA GPUs” than over “every platform measured”. Whenever you quote a
number from these charts, quote the platform set and whether
--non-zero-pp was used along with it.
Selecting data#
Problem sizes#
-s/--size controls how the size axis is collapsed:
Value |
Behaviour |
|---|---|
|
Use every measured size |
exact number |
Restrict everything to that one size |
|
Compute both metrics independently per size, then average arithmetically — every size gets equal weight |
|
Take the per-application maximum / minimum over sizes |
--average-over decides what avg averages for \(\Phi\). The default
pp computes \(\Phi\) at every size and takes the arithmetic mean of the
scores. efficiency treats the size sweep as one benchmark: it averages every
application efficiency over the sizes and computes \(\Phi\) once, as the
harmonic mean of those averages — so the \(\Phi\) panel is exactly the
harmonic mean of the efficiency panel next to it. The two differ whenever an
implementation’s best platform changes with the size, because an arithmetic
mean of harmonic means is not a harmonic mean. Application efficiency and the
per-size scaling heatmap are identical in both modes, and the option is
rejected for every -s other than avg.
Boxplots accept only all or an exact numeric size, because the summary
modes remove the very distribution the boxplot displays. In the combined
chart, the scaling panel always uses all sizes regardless of -s.
Filtering implementations#
-i/--include and -x/--exclude are regular expressions matched against
the Description column. Excluding vendor-optimised references is a common
move, since they otherwise define the efficiency baseline:
ppbcc p2analysis cascade ./Results_* -n MatrixMultiplication -x "Cublas"
--remove-description drops bracketed labels such as [Naive] from
labels and legends; efficiency charts then combine variants by paradigm.
-p/--precision keeps only 32- or 64-bit results.
Joining code complexity#
p3analysis takes the implementation-level CSV described in
Applied example: analysing a whole benchmark suite. Its Name/Framework columns are
matched against the benchmark problem and paradigm labels, and its raw Halstead
counts (n1, n2, N1, N2) are used to derive the metric selected
by -c/--complexity-metric:
Metric |
Accepted aliases |
Definition |
|---|---|---|
|
— |
source lines of code |
|
|
\(n_1 + n_2\) |
|
|
\(N_1 + N_2\) |
|
|
\(N \log_2 \eta\) |
|
|
\(\frac{n_1}{2} \cdot \frac{N_2}{n_2}\) |
|
|
\(D \cdot V\) |
Matching is case-insensitive and treats _, -, and spaces alike, so
Halstead Difficulty and halstead_difficulty both work.
Complexity is put into perspective relative to the plain-C++ implementation
by default: every score is divided by the CPP score, so CPP sits at 100 %.
--complexity-metric-absolute plots the raw scores instead.
CSV export#
-e/--export-to-csv writes application-efficiency and
performance-portability data next to the plot as
<plot-prefix>_application_efficiency.csv and
<plot-prefix>_performance_portability.csv. The export keeps separate rows
per size and precision plus averaged rows, and includes every available
complexity metric for p3analysis — useful when you want the
underlying numbers rather than the figure.
Python API#
from pathlib import Path
from ppbcc.performance_portability import (
calculate_metrics,
load_benchmark_csvs,
select_problem_rows,
)
frame = load_benchmark_csvs([Path("Results_NVIDIA_RTX5080.csv")])
rows, title, description_is_workload = select_problem_rows(
frame,
"MatrixMultiplication",
description_include=None,
description_exclude="Cublas",
problem_size=None,
precision=None,
)
efficiency, portability = calculate_metrics(
rows, description_is_workload, non_zero_pp=True
)
See API — ppbcc.performance_portability for the full signatures.
References#
The \(\Phi\) metric implemented here was introduced in [Pennycook2019];
the Cascade and Navchart visualizations, and the productivity dimension they
add, come from [Pennycook2021]. Both are the work behind the
P3 Analysis Library, which
inspired this part of ppbcc — please have a look at it, and cite the papers
below rather than this tool when you report performance-portability results.
S. J. Pennycook, J. D. Sewall, and V. W. Lee, “Implications of a Metric for Performance Portability,” Future Generation Computer Systems, vol. 92, pp. 947–958, Mar. 2019, doi: 10.1016/j.future.2017.08.007.
S. J. Pennycook, J. D. Sewall, D. W. Jacobsen, T. Deakin, and S. McIntosh-Smith, “Navigating Performance, Portability, and Productivity,” Computing in Science & Engineering, vol. 23, no. 5, pp. 28–38, Sep. 2021, doi: 10.1109/MCSE.2021.3097276.
The complexity axis of the Navchart rests on a separate body of work; see References.