One sample is an anecdote

Application benchmarks run on operating systems that schedule many tasks. A single sample can be disturbed by background work, allocation, cache state or virtual-machine activity. Repeating the operation does not guarantee truth, but it makes instability visible.

ASAMB uses the median as the center because one unusually slow sample cannot pull it as strongly as it pulls an arithmetic mean. The 25th and 75th percentiles describe the middle half of observed samples.

Spread changes the confidence of a ranking

When two medians are close and their sample ranges overlap heavily, the honest conclusion is that the local difference is small or unstable. A large percentage shown to two decimal places is not automatically important if the operation contributes almost nothing to end-to-end latency.

Operational decisions should combine effect size, call frequency, correctness constraints and the cost of extra complexity.

Continue reading

Full benchmark methodology Browse measured studies