Fashion-MNIST · 2,000 samples

Algorithm performance radar · farther from the centre means better performance. Choose methods to compare. Hover or focus a point for its measured value.

Algorithm performance radar Six relative-performance axes, zero at the centre and one at the outer ring. Exact measurements are in the table below.

Hover or focus a chart point to inspect its value.

Measured values

The chart scales each axis from the worst to the best result across all algorithms in this dataset. Changing the selection does not rescale it. Zero is the worst observed result, not absence of capability. Runtime uses reversed log scaling, so faster methods extend farther. Tied axes sit at 0.5.

Single-seed benchmark snapshots. Shapes and polygon areas are not overall scores. Axes are normalized independently for each dataset; compare raw metrics across datasets. Transductive accuracy evaluates a classifier on coordinates fitted to all samples, not a held-out embedding.

Benchmark protocol and metric definitions