Choose an algorithm¶
Squeeze has eleven core reduction methods. Compare an inexpensive linear baseline with methods matching the structure you care about, then measure quality and runtime. The API differences below are part of the contract.
| Method | Implementation | Fit new coordinates | Native transform | Main purpose |
|---|---|---|---|---|
| PCA | Rust | fit_transform; fit |
Yes | Linear variance baseline |
| UMAP | Python/Numba + optional Rust search | fit_transform; fit |
Yes, mode-dependent | Neighbor graph layout |
| t-SNE | Rust | fit_transform |
No | Local probability layout |
| MDS | Rust | fit_transform; from distances |
No | Distance fitting |
| Isomap | Rust | fit_transform |
No | Graph geodesics |
| LLE | Rust | fit_transform |
No | Local linear relations |
| PHATE | Rust | fit_transform |
No | Diffusion geometry |
| TriMap | Rust | fit_transform |
No | Triplet constraints |
| PaCMAP | Rust | fit_transform |
No | Pairwise layout |
| NeighborMap | Experimental Rust | fit_transform |
No | Sampled graph refinement |
| SpectralMap | Experimental Rust | fit_transform |
No | Approximate spectral layout |
NeighborMap and SpectralMap are Euclidean, two-dimensional methods. The other
constructors expose n_components, but that does not establish identical behavior
or performance at every dimension. Rust reducers do not implement sklearn's full
parameter/cloning interface. Composition wrappers have
additional limits when they contain fit-only components.
Start from your requirement¶
- For a reusable linear projection, use PCA as a baseline.
- For neighborhood visualization, compare UMAP, t-SNE, PaCMAP and NeighborMap.
- For global distances or graph geometry, inspect MDS, Isomap, LLE and PHATE.
- For an inexpensive approximate graph layout, evaluate SpectralMap's quality tradeoff.
There is no supported claim here that every Rust method is faster than every alternative. The Digits and Fashion-MNIST heatmaps show measured configurations, including weak results for the current TriMap implementation. Do not infer large-dataset scalability from these small benchmarks.