Skip to content

Evaluation metrics

Measure several properties of an embedding. No single score captures neighborhood fidelity, global geometry, cluster separation, and runtime.

Rank trustworthiness versus neighbor overlap

The refreshed heatmaps use scikit-learn rank trustworthiness. It penalizes embedded neighbors that were far away in the original ranking. Rank continuity is computed by reversing the original and embedded spaces.

The public squeeze.trustworthiness, squeeze.continuity, and squeeze.co_ranking_quality functions currently use neighborhood-overlap calculations. Those names are retained for compatibility. quick_evaluate and DREvaluator inherit those definitions. Do not compare their values directly with rank trustworthiness as if they were the same metric.

import numpy as np
from sklearn.datasets import load_digits
from sklearn.manifold import trustworthiness as rank_trustworthiness
from squeeze import PCA, trustworthiness as neighbor_overlap

X = np.asarray(load_digits().data[:100], dtype=np.float64)
Y = PCA(n_components=2).fit_transform(X)
rank_score = rank_trustworthiness(X, Y, n_neighbors=15)
recall = neighbor_overlap(X, Y, k=15)
assert 0 <= rank_score <= 1 and 0 <= recall <= 1
print({"rank_trustworthiness": rank_score, "neighbor_recall": recall})

Reading the heatmap

Metric Interpretation and limitation
Rank trustworthiness, k=5/15/30; continuity, k=15 Local rank penalties in each direction; higher is better
Neighbor recall, k=15 Fraction of original neighbors recovered; higher is better
Spearman Correlation of pairwise distances; higher is better
Global structure Correlation of distances between class centroids; requires labels
Density Correlation of local density estimates
Reconstruction R² Linear reconstruction from coordinates; not a reducer inverse
Silhouette Separation of KMeans-assigned clusters in the embedding
ARI / NMI Agreement between KMeans assignments and supplied labels
Transductive accuracy RandomForest cross-validation on embeddings already fitted to all rows
Time Full fit after warmup; lower is better; metric computation excluded

Heatmap colors are normalized within each column. Time uses a reversed log scale; cell labels retain the raw values. A green cell is relative to the displayed methods.

Evaluation API

DREvaluator(X_original, X_reduced, labels=None, reducer=None, method_name="Unknown") provides local/global structure, reconstruction, clustering, classification, and optional stability/noise evaluations. evaluate_all(include_stability=False) avoids repeated stability fits. Its EvaluationReport has summary() and to_dict() methods. API details describe the exported metric functions.

For predictive evaluation, split data before fitting the embedding. Cross-validation of a classifier on a full-data embedding is transductive, even when no labels were used to fit the embedding itself.