Skip to content

Streaming observations

StreamingDR embeds the initial batch with a base reducer. Later batches use nearest-neighbor interpolation; they do not retrain the base algorithm or optimize old coordinates again.

import numpy as np
from sklearn.datasets import load_digits
from squeeze import PCA, StreamingDR

X = np.asarray(load_digits().data[:200], dtype=np.float64)
model = StreamingDR(PCA(n_components=2), n_neighbors=5)
initial = model.fit_transform(X[:100])
updated = model.partial_fit_transform(X[100:150])
assert initial.shape == (100, 2)
assert updated.shape == (150, 2)  # All rows accumulated so far.
preview = model.transform(X[150:])
assert preview.shape == (50, 2)   # Does not append these rows.
assert model.embedding_.shape == (150, 2)

partial_fit appends the new rows and their interpolated coordinates, then rebuilds the nearest-neighbor index. The wrapper retains all observations in X_all_, so memory grows with the stream. It is not a bounded-memory online optimizer. For drift or a new population, evaluate whether a complete refit is more appropriate.