Zurück zu den Modellen
Clusterer

SklearnClustererPipeline

Missing valuesPredictOut of sample

Pipeline of transformers and a clusterer.

The SklearnClustererPipeline chains transformers and an single clusterer.

Similar to ClustererPipeline, but uses a tabular sklearn clusterer.

The pipeline is constructed with a list of sktime transformers, plus a clusterer,

i.e., transformers following the BaseTransformer interface, clusterer follows the scikit-learn clusterer interface.

The transformer list can be unnamed - a simple list of transformers -

or string named - a list of pairs of string, estimator.

For a list of transformers trafo1, trafo2, …, trafoN and a clusterer clst,

the pipeline behaves as follows:

fit(X, y) - changes styte by running trafo1.fit_transform on X,

them trafo2.fit_transform on the output of trafo1.fit_transform, etc sequentially, with trafo[i] receiving the output of trafo[i-1], and then running clst.fit with X the output of trafo[N] converted to numpy, and y identical with the input to self.fit. X is converted to numpyflat mtype if X is of Panel scitype; X is converted to numpy2D mtype if X is of Table scitype.

predict(X) - result is of executing trafo1.transform, trafo2.transform, etc

with trafo[i].transform input = output of trafo[i-1].transform, then running clst.predict on the numpy converted output of trafoN.transform, and returning the output of clst.predict. Output of trasfoN.transform is converted to numpy, as in fit.

predict_proba(X) - result is of executing trafo1.transform, trafo2.transform,

etc, with trafo[i].transform input = output of trafo[i-1].transform, then running clst.predict_proba on the output of trafoN.transform, and returning the output of clst.predict_proba. Output of trasfoN.transform is converted to numpy, as in fit.

get_params, set_params uses sklearn compatible nesting interface

if list is unnamed, names are generated as names of classes if names are non-unique, f”_{str(i)}” is appended to each name string

where i is the total count of occurrence of a non-unique string inside the list of names leading up to it (inclusive)

SklearnClustererPipeline can also be created by using the magic multiplication
between sktime transformers and sklearn clusterers,

and my_trafo1, my_trafo2 inherit from BaseTransformer, then, for instance, my_trafo1 * my_trafo2 * my_clst will result in the same object as obtained from the constructor SklearnClustererPipeline(clusterer=my_clst, transformers=[t1, t2])

magic multiplication can also be used with (str, transformer) pairs,

as long as one element in the chain is a transformer

Schnellstart

python
from sktime.clustering.compose import SklearnClustererPipeline

estimator = SklearnClustererPipeline(clusterer, transformers)

Parameter(2)

clusterersklearn clusterer, i.e., inheriting from sklearn ClustererMixin
this is a “blueprint” clusterer, state does not change when fit is called
transformerslist of sktime transformers, or
list of tuples (str, transformer) of sktime transformers these are “blueprint” transformers, states do not change when fit is called

Beispiele

>>> from sklearn.cluster import KMeans
>>> from sktime.transformations.exponent import ExponentTransformer
>>> from sktime.transformations.summarize import SummaryTransformer
>>> from sktime.datasets import load_unit_test
>>> from sktime.clustering.compose import SklearnClustererPipeline
>>> X_train, y_train = load_unit_test (split = "train")
>>> X_test, y_test = load_unit_test (split = "test")
>>> t1 = ExponentTransformer ()
>>> t2 = SummaryTransformer ()
>>> pipeline = SklearnClustererPipeline (KMeans (), [t1, t2 ])
>>> pipeline = pipeline. fit (X_train, y_train)
>>> y_pred = pipeline. predict (X_test) Alternative construction via dunder method:
>>> pipeline = t1 * t2 * KMeans ()