TimeSeriesKMeans
TimeSeriesKMeans
- class TimeSeriesKMeans(n_clusters: int = 8, init_algorithm: str | Callable = 'random', metric: str | Callable = 'dtw', n_init: int = 10, max_iter: int = 300, tol: float = 1e-06, verbose: bool = False, random_state: int | RandomState = None, averaging_method: str | Callable[[ndarray], ndarray] = 'mean', distance_params: dict = None, average_params: dict = None)[source]
Time series K-mean implementation.
- Parameters:
- n_clusters: int, defaults = 8
The number of clusters to form as well as the number of centroids to generate.
- init_algorithm: str, np.ndarray (3d array of shape (n_clusters, n_dimensions,
series_length)), defaults = ‘random’ Method for initializing cluster centers or an array of initial cluster centers. If string, any of the following strings are valid:
[‘kmeans++’, ‘random’, ‘forgy’].
- If 3D np.ndarray, initializes cluster centers with the provided array. The array
must have shape (n_clusters, n_dimensions, series_length) and the number of clusters in the array must be the same as what is provided to the n_clusters argument.
- metric: str or Callable, defaults = ‘dtw’
Distance metric to compute similarity between time series. Any of the following are valid: [‘dtw’, ‘euclidean’, ‘erp’, ‘edr’, ‘lcss’, ‘squared’, ‘ddtw’, ‘wdtw’, ‘wddtw’]
- n_init: int, defaults = 10
Number of times the k-means algorithm will be run with different centroid seeds. The final result will be the best output of n_init consecutive runs in terms of inertia.
- max_iter: int, defaults = 300
Maximum number of iterations of the k-means algorithm for a single run.
- tol: float, defaults = 1e-6
Relative tolerance with regards to Frobenius norm of the difference in the cluster centers of two consecutive iterations to declare convergence.
- verbose: bool, defaults = False
Verbosity mode.
- random_state: int or np.random.RandomState instance or None, defaults = None
Determines random number generation for centroid initialization.
- averaging_method: str or Callable, defaults = ‘mean’
Averaging method to compute the average of a cluster. Any of the following strings are valid: [‘mean’, ‘dba’]. If a Callable is provided must take the form Callable[[np.ndarray], np.ndarray].
- average_params: dict, defaults = None = no parameters
Dictionary containing kwargs for averaging_method.
- distance_params: dict, defaults = None = no parameters
Dictionary containing kwargs for the distance metric being used.
- Attributes:
- cluster_centers_: np.ndarray (3d array of shape (n_clusters, n_dimensions,
series_length)) Time series that represent each of the cluster centers. If the algorithm stops before fully converging these will not be consistent with labels_.
- labels_: np.ndarray (1d array of shape (n_instance,))
Labels that is the index each time series belongs to.
- inertia_: float
Sum of squared distances of samples to their closest cluster center, weighted by the sample weights if provided.
- n_iter_: int
Number of iterations run.
Examples
>>> from sktime.datasets import load_arrow_head >>> from sktime.clustering.k_means import TimeSeriesKMeans >>> X_train, y_train = load_arrow_head(split="train") >>> X_test, y_test = load_arrow_head(split="test") >>> clusterer = TimeSeriesKMeans(n_clusters=3) >>> clusterer.fit(X_train) TimeSeriesKMeans(n_clusters=3) >>> y_pred = clusterer.predict(X_test)
Methods
check_is_fitted([method_name])Check if the estimator has been fitted.
clone()Obtain a clone of the object with same hyper-parameters and config.
clone_tags(estimator[, tag_names])Clone tags from another object as dynamic override.
create_test_instance([parameter_set])Construct an instance of the class, using first test parameter set.
create_test_instances_and_names([parameter_set])Create list of all test instances and a list of names for them.
fit(X[, y])Fit time series clusterer to training data.
fit_predict(X[, y])Compute cluster centers and predict cluster index for each time series.
get_class_tag(tag_name[, tag_value_default])Get class tag value from class, with tag level inheritance from parents.
get_class_tags()Get class tags from class, with tag level inheritance from parent classes.
get_config()Get config flags for self.
get_fitted_params([deep])Get fitted parameters.
get_param_defaults()Get object's parameter defaults.
get_param_names([sort])Get object's parameter names.
get_params([deep])Get a dict of parameters values for this object.
get_tag(tag_name[, tag_value_default, ...])Get tag value from instance, with tag level inheritance and overrides.
get_tags()Get tags from instance, with tag level inheritance and overrides.
get_test_params([parameter_set])Return testing parameter settings for the estimator.
is_composite()Check if the object is composed of other BaseObjects.
load_from_path(serial)Load object from file location.
load_from_serial(serial)Load object from serialized memory container.
predict(X[, y])Predict the closest cluster each sample in X belongs to.
predict_proba(X)Predicts labels probabilities for sequences in X.
reset()Reset the object to a clean post-init state.
save([path, serialization_format])Save serialized self to bytes-like object or to (.zip) file.
score(X[, y])Score the quality of the clusterer.
set_config(**config_dict)Set config flags to given values.
set_params(**params)Set the parameters of this object.
set_random_state([random_state, deep, ...])Set random_state pseudo-random seed parameters for self.
set_tags(**tag_dict)Set instance level tag overrides to given values.

