Skip to content

ShapeletTransformClassifier

ShapeletTransformClassifier

class ShapeletTransformClassifier(n_shapelet_samples=10000, max_shapelets=None, max_shapelet_length=None, estimator=None, transform_limit_in_minutes=0, time_limit_in_minutes=0, contract_max_n_shapelet_samples=inf, save_transformed_data=False, n_jobs=1, batch_size=100, random_state=None)[source]

A shapelet transform classifier (STC).

Implementation of the binary shapelet transform classifier pipeline along the lines of [1][2] but with random shapelet sampling. Transforms the data using the configurable RandomShapeletTransform and then builds a RotationForest classifier.

As some implementations and applications contract the transformation solely, contracting is available for the transform only and both classifier and transform.

Parameters:
n_shapelet_samplesint, default=10000

The number of candidate shapelets to be considered for the final transform. Filtered down to <= max_shapelets, keeping the shapelets with the most information gain.

max_shapeletsint or None, default=None

Max number of shapelets to keep for the final transform. Each class value will have its own max, set to n_classes_ / max_shapelets. If None, uses the minimum between 10 * n_instances_ and 1000.

max_shapelet_lengthint or None, default=None

Lower bound on candidate shapelet lengths for the transform. If None, no max length is used

estimatorBaseEstimator or None, default=None

Base estimator for the ensemble, can be supplied a sklearn BaseEstimator. If None a default RotationForest classifier is used.

transform_limit_in_minutesint, default=0

Time contract to limit transform time in minutes for the shapelet transform, overriding n_shapelet_samples. A value of 0 means n_shapelet_samples is used.

time_limit_in_minutesint, default=0

Time contract to limit build time in minutes, overriding n_shapelet_samples and transform_limit_in_minutes. The estimator will only be contracted if a time_limit_in_minutes parameter is present. Default of 0 means n_shapelet_samples or transform_limit_in_minutes is used.

contract_max_n_shapelet_samplesint, default=np.inf

Max number of shapelets to extract when contracting the transform with transform_limit_in_minutes or time_limit_in_minutes.

save_transformed_databool, default=False

Save the data transformed in fit in transformed_data_ for use in _get_train_probs.

n_jobsint, default=1

The number of jobs to run in parallel for both fit and predict. -1 means using all processors.

batch_sizeint or None, default=100

Number of shapelet candidates processed before being merged into the set of best shapelets in the transform.

random_stateint, RandomState instance or None, default=None

If int, random_state is the seed used by the random number generator; If RandomState instance, random_state is the random number generator; If None, the random number generator is the RandomState instance used by np.random.

Attributes:
classes_list

The unique class labels in the training set.

n_classes_int

The number of unique classes in the training set.

fit_time_int

The time (in milliseconds) for fit to run.

n_instances_int

The number of train cases in the training set.

n_dims_int

The number of dimensions per case in the training set.

series_length_int

The length of each series in the training set.

transformed_data_list of shape (n_estimators) of ndarray

The transformed training dataset for all classifiers. Only saved when save_transformed_data is True.

See also

RandomShapeletTransform

The randomly sampled shapelet transform.

RotationForest

The default rotation forest classifier used.

Notes

For the Java version, see tsml.

References

[1]

Jon Hills et al., “Classification of time series by shapelet transformation”, Data Mining and Knowledge Discovery, 28(4), 851–881, 2014.

[2]

A. Bostrom and A. Bagnall, “Binary Shapelet Transform for Multiclass Time Series Classification”, Transactions on Large-Scale Data and Knowledge Centered Systems, 32, 2017.

Examples

>>> from sktime.classification.shapelet_based import ShapeletTransformClassifier
>>> from sktime.classification.sklearn import RotationForest
>>> from sktime.datasets import load_unit_test
>>> X_train, y_train = load_unit_test(split="train", return_X_y=True)
>>> X_test, y_test = load_unit_test(split="test", return_X_y=True)
>>> clf = ShapeletTransformClassifier(
...     estimator=RotationForest(n_estimators=3),
...     n_shapelet_samples=100,
...     max_shapelets=10,
...     batch_size=20,
... )
>>> clf.fit(X_train, y_train)
ShapeletTransformClassifier(...)
>>> y_pred = clf.predict(X_test)

Methods

check_is_fitted([method_name])

Check if the estimator has been fitted.

clone()

Obtain a clone of the object with same hyper-parameters and config.

clone_tags(estimator[, tag_names])

Clone tags from another object as dynamic override.

create_test_instance([parameter_set])

Construct an instance of the class, using first test parameter set.

create_test_instances_and_names([parameter_set])

Create list of all test instances and a list of names for them.

fit(X, y)

Fit time series classifier to training data.

fit_predict(X, y[, cv, change_state])

Fit and predict labels for sequences in X.

fit_predict_proba(X, y[, cv, change_state])

Fit and predict labels probabilities for sequences in X.

get_class_tag(tag_name[, tag_value_default])

Get class tag value from class, with tag level inheritance from parents.

get_class_tags()

Get class tags from class, with tag level inheritance from parent classes.

get_config()

Get config flags for self.

get_fitted_params([deep])

Get fitted parameters.

get_param_defaults()

Get object's parameter defaults.

get_param_names([sort])

Get object's parameter names.

get_params([deep])

Get a dict of parameters values for this object.

get_tag(tag_name[, tag_value_default, ...])

Get tag value from instance, with tag level inheritance and overrides.

get_tags()

Get tags from instance, with tag level inheritance and overrides.

get_test_params([parameter_set])

Return testing parameter settings for the estimator.

is_composite()

Check if the object is composed of other BaseObjects.

load_from_path(serial)

Load object from file location.

load_from_serial(serial)

Load object from serialized memory container.

predict(X)

Predicts labels for sequences in X.

predict_proba(X)

Predicts labels probabilities for sequences in X.

reset()

Reset the object to a clean post-init state.

save([path, serialization_format])

Save serialized self to bytes-like object or to (.zip) file.

score(X, y)

Scores predicted labels against ground truth labels on X.

set_config(**config_dict)

Set config flags to given values.

set_params(**params)

Set the parameters of this object.

set_random_state([random_state, deep, ...])

Set random_state pseudo-random seed parameters for self.

set_tags(**tag_dict)

Set instance level tag overrides to given values.