Skip to content

TemporalDictionaryEnsemble

TemporalDictionaryEnsemble

class TemporalDictionaryEnsemble(n_parameter_samples=250, max_ensemble_size=50, max_win_len_prop=1, min_window=10, randomly_selected_params=50, bigrams=None, dim_threshold=0.85, max_dims=20, time_limit_in_minutes=0.0, contract_max_n_parameter_samples=inf, typed_dict=True, save_train_predictions=False, n_jobs=1, random_state=None)[source]

Temporal Dictionary Ensemble (TDE).

Implementation of the dictionary based Temporal Dictionary Ensemble as described in [R822586daffca-1].

Overview: Input “n” series length “m” with “d” dimensions TDE searches “k” parameter values selected using a Gaussian processes regressor, evaluating each with a LOOCV. It then retains “s” ensemble members. There are six primary parameters for individual classifiers:

  • alpha: alphabet size

  • w: window length

  • l: word length

  • p: normalise/no normalise

  • h: levels

  • b: MCB/IGB

For any combination, an individual TDE classifier slides a window of length w along the series. The w length window is shortened to an l length word through taking a Fourier transform and keeping the first l/2 complex coefficients. These lcoefficients are then discretised into alpha possible values, to form a word length l using breakpoints found using b. A histogram of words for each series is formed and stored, using a spatial pyramid of h levels. For multivariate series, accuracy from a reduced histogram is used to select dimensions.

fit involves finding n histograms. predict uses 1 nearest neighbour with a the histogram intersection distance function.

Parameters:
n_parameter_samplesint, default=250

Number of parameter combinations to consider for the final ensemble.

max_ensemble_sizeint, default=50

Maximum number of estimators in the ensemble.

max_win_len_propfloat, default=1

Maximum window length as a proportion of series length, must be between 0 and 1.

min_windowint, default=10

Minimum window length.

randomly_selected_params: int, default=50

Number of parameters randomly selected before the Gaussian process parameter selection is used.

bigramsboolean or None, default=None

Whether to use bigrams, defaults to true for univariate data and false for multivariate data.

dim_thresholdfloat, default=0.85

Dimension accuracy threshold for multivariate data, must be between 0 and 1.

max_dimsint, default=20

Max number of dimensions per classifier for multivariate data.

time_limit_in_minutesint, default=0

Time contract to limit build time in minutes, overriding n_parameter_samples. Default of 0 means n_parameter_samples is used.

contract_max_n_parameter_samplesint, default=np.inf

Max number of parameter combinations to consider when time_limit_in_minutes is set.

typed_dictbool, default=True

Use a numba typed Dict to store word counts. May increase memory usage, but will be faster for larger datasets. As the Dict cannot be pickled currently, there will be some overhead converting it to a python dict with multiple threads and pickling.

save_train_predictionsbool, default=False

Save the ensemble member train predictions in fit for use in _get_train_probs leave-one-out cross-validation.

n_jobsint, default=1

The number of jobs to run in parallel for both fit and predict. -1 means using all processors.

random_stateint or None, default=None

Seed for random number generation.

Attributes:
n_classes_int

The number of classes.

classes_list

The classes labels.

n_instances_int

The number of train cases.

n_dims_int

The number of dimensions per case.

series_length_int

The length of each series.

n_estimators_int

The final number of classifiers used (<= max_ensemble_size)

estimators_list of shape (n_estimators) of IndividualTDE

The collections of estimators trained in fit.

weights_list of shape (n_estimators) of float

Weight of each estimator in the ensemble.

Notes

For the Java version, see TSML.

References

Examples

>>> from sktime.classification.dictionary_based import TemporalDictionaryEnsemble
>>> from sktime.datasets import load_unit_test
>>> X_train, y_train = load_unit_test(split="train", return_X_y=True)
>>> X_test, y_test = load_unit_test(split="test", return_X_y=True)
>>> clf = TemporalDictionaryEnsemble(
...     n_parameter_samples=10,
...     max_ensemble_size=3,
...     randomly_selected_params=5,
... )
>>> clf.fit(X_train, y_train)
TemporalDictionaryEnsemble(...)
>>> y_pred = clf.predict(X_test)

Methods

check_is_fitted([method_name])

Check if the estimator has been fitted.

clone()

Obtain a clone of the object with same hyper-parameters and config.

clone_tags(estimator[, tag_names])

Clone tags from another object as dynamic override.

create_test_instance([parameter_set])

Construct an instance of the class, using first test parameter set.

create_test_instances_and_names([parameter_set])

Create list of all test instances and a list of names for them.

fit(X, y)

Fit time series classifier to training data.

fit_predict(X, y[, cv, change_state])

Fit and predict labels for sequences in X.

fit_predict_proba(X, y[, cv, change_state])

Fit and predict labels probabilities for sequences in X.

get_class_tag(tag_name[, tag_value_default])

Get class tag value from class, with tag level inheritance from parents.

get_class_tags()

Get class tags from class, with tag level inheritance from parent classes.

get_config()

Get config flags for self.

get_fitted_params([deep])

Get fitted parameters.

get_param_defaults()

Get object's parameter defaults.

get_param_names([sort])

Get object's parameter names.

get_params([deep])

Get a dict of parameters values for this object.

get_tag(tag_name[, tag_value_default, ...])

Get tag value from instance, with tag level inheritance and overrides.

get_tags()

Get tags from instance, with tag level inheritance and overrides.

get_test_params([parameter_set])

Return testing parameter settings for the estimator.

is_composite()

Check if the object is composed of other BaseObjects.

load_from_path(serial)

Load object from file location.

load_from_serial(serial)

Load object from serialized memory container.

predict(X)

Predicts labels for sequences in X.

predict_proba(X)

Predicts labels probabilities for sequences in X.

reset()

Reset the object to a clean post-init state.

save([path, serialization_format])

Save serialized self to bytes-like object or to (.zip) file.

score(X, y)

Scores predicted labels against ground truth labels on X.

set_config(**config_dict)

Set config flags to given values.

set_params(**params)

Set the parameters of this object.

set_random_state([random_state, deep, ...])

Set random_state pseudo-random seed parameters for self.

set_tags(**tag_dict)

Set instance level tag overrides to given values.