BOSSEnsemble
BOSSEnsemble
- class BOSSEnsemble(threshold=0.92, max_ensemble_size=500, max_win_len_prop=1, min_window=10, save_train_predictions=False, feature_selection='none', use_boss_distance=True, alphabet_size=2, n_jobs=1, random_state=None)[source]
Ensemble of Bag of Symbolic Fourier Approximation Symbols (BOSS).
Implementation of BOSS Ensemble from Schäfer (2015). [1]
Overview: Input n series of length m and BOSS performs a grid search over a set of parameter values, evaluating each with a LOOCV. It then retains all ensemble members within 92% of the best by default for use in the ensemble. There are three primary parameters:
alpha: alphabet size
w: window length
l: word length.
For any combination, a single BOSS slides a window length w along the series. The w length window is shortened to an l length word through taking a Fourier transform and keeping the first l/2 complex coefficients. These l coefficients are then discretized into alpha possible values, to form a word length l. A histogram of words for each series is formed and stored.
Fit involves finding “n” histograms.
Predict uses 1 nearest neighbor with a bespoke BOSS distance function.
- Parameters:
- thresholdfloat, default=0.92
Threshold used to determine which classifiers to retain. All classifiers within percentage
thresholdof the best one are retained.- max_ensemble_sizeint or None, default=500
Maximum number of classifiers to retain. Will limit number of retained classifiers even if more than
max_ensemble_sizeare within threshold.- max_win_len_propint or float, default=1
Maximum window length as a proportion of the series length.
- min_windowint, default=10
Minimum window size.
- save_train_predictionsbool, default=False
Save the ensemble member train predictions in fit for use in _get_train_probs leave-one-out cross-validation.
- alphabet_sizedefault = 2
Number of possible letters (values) for each word.
- n_jobsint, default=1
The number of jobs to run in parallel for both
fitandpredict.-1means using all processors.- use_boss_distanceboolean, default=True
The Boss-distance is an asymmetric distance measure. It provides higher accuracy, yet is signifaicantly slower to compute.
- feature_selection: {“chi2”, “none”, “random”}, default: none
Sets the feature selections strategy to be used. Chi2 reduces the number of words significantly and is thus much faster (preferred). Random also reduces the number significantly. None applies not feature selectiona and yields large bag of words, e.g. much memory may be needed.
- random_stateint or None, default=None
Seed for random, integer.
- Attributes:
- n_classes_int
Number of classes. Extracted from the data.
- classes_list
The classes labels.
- n_instances_int
Number of instances. Extracted from the data.
- n_estimators_int
The final number of classifiers used. Will be <=
max_ensemble_sizeifmax_ensemble_sizehas been specified.- series_length_int
Length of all series (assumed equal).
- estimators_list
List of DecisionTree classifiers.
See also
Notes
For the Java version, see - Original Publication. - TSML.
References
[1]Patrick Schäfer, “The BOSS is concerned with time series classification in the presence of noise”, Data Mining and Knowledge Discovery, 29(6): 2015 https://link.springer.com/article/10.1007/s10618-014-0377-7
Examples
>>> from sktime.classification.dictionary_based import BOSSEnsemble >>> from sktime.datasets import load_unit_test >>> X_train, y_train = load_unit_test(split="train", return_X_y=True) >>> X_test, y_test = load_unit_test(split="test", return_X_y=True) >>> clf = BOSSEnsemble(max_ensemble_size=3) >>> clf.fit(X_train, y_train) BOSSEnsemble(...) >>> y_pred = clf.predict(X_test)
Methods
check_is_fitted([method_name])Check if the estimator has been fitted.
clone()Obtain a clone of the object with same hyper-parameters and config.
clone_tags(estimator[, tag_names])Clone tags from another object as dynamic override.
create_test_instance([parameter_set])Construct an instance of the class, using first test parameter set.
create_test_instances_and_names([parameter_set])Create list of all test instances and a list of names for them.
fit(X, y)Fit time series classifier to training data.
fit_predict(X, y[, cv, change_state])Fit and predict labels for sequences in X.
fit_predict_proba(X, y[, cv, change_state])Fit and predict labels probabilities for sequences in X.
get_class_tag(tag_name[, tag_value_default])Get class tag value from class, with tag level inheritance from parents.
get_class_tags()Get class tags from class, with tag level inheritance from parent classes.
get_config()Get config flags for self.
get_fitted_params([deep])Get fitted parameters.
get_param_defaults()Get object's parameter defaults.
get_param_names([sort])Get object's parameter names.
get_params([deep])Get a dict of parameters values for this object.
get_tag(tag_name[, tag_value_default, ...])Get tag value from instance, with tag level inheritance and overrides.
get_tags()Get tags from instance, with tag level inheritance and overrides.
get_test_params([parameter_set])Return testing parameter settings for the estimator.
is_composite()Check if the object is composed of other BaseObjects.
load_from_path(serial)Load object from file location.
load_from_serial(serial)Load object from serialized memory container.
predict(X)Predicts labels for sequences in X.
predict_proba(X)Predicts labels probabilities for sequences in X.
reset()Reset the object to a clean post-init state.
save([path, serialization_format])Save serialized self to bytes-like object or to (.zip) file.
score(X, y)Scores predicted labels against ground truth labels on X.
set_config(**config_dict)Set config flags to given values.
set_params(**params)Set the parameters of this object.
set_random_state([random_state, deep, ...])Set random_state pseudo-random seed parameters for self.
set_tags(**tag_dict)Set instance level tag overrides to given values.

