Skip to content

TSCOptCV

TSCOptCV

class TSCOptCV(estimator, optimizer, cv=None, scoring=None, refit=True, error_score=nan, backend=None, backend_params=None)[source]

Tune an sktime classifier via any optimizer in the hyperactive toolbox.

TSCOptCV uses any available tuning engine from hyperactive to tune a classifier by backtesting.

It passes backtesting results as scores to the tuning engine, which identifies the best hyperparameters.

Any available tuning engine from hyperactive can be used, for example:

  • grid search - from hyperactive.opt import GridSearchSk as GridSearch, this results in the same algorithm as TSCGridSearchCV

  • hill climbing - from hyperactive.opt import HillClimbing

  • optuna parzen-tree search - from hyperactive.opt.optuna import TPEOptimizer

Configuration of the tuning engine is as per the respective documentation.

Formally, TSCOptCV does the following:

In fit:

  • wraps the estimator, scoring, and other parameters into a SktimeClassificationExperiment instance, which is passed to the optimizer optimizer as the experiment argument.

  • Optimal parameters are then obtained from optimizer.solve, and set as best_params_ and best_estimator_ attributes.

  • If refit=True, best_estimator_ is fitted to the entire y and X.

In predict and predict-like methods, calls the respective method of the best_estimator_ if refit=True.

Parameters:
estimatorsktime classifier, BaseClassifier instance or interface compatible

The classifier to tune, must implement the sktime classifier interface.

optimizerhyperactive BaseOptimizer

The optimizer to be used for hyperparameter search.

cvint, sklearn cross-validation generator or an iterable, default=3-fold CV

Determines the cross-validation splitting strategy. Possible inputs for cv are:

  • None = default = KFold(n_splits=3, shuffle=True)

  • integer, number of folds folds in a KFold splitter, shuffle=True

  • An iterable yielding (train, test) splits as arrays of indices.

For integer/None inputs, if the estimator is a classifier and y is either binary or multiclass, StratifiedKFold is used. In all other cases, KFold is used. These splitters are instantiated with shuffle=False so the splits will be the same across calls.

scoringstr, callable, default=None

Strategy to evaluate the performance of the cross-validated model on the test set. Can be:

  • a single string resolvable to an sklearn scorer

  • a callable that returns a single value;

  • None = default = accuracy_score

refitbool, optional (default=True)

True = refit the forecaster with the best parameters on the entire data in fit False = no refitting takes place. The forecaster cannot be used to predict. This is to be used to tune the hyperparameters, and then use the estimator as a parameter estimator, e.g., via get_fitted_params or PluginParamsForecaster.

error_score“raise” or numeric, default=np.nan

Value to assign to the score if an exception occurs in estimator fitting. If set to “raise”, the exception is raised. If a numeric value is given, FitFailedWarning is raised.

backendstring, by default “None”.

Parallelization backend to use for runs. Runs parallel evaluate if specified and strategy="refit".

  • “None”: executes loop sequentially, simple list comprehension

  • “loky”, “multiprocessing” and “threading”: uses joblib.Parallel loops

  • “joblib”: custom and 3rd party joblib backends, e.g., spark

  • “dask”: uses dask, requires dask package in environment

  • “dask_lazy”: same as “dask”, but changes the return to (lazy) dask.dataframe.DataFrame.

  • “ray”: uses ray, requires ray package in environment

Recommendation: Use “dask” or “loky” for parallel evaluate. “threading” is unlikely to see speed ups due to the GIL and the serialization backend (cloudpickle) for “dask” and “loky” is generally more robust than the standard pickle library used in “multiprocessing”.

backend_paramsdict, optional

additional parameters passed to the backend as config. Directly passed to utils.parallel.parallelize. Valid keys depend on the value of backend:

  • “None”: no additional parameters, backend_params is ignored

  • “loky”, “multiprocessing” and “threading”: default joblib backends any valid keys for joblib.Parallel can be passed here, e.g., n_jobs, with the exception of backend which is directly controlled by backend. If n_jobs is not passed, it will default to -1, other parameters will default to joblib defaults.

  • “joblib”: custom and 3rd party joblib backends, e.g., spark. any valid keys for joblib.Parallel can be passed here, e.g., n_jobs, backend must be passed as a key of backend_params in this case. If n_jobs is not passed, it will default to -1, other parameters will default to joblib defaults.

  • “dask”: any valid keys for dask.compute can be passed, e.g., scheduler

  • “ray”: The following keys can be passed:

    • “ray_remote_args”: dictionary of valid keys for ray.init

    • “shutdown_ray”: bool, default=True; False prevents ray from shutting

      down after parallelization.

    • “logger_name”: str, default=”ray”; name of the logger to use.

    • “mute_warnings”: bool, default=False; if True, suppresses warnings

Attributes:
is_fitted

Whether fit has been called.

Methods

check_is_fitted([method_name])

Check if the estimator has been fitted.

clone()

Obtain a clone of the object with same hyper-parameters and config.

clone_tags(estimator[, tag_names])

Clone tags from another object as dynamic override.

create_test_instance([parameter_set])

Construct an instance of the class, using first test parameter set.

create_test_instances_and_names([parameter_set])

Create list of all test instances and a list of names for them.

fit(X, y)

Fit time series classifier to training data.

fit_predict(X, y[, cv, change_state])

Fit and predict labels for sequences in X.

fit_predict_proba(X, y[, cv, change_state])

Fit and predict labels probabilities for sequences in X.

get_class_tag(tag_name[, tag_value_default])

Get class tag value from class, with tag level inheritance from parents.

get_class_tags()

Get class tags from class, with tag level inheritance from parent classes.

get_config()

Get config flags for self.

get_fitted_params([deep])

Get fitted parameters.

get_param_defaults()

Get object's parameter defaults.

get_param_names([sort])

Get object's parameter names.

get_params([deep])

Get a dict of parameters values for this object.

get_tag(tag_name[, tag_value_default, ...])

Get tag value from instance, with tag level inheritance and overrides.

get_tags()

Get tags from instance, with tag level inheritance and overrides.

get_test_params([parameter_set])

Return testing parameter settings for the estimator.

is_composite()

Check if the object is composed of other BaseObjects.

load_from_path(serial)

Load object from file location.

load_from_serial(serial)

Load object from serialized memory container.

predict(X)

Predicts labels for sequences in X.

predict_proba(X)

Predicts labels probabilities for sequences in X.

reset()

Reset the object to a clean post-init state.

save([path, serialization_format])

Save serialized self to bytes-like object or to (.zip) file.

score(X, y)

Scores predicted labels against ground truth labels on X.

set_config(**config_dict)

Set config flags to given values.

set_params(**params)

Set the parameters of this object.

set_random_state([random_state, deep, ...])

Set random_state pseudo-random seed parameters for self.

set_tags(**tag_dict)

Set instance level tag overrides to given values.