ForecastingOptCV
Tune an sktime forecaster via any optimizer in the hyperactive toolbox.
ForecastingOptCV uses any available tuning engine from hyperactive to tune a forecaster by backtesting.
It passes backtesting results as scores to the tuning engine, which identifies the best hyperparameters.
Any available tuning engine from hyperactive can be used, for example:
grid search -
from hyperactive.opt import GridSearchSk as GridSearch, this results in the same algorithm asForecastingGridSearchCVhill climbing -
from hyperactive.opt import HillClimbingoptuna parzen-tree search -
from hyperactive.opt.optuna import TPEOptimizer
Configuration of the tuning engine is as per the respective documentation.
Formally, ForecastingOptCV does the following:
In fit:
wraps the
forecaster,scoring, and other parameters into aSktimeForecastingExperimentinstance, which is passed to the optimizeroptimizeras theexperimentargument.Optimal parameters are then obtained from
optimizer.solve, and set asbest_params_andbest_forecaster_attributes.If
refit=True,best_forecaster_is fitted to the entireyandX.
In predict and predict-like methods, calls the respective method of the best_forecaster_ if refit=True.
Quickstart
from sktime.forecasting.model_selection import ForecastingOptCV
estimator = ForecastingOptCV(forecaster, optimizer, cv, strategy='refit', update_behaviour='full_refit', scoring=None, refit=True, error_score=nan, cv_X=None, backend=None, backend_params=None)Parameters(11)
- forecastersktime forecaster, BaseForecaster instance or interface compatible
- The forecaster to tune, must implement the sktime forecaster interface.
- optimizerhyperactive BaseOptimizer
- The optimizer to be used for hyperparameter search.
- cvsktime BaseSplitter descendant
determines split of
yand possiblyXinto test and train folds y is always split according tocv, see aboveif
cv_Xis not passed,Xsplits are subset tolocequal toyif
cv_Xis passed,Xis split according tocv_X
- strategy{“refit”, “update”, “no-update_params”}, optional, default=”refit”
data ingestion strategy in fitting cv, passed to
evaluateinternally defines the ingestion mode when the forecaster sees new data when window expands"refit"= a new copy of the forecaster is fitted to each training window"update"= forecaster is updated with training window data, in sequence provided"no-update_params"= fit to first training window, re-used without fit or update
- update_behaviourstr, optional, default = “full_refit”
one of {“full_refit”, “inner_only”, “no_update”} behaviour of the forecaster when calling update
"full_refit"= both tuning parameters and inner estimator refit on all data seen"inner_only"= tuning parameters are not re-tuned, inner estimator is updated"no_update"= neither tuning parameters nor inner estimator are updated
- scoringsktime metric (BaseMetric), str, or callable, optional (default=None)
scoring metric to use in tuning the forecaster
sktime metric objects (BaseMetric) descendants can be searched
with the
registry.all_estimatorssearch utility, for instance viaall_estimators("metric", as_dataframe=True)If callable, must have signature
(y_true: 1D np.ndarray, y_pred: 1D np.ndarray) -> float, assuming np.ndarrays being of the same length, and lower being better. Metrics in sktime.performance_metrics.forecasting are all of this form.If str, uses registry.resolve_alias to resolve to one of the above. Valid strings are valid registry.craft specs, which include string repr-s of any BaseMetric object, e.g., “MeanSquaredError()”; and keys of registry.ALIAS_DICT referring to metrics.
If None, defaults to MeanAbsolutePercentageError()
- refitbool, optional (default=True)
Whether to refit the forecaster with the best parameters on the entire data.
True = refit the forecaster with the best parameters on the entire data in
fitFalse = no refitting takes place. The forecaster cannot be used to predict. This is to be used to tune the hyperparameters, and then use the estimator as a parameter estimator, e.g., via
get_fitted_paramsorPluginParamsForecaster.
- error_score“raise” or numeric, default=np.nan
- Value to assign to the score if an exception occurs in estimator fitting. If set to “raise”, the exception is raised. If a numeric value is given, FitFailedWarning is raised.
- cv_Xsktime BaseSplitter descendant, optional
determines split of
Xinto test and train folds default isXbeing split to identicallocindices asyif passed, must have same number of splits ascv- backendstring, by default “None”.
Parallelization backend to use for runs. Runs parallel evaluate if specified and
strategy="refit".“None”: executes loop sequentially, simple list comprehension
“loky”, “multiprocessing” and “threading”: uses
joblib.Parallelloops“joblib”: custom and 3rd party
joblibbackends, e.g.,spark“dask”: uses
dask, requiresdaskpackage in environment“dask_lazy”: same as “dask”, but changes the return to (lazy)
dask.dataframe.DataFrame.“ray”: uses
ray, requiresraypackage in environment
Recommendation: Use “dask” or “loky” for parallel evaluate. “threading” is unlikely to see speed ups due to the GIL and the serialization backend (
cloudpickle) for “dask” and “loky” is generally more robust than the standardpicklelibrary used in “multiprocessing”.- backend_paramsdict, optional
additional parameters passed to the backend as config. Directly passed to
utils.parallel.parallelize. Valid keys depend on the value ofbackend:“None”: no additional parameters,
backend_paramsis ignored“loky”, “multiprocessing” and “threading”: default
joblibbackends any valid keys forjoblib.Parallelcan be passed here, e.g.,n_jobs, with the exception ofbackendwhich is directly controlled bybackend. Ifn_jobsis not passed, it will default to-1, other parameters will default tojoblibdefaults.“joblib”: custom and 3rd party
joblibbackends, e.g.,spark. any valid keys forjoblib.Parallelcan be passed here, e.g.,n_jobs,backendmust be passed as a key ofbackend_paramsin this case. Ifn_jobsis not passed, it will default to-1, other parameters will default tojoblibdefaults.“dask”: any valid keys for
dask.computecan be passed, e.g.,scheduler“ray”: The following keys can be passed:
“ray_remote_args”: dictionary of valid keys for
ray.init- “shutdown_ray”: bool, default=True; False prevents
rayfrom shutting down after parallelization.
- “shutdown_ray”: bool, default=True; False prevents