HFTransformersForecaster
HFTransformersForecaster
- class HFTransformersForecaster(model_path: str = None, fit_strategy='minimal', validation_split=0.2, config=None, training_args=None, compute_metrics=None, deterministic=False, callbacks=None, peft_config=None, device=None)[source]
Forecaster that uses a huggingface model for forecasting.
This forecaster fetches the model from the huggingface model hub. Note, this forecaster is in an experimental state. It is currently only working for Informer, Autoformer, and TimeSeriesTransformer.
- Parameters:
- model_pathstr or PreTrainedModel
Path to the huggingface model to use for forecasting. Currently, Informer, Autoformer, and TimeSeriesTransformer are supported. This can be one of the following: - A string specifying the Hugging Face model name or path
(e.g., “huggingface/autoformer-tourism-monthly”).
An instance of a PreTrainedModel, allowing manual initialization and configuration.
- fit_strategystr, default=”minimal”
Strategy to use for fitting (fine-tuning) the model. This can be one of the following:
“minimal”: Fine-tunes only a small subset of the model parameters, allowing for quick adaptation with limited computational resources.
“full”: Fine-tunes all model parameters, which may result in better performance but requires more computational power and time.
“peft”: Applies Parameter-Efficient Fine-Tuning (PEFT) techniques to adapt the model with fewer trainable parameters, saving computational resources.
Note: If the ‘peft’ package is not available, a ModuleNotFoundError will be raised, indicating that the ‘peft’ package is required. Please install it using pip install peft to use this fit strategy.
- validation_splitfloat, default=0.2
Fraction of the data to use for validation
- configdict, default={}
Configuration to use for the model. Configuration objects inherit from
PreTrainedConfigand can be used to control the model outputs and architecture. Refer to the individual model config for particular model-specific config params.PreTrainedConfigis the base class for all configuration classes. It handles a few parameters common to all models’ configurations as well as methods for loading/downloading/saving configurations. A configuration file can be loaded and saved to disk. Loading the configuration file and using this file to initialize a model does not load the model weights. It only affects the model’s configuration.Keys supported by
PreTrainedConfiginclude:- name_or_pathstr, optional, default=””
Store the string that was passed to
PreTrainedModel.from_pretrained()aspretrained_model_name_or_pathif the configuration was created with such a method.- output_hidden_statesbool, optional, default=False
Whether or not the model should return all hidden-states.
- output_attentionsbool, optional, default=False
Whether or not the model should return all attentions.
- return_dictbool, optional, default=True
Whether or not the model should return a
ModelOutputinstead of a plain tuple.- is_encoder_decoderbool, optional, default=False
Whether the model is used as an encoder/decoder or not.
- chunk_size_feed_forwardint, optional, default=0
The chunk size of all feed forward layers in the residual attention blocks. A chunk size of
0means that the feed forward layer is not chunked. A chunk size ofnmeans that the feed forward layer processesn < sequence_lengthembeddings at a time.- per_layer_configdict, optional
A sparse mapping from layer indices to configuration attribute overrides. Each key is a layer index, and each value contains the attributes that differ from the global config for that layer.
Parameters for fine-tuning tasks:
- architectureslist of str, optional
Model architectures that can be used with the model pretrained weights.
- id2labeldict of int to str, optional
A map from index (for instance prediction index, or target index) to label.
- label2iddict of str to int, optional
A map from label to index for the model.
- num_labelsint, optional
Number of labels to use in the last layer added to the model, typically for a classification task.
- problem_typestr, optional
Problem type for
XxxForSequenceClassificationmodels. Can be one of"regression","single_label_classification"or"multi_label_classification".
PyTorch specific parameters:
- dtypestr, optional
The dtype of the weights. This attribute can be used to initialize the model to a non-default dtype (which is normally
float32) and thus allow for optimal storage allocation. For example, if the saved model isfloat16, ideally we want to load it back using the minimal amount of memory needed to loadfloat16weights.
Class attributes (overridden by derived classes):
- model_typestr
An identifier for the model type, serialized into the JSON file, and used to recreate the correct object in
AutoConfig.- has_no_defaults_at_initbool
Whether the config class can be initialized without providing input arguments. Some configurations require inputs to be defined at init and have no default values, usually these are composite configs (but not necessarily) such as
EncoderDecoderConfigorRagConfig. They have to be initialized from two or more configs of typePreTrainedConfig.- keys_to_ignore_at_inferencelist of str
A list of keys to ignore by default when looking at dictionary outputs of the model during inference.
- attribute_mapdict of str to str
A dict that maps model specific attribute names to the standardized naming of attributes.
- base_model_tp_plandict
A dict that maps sub-modules FQNs of a base model to a tensor parallel plan applied to the sub-module when
model.tensor_parallelis called.- base_model_fsdp_plandict
A dict that maps sub-modules of a base model to an FSDP2 sharding strategy (e.g.
"free_full_weight"/"keep_full_weight"). Keys can be wildcard module paths (e.g."layers.*") or tuples of paths (grouped into a singlefully_shardcall).- base_model_pp_plandict of str to tuple of list of str
A dict that maps child-modules of a base model to a pipeline parallel plan that enables users to place the child-module on the appropriate device.
Common attributes (present in all subclasses):
- vocab_sizeint
The number of tokens in the vocabulary, which is also the first dimension of the embeddings matrix (this attribute may be missing for models that don’t have a text modality like ViT).
- hidden_sizeint
The hidden size of the model.
- num_attention_headsint
The number of attention heads used in the multi-head attention layers of the model.
- num_hidden_layersint
The number of blocks in the model.
- training_argsdict, default={}
Training arguments to use for the model. See
transformers.TrainingArgumentsfor details [1]. Note that theoutput_dirargument is required.- compute_metricslist, default=None
List of metrics to compute during training. See
transformers.Trainerfor details.- deterministicbool, default=False
Whether the predictions should be deterministic or not.
- callbackslist, default=[]
List of callbacks to use during training. See
transformers.Trainer- peft_configpeft.PeftConfig, default=None
Configuration for Parameter-Efficient Fine-Tuning. When
fit_strategyis set to “peft”, this will be used to set up PEFT parameters for the model. See thepeftdocumentation for details [2].- devicestr, optional (default=None)
Device on which to load the model, passed to the transformers
device_map, for example"cpu","cuda", or"auto"."auto"selects an available accelerator. IfNone, the transformers default placement is used. Ignored whenmodel_pathis an already initialized model object, which keeps its own device. Settingdevicerequires theacceleratepackage, which is not part of the basetransformersinstall. Install it withpip install accelerateorpip install "transformers[torch]".
- Attributes:
cutoffCut-off = “present time” state of forecaster.
fhForecasting horizon that was passed.
is_fittedWhether
fithas been called.stateState of the estimator.
References
[1]Examples
Using a Pretrained Model from Hugging Face
>>> from sktime.forecasting.hf_transformers import HFTransformersForecaster >>> from sktime.datasets import load_airline >>> y = load_airline() >>> forecaster = HFTransformersForecaster( ... model_path="huggingface/autoformer-tourism-monthly", ... training_args ={ ... "num_train_epochs": 20, ... "output_dir": "test_output", ... "per_device_train_batch_size": 32, ... }, ... config={ ... "lags_sequence": [1, 2, 3], ... "context_length": 2, ... "prediction_length": 4, ... "use_cpu": True, ... "label_length": 2, ... }, ... ) >>> forecaster.fit(y) >>> fh = [1, 2, 3] >>> y_pred = forecaster.predict(fh)
Using PEFT for Fine-Tuning
>>> from sktime.forecasting.hf_transformers import HFTransformersForecaster >>> from sktime.datasets import load_airline >>> from peft import LoraConfig >>> y = load_airline() >>> forecaster = HFTransformersForecaster( ... model_path="huggingface/autoformer-tourism-monthly", ... fit_strategy="peft", ... training_args={ ... "num_train_epochs": 20, ... "output_dir": "test_output", ... "per_device_train_batch_size": 32, ... }, ... config={ ... "lags_sequence": [1, 2, 3], ... "context_length": 2, ... "prediction_length": 4, ... "use_cpu": True, ... "label_length": 2, ... }, ... peft_config=LoraConfig( ... r=8, ... lora_alpha=32, ... target_modules=["q_proj", "v_proj"], ... lora_dropout=0.01, ... ) ... ) >>> forecaster.fit(y) >>> fh = [1, 2, 3] >>> y_pred = forecaster.predict(fh)
Using an Initialized Model
>>> from sktime.datasets import load_airline >>> from transformers import AutoformerConfig, AutoformerForPrediction >>> from sktime.forecasting.hf_transformers import HFTransformersForecaster >>> y = load_airline()
>>> # Define model configuration >>> config = AutoformerConfig( ... num_dynamic_real_features=0, ... num_static_real_features=0, ... num_static_categorical_features=0, ... num_time_features=0, ... context_length=32, ... prediction_length=8, ... lags_sequence=[1, 2, 3], ... )
>>> # Initialize the model >>> model = AutoformerForPrediction(config)
>>> # Initialize the forecaster with the model >>> forecaster = HFTransformersForecaster( ... model_path=model, ... fit_strategy="minimal", ... training_args={ ... "num_train_epochs": 10, ... "output_dir": "output", ... "per_device_train_batch_size": 4 ... }, ... )
>>> forecaster.fit(y) >>> fh = [1, 2, 3] >>> y_pred = forecaster.predict(fh)
Methods
check_is_fitted([method_name])Check if the estimator has been fitted.
clone()Obtain a clone of the object with same hyper-parameters and config.
clone_tags(estimator[, tag_names])Clone tags from another object as dynamic override.
create_test_instance([parameter_set])Construct an instance of the class, using first test parameter set.
create_test_instances_and_names([parameter_set])Create list of all test instances and a list of names for them.
fit(y[, X, fh])Fit forecaster to training data.
fit_predict(y[, X, fh, X_pred])Fit and forecast time series at future horizon.
get_class_tag(tag_name[, tag_value_default])Get class tag value from class, with tag level inheritance from parents.
get_class_tags()Get class tags from class, with tag level inheritance from parent classes.
get_config()Get config flags for self.
get_fitted_params([deep])Get fitted parameters.
get_param_defaults()Get object's parameter defaults.
get_param_names([sort])Get object's parameter names.
get_params([deep])Get a dict of parameters values for this object.
get_pretrained_params([deep])Get pretrained parameters of this estimator.
get_tag(tag_name[, tag_value_default, ...])Get tag value from instance, with tag level inheritance and overrides.
get_tags()Get tags from instance, with tag level inheritance and overrides.
get_test_params([parameter_set])Return testing parameter settings for the estimator.
is_composite()Check if the object is composed of other BaseObjects.
load_from_path(serial)Load object from file location.
load_from_serial(serial)Load object from serialized memory container.
predict([fh, X])Forecast time series at future horizon.
predict_interval([fh, X, coverage])Compute/return prediction interval forecasts.
predict_proba([fh, X, marginal])Compute/return fully probabilistic forecasts.
predict_quantiles([fh, X, alpha])Compute/return quantile forecasts.
predict_residuals([y, X])Return residuals of time series forecasts.
predict_var([fh, X, cov])Compute/return variance forecasts.
pretrain(y[, X, fh])Pre-train forecaster on panel (global) data.
reset()Reset the object to a clean post-init state.
save([path, serialization_format])Save serialized self to bytes-like object or to (.zip) file.
score(y[, X, fh])Scores forecast against ground truth, using MAPE (non-symmetric).
set_config(**config_dict)Set config flags to given values.
set_params(**params)Set the parameters of this object.
set_random_state([random_state, deep, ...])Set random_state pseudo-random seed parameters for self.
set_tags(**tag_dict)Set instance level tag overrides to given values.
update(y[, X, update_params])Update cutoff value and, optionally, fitted parameters.
update_predict(y[, cv, X, update_params, ...])Make predictions and update model iteratively over the test set.
update_predict_single([y, fh, X, update_params])Update model with new data and make forecasts.

