TinyTimeMixerForecaster
TinyTimeMixer Forecaster for Zero-Shot Forecasting of Multivariate Time Series.
Wrapping implementation in [1] of method proposed in [2]. See [3] for tutorial by creators.
TinyTimeMixer (TTM) are compact pre-trained models for Time-Series Forecasting, open-sourced by IBM Research. With less than 1 Million parameters, TTM introduces the notion of the first-ever “tiny” pre-trained models for Time-Series Forecasting.
Fit Strategies: Full, Minimal, and Zero-shot
This model supports three fit strategies: zero-shot for direct predictions without training, minimal fine-tuning for lightweight adaptation to new data, and full fine-tuning for comprehensive model training. The selected strategy is determined by the model’s fit_strategy parameter
Initialization Process:
Model Path: The
model_pathparameter points to a local folder or huggingface repo that contains both configuration files and pretrained weights.Default Configuration: The model loads its default configuration from the configuration files.
Custom Configuration: Users can provide a custom configuration via the
configparameter during model initialization.Configuration Override: If custom configuration is provided, it overrides the default configuration.
Forecasting Horizon: If the forecasting horizon (
fh) specified duringfitexceeds the defaultconfig.prediction_length, the configuration is updated to reflectmax(fh).Model Architecture: The final configuration is used to construct the model architecture.
Pretrained Weights: pretrained weights are loaded from the
model_path, these weights are then aligned and loaded into the model architecture.Weight Alignment: However sometimes, pretrained weights do not align with the model architecture, because the config was changed which created a model architecture of different size than the default one. This causes some of the weights in model architecture to be reinitialized randomly instead of using the pre-trained weights.
Training Strategies:
Zero-shot Forecasting: When all the pre-trained weights are correctly aligned with the model architecture, fine-tuing part is bypassed and the model preforms zero-short forecasting.
Minimal Fine-tuning: When not all the pre-trained weights are correctly aligned with the model architecture, rather some weights are re-initialized, these re-initialized weights are fine-tuned on the provided data.
Full Fine-tuning: The model is fully fine-tuned on new data, updating all parameters. This approach offers maximum adaptation to the dataset but requires more computational resources.
Exogenous Variables Support
TTM supports exogenous variables (external factors) that can improve forecasting accuracy. The model accepts exogenous variables that are known for both the historical period and the future forecasting horizon.
When using exogenous variables: - The X parameter should contain exogenous data covering both
past and future periods
Exogenous variables must have the same index structure as the target series
For prediction, exogenous data must extend into the forecasting horizon
Schnellstart
from sktime.forecasting.ttm import TinyTimeMixerForecaster
estimator = TinyTimeMixerForecaster(model_path='ibm/TTM', revision='main', validation_split=0.2, config=None, training_args=None, compute_metrics=None, callbacks=None, broadcasting=False, use_source_package=False, fit_strategy='minimal')Parameter(10)
- model_pathstr, default=”ibm/TTM”
Path to the Huggingface model to use for forecasting. This can be either:
The name of a Huggingface repository (e.g., “ibm/TTM”)
A local path to a folder containing model files in a format supported by transformers. In this case, ensure that the directory contains all necessary files (e.g., configuration, tokenizer, and model weights).
If this parameter is None, fit_strategy should be full to allow full fine tuning of the model loaded from pretrained/provided config, else ValueError is raised.
- revision: str, default=”main”
Revision of the model to use:
“main”: For loading model with context_length of 512 and prediction_length of 96.
“1024_96_v1”: For loading model with context_length of 1024 and prediction_length of 96.
This param becomes irrelevant when model_path is None
- validation_splitfloat, default=0.2
- Fraction of the data to use for validation
- configdict, default={}
Configuration to use for the model. See the
transformersdocumentation for details.- training_argsdict, default={}
Training arguments to use for the model. See
transformers.TrainingArgumentsfor details. Note that theoutput_dirargument is required.- compute_metricslist, default=None
List of metrics to compute during training. See
transformers.Trainerfor details.- callbackslist, default=[]
List of callbacks to use during training. See
transformers.Trainer- broadcastingbool, default=False
if True, multiindex data input will be broadcasted to single series. For each single series, one copy of this forecaster will try to fit and predict on it. The broadcasting is happening inside automatically, from the outerside api perspective, the input and output are the same, only one multiindex output from
predict.- use_source_packagebool, default=False
If True, the model and configuration will be loaded directly from the source package
tsfm_public.models.tinytimemixer. This is useful if you want to bypass the local version of the package or when working in an environment where the latest updates from the source package are needed. If False, the model and configuration will be loaded from the local version of package maintained in sktime because of model’s unavailability on pypi. To install the source package, follow the instructions here [4].- fit_strategystr, default=”minimal”
Strategy to use for fitting (fine-tuning) the model. This can be one of the following: - “zero-shot”: Uses pre-trained model as it is. If model path is None
with this strategy, ValueError is raised.
“minimal”: Fine-tunes only a small subset of the model parameters, allowing for quick adaptation with limited computational resources. If model path is None with this strategy, ValueError is raised.
“full”: Fine-tunes all model parameters, which may result in better performance but requires more computational power and time. Allows model path to be None.
Beispiele
>>> from sktime.forecasting.ttm import TinyTimeMixerForecaster
>>> from sktime.datasets import load_airline
>>> y = load_airline ()
>>> forecaster = TinyTimeMixerForecaster ()
>>> # performs zero-shot forecasting, as default config (unchanged) is used
>>> forecaster. fit (y, fh = [1, 2, 3 ]) TinyTimeMixerForecaster(
... )
>>> y_pred = forecaster. predict ()
>>> from sktime.forecasting.ttm import TinyTimeMixerForecaster
>>> from sktime.datasets import load_tecator
>>>
>>> # load multi-index dataset
>>> y = load_tecator (
... return_type = "pd-multiindex",
... return_X_y = False
... )
>>> y. drop (['class_val' ], axis = 1, inplace = True)
>>>
>>> # global forecasting on multi-index dataset
>>> forecaster = TinyTimeMixerForecaster (
... model_path = None,
... fit_strategy = "full",
... config = {
... "context_length": 8,
... "prediction_length": 2
... },
... training_args = {
... "num_train_epochs": 1,
... "output_dir": "test_output",
... "per_device_train_batch_size": 32,
... },
... )
>>>
>>> # model initialized with random weights due to None model_path
>>> # and trained with the full strategy.
>>> forecaster. fit (y, fh = [1, 2, 3 ]) TinyTimeMixerForecaster(
... )
>>> y_pred = forecaster. predict () Example with exogenous variables:
>>> from sktime.forecasting.ttm import TinyTimeMixerForecaster
>>> from sktime.datasets import load_longley
>>> from sktime.split import temporal_train_test_split
>>> y, X = load_longley ()
>>> y_train, _, X_train, X_future = temporal_train_test_split (y, X, test_size = 2)
>>>
>>> # Initialize forecaster
>>> forecaster = TinyTimeMixerForecaster (
... model_path = None,
... fit_strategy = "full",
... config = {
... "context_length": 8,
... "prediction_length": 2
... },
... training_args = {
... "max_steps": 10,
... "output_dir": "test_output",
... "per_device_train_batch_size": 4,
... "report_to": "none",
... },
... )
>>>
>>> # Fit with exogenous variables
>>> forecaster. fit (y_train, X = X_train, fh = [1, 2 ]) TinyTimeMixerForecaster(
... )
>>>
>>> # Predict with exogenous variables
>>> y_pred = forecaster. predict (X = X_future)Referenzen
Ekambaram, V., Jati, A., Dayama, P., Mukherjee, S., Nguyen, N.H., Gifford, W.M., Reddy, C. and Kalagnanam, J., 2024. Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series. CoRR.