TimeMoEForecaster
TimeMoEForecaster
- class TimeMoEForecaster(model_path: str | None = 'Maple728/TimeMoE-50M', config: dict = None, seed: int = None, use_source_package: bool = False, ignore_deps: bool = False, context_length: int = 1024, stride: int = None, training_args: dict = None, validation_split: float | None = 0.2, device: str | None = None, dtype=None)[source]
Interface for TimeMOE forecaster.
TimeMoE is a decoder-only time series foundational model that uses a mixture of experts algorithm to make predictions. designed to operate in an auto-regressive manner, enabling universal forecasting with arbitrary prediction horizons and context lengths of up to 4096. This method has been proposed in [2] and the official code is available at [2].
Supports:
zero-shot forecasting via
fit+predictfine-tuning of a pretrained checkpoint via
pretraintraining from scratch via
pretrainwithmodel_path=None
- Parameters:
- model_path: str or None, default=”Maple728/TimeMoE-50M”
Path to the TimeMOE model. This can be:
A model ID from the HuggingFace Hub, e.g., “Maple728/TimeMoE-50M”
A local directory containing the model files, specified as an absolute or relative path to the current working directory The path should point to a directory containing the model weights and configuration files in the format expected by the HuggingFace Transformers library.
Noneto initialize fromconfigwith random weights (from-scratch)
- config: dict, optional
A dictionary specifying the configuration of the TimeMOE model. The available configuration options include hyperparameters that control the prediction behavior, sampling, and hardware utilization.
- input_size: int, default=1
The size of the input time series.
- hidden_size: int, default=4096
The size of the hidden layers in the TimeMOE model.
- intermediate_size: int, default=22016
The size of the intermediate layers in the TimeMOE model.
- horizon_lengths: list[int], default=[1]
The prediction horizon length.
- num_hidden_layers: int, default=32
The number of hidden layers in the TimeMOE model.
- num_attention_heads: int, default=32
The number of attention heads in the TimeMOE model.
- num_experts_per_tok: int, default=2
The number of experts per token in the TimeMOE model.
- num_experts: int, default=1
The number of experts in the TimeMOE model.
- max_position_embeddings: int, default=32768
The maximum position embeddings in the TimeMOE model.
- rms_norm_eps: float, default=1e-6
The epsilon value for RMS normalization in the TimeMOE model.
- rope_theta: int, default=10000
Initialise theta for RoPE (Rotational Positional Embeddings).
- attention_dropout: float, default=0.1
The dropout rate for attention layers in the TimeMOE model.
- apply_aux_loss: bool, default=True
Whether to apply auxiliary loss in the TimeMOE model.
- router_aux_loss_factor: float, default=0.02
The auxiliary loss factor for the router in the TimeMOE model.
- tie_word_embeddings: bool, default=False
Whether to tie word embeddings in the TimeMOE model.
When
model_path=None, these keys initialize the model architecture from scratch (random weights). Whenmodel_pathis set, user-provided keys override the checkpoint config. Shape-changing overrides cause mismatched weights to be reinitialized randomly; those layers needpretrain/ fine-tuning before they are useful. Ifconfigis omitted with a pretrained path, the checkpoint config is used as-is.- seed: int, optional (default=None)
Seed for reproducibility.
- use_source_package: bool, optional (default=False)
If True, the model will be loaded directly from the source package
TimeMoE. This is useful if you want to bypass the local version of the package or when working in an environment where the latest updates from the source package are needed. If False, the model will be loaded from the local version of package maintained in sktime. To install the source package, follow the instructions here [1].- ignore_deps: bool, optional, default=False
If True, dependency checks will be ignored, and the user is expected to handle the installation of required packages manually. If False, the class will enforce the default dependencies required for Chronos.
- context_lengthint, optional (default=1024)
Sliding-window length for
pretrain. For small datasets, use a shorter length withstride=1.- strideint, optional (default=None)
Sliding-window stride for
pretrain. Defaults tocontext_length.- training_argsdict, optional (default=None)
Keyword arguments used for training. Supports all arguments by
transformers.TrainingArguments[Rf74a85329ff4-3].Additionally, the following arguments are supported: - min_learning_rate: float, default=0
Minimum learning rate for cosine_schedule
- validation_splitfloat or None, default=0.2
Fraction of data reserved for evaluation when
pretrainis used. IfNone, no evaluation dataset is created.- devicestr, optional (default=None)
Device placement passed to transformers
device_map, for example"cpu","cuda", or"auto". IfNone,config["device_map"]is used.- dtypetorch.dtype, optional (default=None)
Torch dtype used when loading the model and preparing prediction inputs.
Nonekeeps the checkpoint’s native dtype forfrom_pretrained(and the initialized dtype for from-scratch models); prediction inputs then follow the loaded model dtype.
- Attributes:
cutoffCut-off = “present time” state of forecaster.
fhForecasting horizon that was passed.
is_fittedWhether
fithas been called.stateState of the estimator.
References
Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts .. [Rf74a85329ff4-3] Trainer/TrainingArguments docs:
Examples
Zero-shot forecasting:
>>> from sktime.forecasting.timemoe import TimeMoEForecaster >>> from sktime.datasets import load_airline >>> from sktime.forecasting.model_selection import temporal_train_test_split >>> y = load_airline() >>> forecaster = TimeMoEForecaster("Maple728/TimeMoE-50M") >>> forecaster.fit(y_train) >>> y_pred = forecaster.predict(fh=[1, 2, 3], y = y_test)
Fine-tuning of a pretrained checkpoint:
>>> from sktime.utils._testing.hierarchical import _make_hierarchical >>> y_panel = _make_hierarchical( ... hierarchy_levels=(3,), min_timepoints=64, max_timepoints=128, ... ) >>> forecaster = TimeMoEForecaster( ... model_path="Maple728/TimeMoE-50M", ... context_length=32, ... stride=1, ... training_args={"max_steps": 10, "per_device_train_batch_size": 2}, ... ) >>> forecaster.pretrain(y_panel) >>> forecaster.fit(load_airline()) >>> y_pred = forecaster.predict(fh=[1, 2, 3])
Training from scratch:
>>> forecaster = TimeMoEForecaster( ... model_path=None, ... config={ ... "hidden_size": 64, ... "intermediate_size": 128, ... "num_hidden_layers": 2, ... "num_attention_heads": 4, ... "num_experts": 2, ... "num_experts_per_tok": 1, ... "horizon_lengths": [1], ... "max_position_embeddings": 128, ... }, ... context_length=32, ... stride=1, ... training_args={"max_steps": 10}, ... ) >>> forecaster.pretrain(y_panel) >>> forecaster.fit(load_airline()) >>> y_pred = forecaster.predict(fh=[1, 2, 3])
Methods
check_is_fitted([method_name])Check if the estimator has been fitted.
clone()Obtain a clone of the object with same hyper-parameters and config.
clone_tags(estimator[, tag_names])Clone tags from another object as dynamic override.
create_test_instance([parameter_set])Construct an instance of the class, using first test parameter set.
create_test_instances_and_names([parameter_set])Create list of all test instances and a list of names for them.
fit(y[, X, fh])Fit forecaster to training data.
fit_predict(y[, X, fh, X_pred])Fit and forecast time series at future horizon.
get_class_tag(tag_name[, tag_value_default])Get class tag value from class, with tag level inheritance from parents.
get_class_tags()Get class tags from class, with tag level inheritance from parent classes.
get_config()Get config flags for self.
get_fitted_params([deep])Get fitted parameters.
get_param_defaults()Get object's parameter defaults.
get_param_names([sort])Get object's parameter names.
get_params([deep])Get a dict of parameters values for this object.
get_pretrained_params([deep])Get pretrained parameters of this estimator.
get_tag(tag_name[, tag_value_default, ...])Get tag value from instance, with tag level inheritance and overrides.
get_tags()Get tags from instance, with tag level inheritance and overrides.
get_test_params([parameter_set])Get the test parameters for the forecaster.
is_composite()Check if the object is composed of other BaseObjects.
load_from_path(serial)Load object from file location.
load_from_serial(serial)Load object from serialized memory container.
predict([fh, X])Forecast time series at future horizon.
predict_interval([fh, X, coverage])Compute/return prediction interval forecasts.
predict_proba([fh, X, marginal])Compute/return fully probabilistic forecasts.
predict_quantiles([fh, X, alpha])Compute/return quantile forecasts.
predict_residuals([y, X])Return residuals of time series forecasts.
predict_var([fh, X, cov])Compute/return variance forecasts.
pretrain(y[, X, fh])Pre-train forecaster on panel (global) data.
reset()Reset the object to a clean post-init state.
save([path, serialization_format])Save serialized self to bytes-like object or to (.zip) file.
score(y[, X, fh])Scores forecast against ground truth, using MAPE (non-symmetric).
set_config(**config_dict)Set config flags to given values.
set_params(**params)Set the parameters of this object.
set_random_state([random_state, deep, ...])Set random_state pseudo-random seed parameters for self.
set_tags(**tag_dict)Set instance level tag overrides to given values.
update(y[, X, update_params])Update cutoff value and, optionally, fitted parameters.
update_predict(y[, cv, X, update_params, ...])Make predictions and update model iteratively over the test set.
update_predict_single([y, fh, X, update_params])Update model with new data and make forecasts.

