Skip to content

SyncToLongest

SyncToLongest

class SyncToLongest(base_cv: BaseSplitter, min_length: int | None = None)[source]

Wrap a temporal splitter to make it work with unequal-length panels.

Folds (cutoffs) come from the longest series in the panel. For each fold, every other instance is included with as much of its own history as is available up to that cutoff, as long as it covers the full forecasting horizon and has enough training history. Instances that don’t qualify for a fold are excluded from it, rather than truncated to the shortest instance the way base_cv would by default on a panel.

Each split is itself a panel containing whichever instances qualify for that fold, so different folds can contain a different number of instances.

This assumes all instances share a common, comparable time index, for example because they’re all cut off at the same latest timestamp. If y is a single (non-panel) series, or every instance has equal length and shares the same time index, SyncToLongest behaves exactly like base_cv.

Parameters:
base_cvsktime splitter, BaseSplitter descendant instance

the underlying temporal splitter, e.g., SlidingWindowSplitter

min_lengthint, optional, default=None

minimum number of training observations an instance needs to be included in a fold. Must be > 0. If None (default), only instances with a full training window (as long as the longest instance’s for that fold) are included. If set, instances with at least min_length observations before the cutoff are included too, with all the history they have.

Notes

Which instances qualify for a fold is decided from whatever panel is passed to split/split_series. If this splitter is used as (part of) an explicit cv_X in evaluate, wrap it with SameLocSplitter around y (e.g. SameLocSplitter(TestPlusTrainSplitter(cv), y), the same composition evaluate uses by default) so X gets the same per-fold instance selection as y, instead of X being split independently based on its own coverage. Without that, an X whose per-instance coverage differs from y (e.g. exogenous data starting later for some instance) silently yields folds where X and y disagree on which instances are included.

Examples

>>> from sktime.split import SlidingWindowSplitter
>>> from sktime.split.compose import SyncToLongest
>>> from sktime.utils._testing.hierarchical import _make_hierarchical
>>> y = _make_hierarchical(hierarchy_levels=(2,), max_timepoints=10,
...     min_timepoints=7, random_state=42)
>>> cv = SyncToLongest(SlidingWindowSplitter(window_length=3, fh=1))
>>> def train_ranges(train):
...     # per-series (start..end) of the train window, keyed by series id
...     idx = y.index[train]
...     times = idx.get_level_values(-1)
...     series = idx.droplevel(-1)
...     return ", ".join(
...         f"{s}: {times[series == s].min():%m-%d}"
...         f"..{times[series == s].max():%m-%d}"
...         for s in sorted(series.unique())
...     )
>>> for i, (train, test) in enumerate(cv.split(y)):
...     print(f"fold {i}: {train_ranges(train)}")
fold 0: h0_0: 01-02..01-04
fold 1: h0_0: 01-03..01-05
fold 2: h0_0: 01-04..01-06, h0_1: 01-04..01-06
fold 3: h0_0: 01-05..01-07, h0_1: 01-05..01-07
fold 4: h0_0: 01-06..01-08, h0_1: 01-06..01-08
fold 5: h0_0: 01-07..01-09, h0_1: 01-07..01-09

Every included series gets the same 3-day calendar window. base_cv alone positions each series’ window relative to its own local index instead, so windows of equal length land on different calendar dates per series (h0_1 starts 2 days later than h0_0, and that offset carries through every fold):

>>> for i, (train, test) in enumerate(cv.base_cv.split(y)):
...     print(f"fold {i}: {train_ranges(train)}")
fold 0: h0_0: 01-02..01-04, h0_1: 01-04..01-06
fold 1: h0_0: 01-03..01-05, h0_1: 01-05..01-07
fold 2: h0_0: 01-04..01-06, h0_1: 01-06..01-08
fold 3: h0_0: 01-05..01-07, h0_1: 01-07..01-09

Methods

clone()

Obtain a clone of the object with same hyper-parameters and config.

clone_tags(estimator[, tag_names])

Clone tags from another object as dynamic override.

create_test_instance([parameter_set])

Construct an instance of the class, using first test parameter set.

create_test_instances_and_names([parameter_set])

Create list of all test instances and a list of names for them.

get_class_tag(tag_name[, tag_value_default])

Get class tag value from class, with tag level inheritance from parents.

get_class_tags()

Get class tags from class, with tag level inheritance from parent classes.

get_config()

Get config flags for self.

get_cutoffs([y])

Return the cutoff points in .iloc[] context.

get_fh()

Forecasting horizon, in ForecastingHorizon format, relative to cutoff.

get_n_splits(y)

Return the number of splits.

get_param_defaults()

Get object's parameter defaults.

get_param_names([sort])

Get object's parameter names.

get_params([deep])

Get a dict of parameters values for this object.

get_tag(tag_name[, tag_value_default, ...])

Get tag value from instance, with tag level inheritance and overrides.

get_tags()

Get tags from instance, with tag level inheritance and overrides.

get_test_params([parameter_set])

Return testing parameter settings for the splitter.

is_composite()

Check if the object is composed of other BaseObjects.

load_from_path(serial)

Load object from file location.

load_from_serial(serial)

Load object from serialized memory container.

reset()

Reset the object to a clean post-init state.

save([path, serialization_format])

Save serialized self to bytes-like object or to (.zip) file.

set_config(**config_dict)

Set config flags to given values.

set_params(**params)

Set the parameters of this object.

set_random_state([random_state, deep, ...])

Set random_state pseudo-random seed parameters for self.

set_tags(**tag_dict)

Set instance level tag overrides to given values.

split(y)

Get iloc references to train/test splits of y.

split_loc(y)

Get loc references to train/test splits of y.

split_series(y)

Split y into training and test windows.