SyncToLongest
Wrap a temporal splitter to make it work with unequal-length panels.
Folds (cutoffs) come from the longest series in the panel. For each fold, every other instance is included with as much of its own history as is available up to that cutoff, as long as it covers the full forecasting horizon and has enough training history. Instances that don’t qualify for a fold are excluded from it, rather than truncated to the shortest instance the way base_cv would by default on a panel.
Each split is itself a panel containing whichever instances qualify for that fold, so different folds can contain a different number of instances.
This assumes all instances share a common, comparable time index, for example because they’re all cut off at the same latest timestamp. If y is a single (non-panel) series, or every instance has equal length and shares the same time index, SyncToLongest behaves exactly like base_cv.
Quickstart
from sktime.split.compose import SyncToLongest
estimator = SyncToLongest(base_cv: BaseSplitter, min_length: int | None=None)Parameters(2)
- base_cvsktime splitter, BaseSplitter descendant instance
the underlying temporal splitter, e.g.,
SlidingWindowSplitter- min_lengthint, optional, default=None
minimum number of training observations an instance needs to be included in a fold. Must be > 0. If None (default), only instances with a full training window (as long as the longest instance’s for that fold) are included. If set, instances with at least
min_lengthobservations before the cutoff are included too, with all the history they have.
Examples
>>> from sktime.split import SlidingWindowSplitter
>>> from sktime.split.compose import SyncToLongest
>>> from sktime.utils._testing.hierarchical import _make_hierarchical
>>> y = _make_hierarchical (hierarchy_levels = (2,), max_timepoints = 10,
... min_timepoints = 7, random_state = 42)
>>> cv = SyncToLongest (SlidingWindowSplitter (window_length = 3, fh = 1))
>>> def train_ranges (train):
... # per-series (start..end) of the train window, keyed by series id
... idx = y. index [train ]
... times = idx. get_level_values (- 1)
... series = idx. droplevel (- 1)
... return ", ". join (
... f " { s }: { times [series == s ]. min (): %m-%d } "
... f ".. { times [series == s ]. max (): %m-%d } "
... for s in sorted (series. unique ())
... )
>>> for i, (train, test) in enumerate (cv. split (y)):
... print (f "fold { i }: { train_ranges (train) } ") fold 0: h0_0: 01-02..01-04 fold 1: h0_0: 01-03..01-05 fold 2: h0_0: 01-04..01-06, h0_1: 01-04..01-06 fold 3: h0_0: 01-05..01-07, h0_1: 01-05..01-07 fold 4: h0_0: 01-06..01-08, h0_1: 01-06..01-08 fold 5: h0_0: 01-07..01-09, h0_1: 01-07..01-09 Every included series gets the same 3-day calendar window. base_cv alone positions each series’ window relative to its own local index instead, so windows of equal length land on different calendar dates per series (h0_1 starts 2 days later than h0_0, and that offset carries through every fold):
>>> for i, (train, test) in enumerate (cv. base_cv. split (y)):
... print (f "fold { i }: { train_ranges (train) } ") fold 0: h0_0: 01-02..01-04, h0_1: 01-04..01-06 fold 1: h0_0: 01-03..01-05, h0_1: 01-05..01-07 fold 2: h0_0: 01-04..01-06, h0_1: 01-06..01-08 fold 3: h0_0: 01-05..01-07, h0_1: 01-07..01-09