SyncToLongest
SyncToLongest
- class SyncToLongest(base_cv: BaseSplitter, min_length: int | None = None)[source]
Wrap a temporal splitter to make it work with unequal-length panels.
Folds (cutoffs) come from the longest series in the panel. For each fold, every other instance is included with as much of its own history as is available up to that cutoff, as long as it covers the full forecasting horizon and has enough training history. Instances that don’t qualify for a fold are excluded from it, rather than truncated to the shortest instance the way
base_cvwould by default on a panel.Each split is itself a panel containing whichever instances qualify for that fold, so different folds can contain a different number of instances.
This assumes all instances share a common, comparable time index, for example because they’re all cut off at the same latest timestamp. If
yis a single (non-panel) series, or every instance has equal length and shares the same time index,SyncToLongestbehaves exactly likebase_cv.- Parameters:
- base_cvsktime splitter, BaseSplitter descendant instance
the underlying temporal splitter, e.g.,
SlidingWindowSplitter- min_lengthint, optional, default=None
minimum number of training observations an instance needs to be included in a fold. Must be > 0. If None (default), only instances with a full training window (as long as the longest instance’s for that fold) are included. If set, instances with at least
min_lengthobservations before the cutoff are included too, with all the history they have.
Notes
Which instances qualify for a fold is decided from whatever panel is passed to
split/split_series. If this splitter is used as (part of) an explicitcv_Xinevaluate, wrap it withSameLocSplitteraroundy(e.g.SameLocSplitter(TestPlusTrainSplitter(cv), y), the same compositionevaluateuses by default) soXgets the same per-fold instance selection asy, instead ofXbeing split independently based on its own coverage. Without that, anXwhose per-instance coverage differs fromy(e.g. exogenous data starting later for some instance) silently yields folds whereXandydisagree on which instances are included.Examples
>>> from sktime.split import SlidingWindowSplitter >>> from sktime.split.compose import SyncToLongest >>> from sktime.utils._testing.hierarchical import _make_hierarchical >>> y = _make_hierarchical(hierarchy_levels=(2,), max_timepoints=10, ... min_timepoints=7, random_state=42) >>> cv = SyncToLongest(SlidingWindowSplitter(window_length=3, fh=1)) >>> def train_ranges(train): ... # per-series (start..end) of the train window, keyed by series id ... idx = y.index[train] ... times = idx.get_level_values(-1) ... series = idx.droplevel(-1) ... return ", ".join( ... f"{s}: {times[series == s].min():%m-%d}" ... f"..{times[series == s].max():%m-%d}" ... for s in sorted(series.unique()) ... ) >>> for i, (train, test) in enumerate(cv.split(y)): ... print(f"fold {i}: {train_ranges(train)}") fold 0: h0_0: 01-02..01-04 fold 1: h0_0: 01-03..01-05 fold 2: h0_0: 01-04..01-06, h0_1: 01-04..01-06 fold 3: h0_0: 01-05..01-07, h0_1: 01-05..01-07 fold 4: h0_0: 01-06..01-08, h0_1: 01-06..01-08 fold 5: h0_0: 01-07..01-09, h0_1: 01-07..01-09
Every included series gets the same 3-day calendar window.
base_cvalone positions each series’ window relative to its own local index instead, so windows of equal length land on different calendar dates per series (h0_1starts 2 days later thanh0_0, and that offset carries through every fold):>>> for i, (train, test) in enumerate(cv.base_cv.split(y)): ... print(f"fold {i}: {train_ranges(train)}") fold 0: h0_0: 01-02..01-04, h0_1: 01-04..01-06 fold 1: h0_0: 01-03..01-05, h0_1: 01-05..01-07 fold 2: h0_0: 01-04..01-06, h0_1: 01-06..01-08 fold 3: h0_0: 01-05..01-07, h0_1: 01-07..01-09
Methods
clone()Obtain a clone of the object with same hyper-parameters and config.
clone_tags(estimator[, tag_names])Clone tags from another object as dynamic override.
create_test_instance([parameter_set])Construct an instance of the class, using first test parameter set.
create_test_instances_and_names([parameter_set])Create list of all test instances and a list of names for them.
get_class_tag(tag_name[, tag_value_default])Get class tag value from class, with tag level inheritance from parents.
get_class_tags()Get class tags from class, with tag level inheritance from parent classes.
get_config()Get config flags for self.
get_cutoffs([y])Return the cutoff points in .iloc[] context.
get_fh()Forecasting horizon, in ForecastingHorizon format, relative to cutoff.
get_n_splits(y)Return the number of splits.
get_param_defaults()Get object's parameter defaults.
get_param_names([sort])Get object's parameter names.
get_params([deep])Get a dict of parameters values for this object.
get_tag(tag_name[, tag_value_default, ...])Get tag value from instance, with tag level inheritance and overrides.
get_tags()Get tags from instance, with tag level inheritance and overrides.
get_test_params([parameter_set])Return testing parameter settings for the splitter.
is_composite()Check if the object is composed of other BaseObjects.
load_from_path(serial)Load object from file location.
load_from_serial(serial)Load object from serialized memory container.
reset()Reset the object to a clean post-init state.
save([path, serialization_format])Save serialized self to bytes-like object or to (.zip) file.
set_config(**config_dict)Set config flags to given values.
set_params(**params)Set the parameters of this object.
set_random_state([random_state, deep, ...])Set random_state pseudo-random seed parameters for self.
set_tags(**tag_dict)Set instance level tag overrides to given values.
split(y)Get iloc references to train/test splits of y.
split_loc(y)Get loc references to train/test splits of y.
split_series(y)Split y into training and test windows.

