Skip to content

ExpandingGreedySplitter

ExpandingGreedySplitter

class ExpandingGreedySplitter(test_size: int | float, folds: int = 5, step_length: int | None = None)[source]

Splitter that successively cuts test folds off the end of the series.

Takes an integer test_size that defines the number of steps included in the test set of each fold. The train set of each fold will contain all data before the test set. If the data contains multiple instances, test_size is _per instance_.

If no step_length is defined, the test sets (one for each fold) will be adjacent and disjoint, taken from the end of the dataset.

For example, with test_size=7 and folds=5, the test sets in total will cover the last 35 steps of the data with no overlap.

Parameters:
test_sizeint or float
If int: the number of steps included in the test set of each fold.

Formally, steps are consecutive iloc indices.

If float: the proportion of steps included in the test set of each fold,

as a proportion of the total number of consecutive iloc indices. Must be between 0.0 and 1.0. Proportions are rounded to the next higher integer count of samples (ceil). Cave: not the loc proportion between start and end locations, but a proportion of total number of consecutive iloc indices.

foldsint, default = 5

The number of folds.

step_lengthint, optional

The number of steps advanced for each fold. Defaults to test_size.

Examples

>>> import numpy as np
>>> from sktime.split import ExpandingGreedySplitter
>>> ts = np.arange(10)
>>> splitter = ExpandingGreedySplitter(test_size=3, folds=2)
>>> list(splitter.split(ts))
[
    (array([0, 1, 2, 3]), array([4, 5, 6])),
    (array([0, 1, 2, 3, 4, 5, 6]), array([7, 8, 9]))
]

Methods

clone()

Obtain a clone of the object with same hyper-parameters and config.

clone_tags(estimator[, tag_names])

Clone tags from another object as dynamic override.

create_test_instance([parameter_set])

Construct an instance of the class, using first test parameter set.

create_test_instances_and_names([parameter_set])

Create list of all test instances and a list of names for them.

get_class_tag(tag_name[, tag_value_default])

Get class tag value from class, with tag level inheritance from parents.

get_class_tags()

Get class tags from class, with tag level inheritance from parent classes.

get_config()

Get config flags for self.

get_cutoffs([y])

Return the cutoff points in .iloc[] context.

get_fh()

Return the forecasting horizon.

get_n_splits([y])

Return the number of splits.

get_param_defaults()

Get object's parameter defaults.

get_param_names([sort])

Get object's parameter names.

get_params([deep])

Get a dict of parameters values for this object.

get_tag(tag_name[, tag_value_default, ...])

Get tag value from instance, with tag level inheritance and overrides.

get_tags()

Get tags from instance, with tag level inheritance and overrides.

get_test_params([parameter_set])

Return testing parameter settings for the splitter.

is_composite()

Check if the object is composed of other BaseObjects.

load_from_path(serial)

Load object from file location.

load_from_serial(serial)

Load object from serialized memory container.

reset()

Reset the object to a clean post-init state.

save([path, serialization_format])

Save serialized self to bytes-like object or to (.zip) file.

set_config(**config_dict)

Set config flags to given values.

set_params(**params)

Set the parameters of this object.

set_random_state([random_state, deep, ...])

Set random_state pseudo-random seed parameters for self.

set_tags(**tag_dict)

Set instance level tag overrides to given values.

split(y)

Get iloc references to train/test splits of y.

split_loc(y)

Get loc references to train/test splits of y.

split_series(y)

Split y into training and test windows.