Skip to content

Imputer

Imputer

class Imputer(method='drift', random_state=None, value=None, forecaster=None, missing_values=None)[source]

Missing value imputation.

The Imputer transforms input series by replacing missing values according to an imputation strategy specified by method.

Parameters:
methodstr, default=”drift”

Method to fill the missing values. Not all methods can extrapolate, so after method is applied the remaining missing values are filled with ffill then bfill.

  • “drift” : drift/trend values by sktime.PolynomialTrendForecaster(degree=1) first, X in transform() is filled with ffill then bfill then PolynomialTrendForecaster(degree=1) is fitted to filled X, and predict values are queried at indices which had missing values

  • “linear”linear interpolation, uses pd.Series.interpolate()

    WARNING: This method can not extrapolate, so it is fitted always on the data given to transform().

  • “nearest” : use nearest value, uses pd.Series.interpolate()

  • “constant” : same constant value (given in arg value) for all NaN

  • “mean” : pd.Series.mean() of data seen in fit to use data in transform, wrap this estimator in FitInTransform

  • “median” : pd.Series.median() of data seen in fit to use data in transform, wrap this estimator in FitInTransform

  • “backfill” to “bfill” : applies pd.Series.bfill to all data

  • “pad” or “ffill” : applies pd.Series.ffill to all data

  • “random” : random values between pd.Series.min() and .max() of fit data if pd.Series dtype is int, sample is uniform discrete if pd.Series dtype is float, sample is uniform continuous

  • “forecaster” : use an sktime forecaster, given in param forecaster. First, X seed in fit is filled with ffill then bfill then forecaster is fitted to filled X, and predict values are queried at indices of X data in transform which had missing values. forecaster is always applied by variable and instance.

The following methods, fit non-trivially to the data seen in fit: “drift”, “mean”, “median”, “random”. All other methods do not depend on values seen in fit.

random_stateint/float/str, optional

Value to set random.seed() if method=”random”, default None

valueint/float, default=None

Value to use to fill missing values when method=”constant”. Only used if method="constant", otherwise ignored.

forecasterAny Forecaster based on sktime.BaseForecaster, default=None

Use a given Forecaster to impute by insample predictions when method="forecaster". Before fitting, missing data is imputed with method="ffill" or "bfill" as heuristic. In case of multivariate X, a clone of forecaster is applied per column. Only used if method="forecaster", otherwise ignored.

missing_valuesstr, int, float, regex, list, or None, default=None

Value to consider as np.nan` and impute, passed to DataFrame.replace If str, int, float, all entries equal to missing_values will be imputed, in addition to np.nan. If regex, all entries matching regex will be imputed, in addition to np.nan. If list, must be list of str, int, float, or regex. Values matching any list element by above rules will be imputed, in addition to np.nan. If None, then only np.nan values are imputed.

Attributes:
is_fitted

Whether fit has been called.

Examples

>>> from sktime.transformations.impute import Imputer
>>> from sktime.datasets import load_airline
>>> from sktime.split import temporal_train_test_split
>>> y = load_airline()
>>> y_train, y_test = temporal_train_test_split(y)
>>> transformer = Imputer(method="drift")
>>> transformer.fit(y_train)
Imputer(...)
>>> import numpy as np
>>> y_test.iloc[3] = np.nan
>>> y_hat = transformer.transform(y_test)

Methods

check_is_fitted([method_name])

Check if the estimator has been fitted.

clone()

Obtain a clone of the object with same hyper-parameters and config.

clone_tags(estimator[, tag_names])

Clone tags from another object as dynamic override.

create_test_instance([parameter_set])

Construct an instance of the class, using first test parameter set.

create_test_instances_and_names([parameter_set])

Create list of all test instances and a list of names for them.

fit(X[, y])

Fit transformer to X, optionally to y.

fit_transform(X[, y])

Fit to data, then transform it.

get_class_tag(tag_name[, tag_value_default])

Get class tag value from class, with tag level inheritance from parents.

get_class_tags()

Get class tags from class, with tag level inheritance from parent classes.

get_config()

Get config flags for self.

get_fitted_params([deep])

Get fitted parameters.

get_param_defaults()

Get object's parameter defaults.

get_param_names([sort])

Get object's parameter names.

get_params([deep])

Get a dict of parameters values for this object.

get_tag(tag_name[, tag_value_default, ...])

Get tag value from instance, with tag level inheritance and overrides.

get_tags()

Get tags from instance, with tag level inheritance and overrides.

get_test_params([parameter_set])

Return testing parameter settings for the estimator.

inverse_transform(X[, y])

Inverse transform X and return an inverse transformed version.

is_composite()

Check if the object is composed of other BaseObjects.

load_from_path(serial)

Load object from file location.

load_from_serial(serial)

Load object from serialized memory container.

reset()

Reset the object to a clean post-init state.

save([path, serialization_format])

Save serialized self to bytes-like object or to (.zip) file.

set_config(**config_dict)

Set config flags to given values.

set_params(**params)

Set the parameters of this object.

set_random_state([random_state, deep, ...])

Set random_state pseudo-random seed parameters for self.

set_tags(**tag_dict)

Set instance level tag overrides to given values.

transform(X[, y])

Transform X and return a transformed version.

update(X[, y, update_params])

Update transformer with X, optionally y.