Skip to content

AutoResearchForecaster

AutoResearchForecaster

class AutoResearchForecaster(cv, model='openai/gpt-4o-mini', n_iterations=3, n_blueprints=5, n_fix_attempts=0, api_params=None, system_prompt=None, refinement_prompt=None, llm_func=None, description_method='basic', estimator_info='names')[source]

Forecaster that uses an LLM to generate and refine sktime pipeline blueprints.

Inspired by Karpathy’s autoresearch project, this forecaster:

  1. Asks an LLM to propose diverse forecasting pipeline blueprints

  2. Evaluates each blueprint on a validation split of the training data

  3. Feeds results back to the LLM for iterative refinement

  4. Selects the best-performing blueprint as the final forecaster

Parameters:
modelstr, default=”openai/gpt-4o-mini”

LLM model identifier compatible with litellm (e.g., “openai/gpt-4o-mini”, “anthropic/claude-sonnet-4-20250514”).

n_iterationsint, default=3

Number of generate-evaluate-refine iterations.

n_blueprintsint, default=5

Number of blueprints to generate per iteration.

n_fix_attemptsint, default=0

Number of additional LLM calls to attempt fixing each failed blueprint. For each blueprint that fails evaluation, the LLM is asked to correct the spec up to n_fix_attempts times. Set to 0 to disable.

api_paramsdict or None, default=None

Additional keyword arguments passed to litellm.completion (e.g., temperature, max_tokens, api_key).

system_promptstr or None, default=None

Custom system prompt for the LLM. If None, uses the default prompt. The prompt should contain placeholders for {n_blueprints}, {forecaster_names}, and {transformer_names} which are filled in automatically.

refinement_promptstr or None, default=None

Custom refinement prompt template for the LLM. If None, uses the default prompt. The prompt should contain placeholders for {results_summary}, {all_results_ranked}, {best_name}, {best_score}, and {n_blueprints} which are filled in automatically.

llm_funccallable or None, default=None

Custom callable to invoke the LLM. If None, uses litellm.completion via the internal _call_llm function. The callable must have the signature llm_func(messages, model, api_params) -> str, where messages is a list of chat message dicts, model is the model identifier string, api_params is a dict of extra kwargs, and the return value is the raw response text. Primarily useful for testing without an API key.

description_methodstr, default=”basic”

Method for generating dataset description for the LLM. Options:

  • “basic”: Text-only statistics (length, frequency, mean, std, etc.)

  • “described_plot”: Generates a plot and uses a vision LLM to describe it, combined with basic statistics.

  • “image”: Generates a plot and provides it as an image to the blueprint generation LLM (requires vision-capable model).

Vision-based methods (“described_plot”, “image”) require a model that supports image input, which is checked during initialization.

Attributes:
best_blueprint_dict

The best-performing blueprint found during fitting.

best_score_float

The validation score of the best blueprint.

best_forecaster_BaseForecaster

The fitted forecaster from the best blueprint.

blueprint_history_list of dict

History of all evaluated blueprints and their scores.

llm_conversation_list of dict

The full LLM conversation history.

Examples

>>> from sktime_autoresearch import AutoResearchForecaster
>>> from sktime.datasets import load_airline
>>> from sktime.split import SingleWindowSplitter
>>> y = load_airline()
>>> forecaster = AutoResearchForecaster(
...     cv=SingleWindowSplitter(fh=[1, 2, 3]),
...     model="openai/gpt-4o-mini",
...     n_iterations=2,
...     n_blueprints=3,
... )
>>> forecaster.fit(y, fh=[1, 2, 3])
>>> y_pred = forecaster.predict(fh=[1, 2, 3])

Methods

check_is_fitted([method_name])

Check if the estimator has been fitted.

clone()

Obtain a clone of the object with same hyper-parameters and config.

clone_tags(estimator[, tag_names])

Clone tags from another object as dynamic override.

create_test_instance([parameter_set])

Construct an instance of the class, using first test parameter set.

create_test_instances_and_names([parameter_set])

Create list of all test instances and a list of names for them.

fit(y[, X, fh])

Fit forecaster to training data.

fit_predict(y[, X, fh, X_pred])

Fit and forecast time series at future horizon.

get_class_tag(tag_name[, tag_value_default])

Get class tag value from class, with tag level inheritance from parents.

get_class_tags()

Get class tags from class, with tag level inheritance from parent classes.

get_config()

Get config flags for self.

get_fitted_params([deep])

Get fitted parameters.

get_param_defaults()

Get object's parameter defaults.

get_param_names([sort])

Get object's parameter names.

get_params([deep])

Get a dict of parameters values for this object.

get_pretrained_params([deep])

Get pretrained parameters of this estimator.

get_tag(tag_name[, tag_value_default, ...])

Get tag value from instance, with tag level inheritance and overrides.

get_tags()

Get tags from instance, with tag level inheritance and overrides.

get_test_params([parameter_set])

Return testing parameter settings for the estimator.

is_composite()

Check if the object is composed of other BaseObjects.

load_from_path(serial)

Load object from file location.

load_from_serial(serial)

Load object from serialized memory container.

predict([fh, X])

Forecast time series at future horizon.

predict_interval([fh, X, coverage])

Compute/return prediction interval forecasts.

predict_proba([fh, X, marginal])

Compute/return fully probabilistic forecasts.

predict_quantiles([fh, X, alpha])

Compute/return quantile forecasts.

predict_residuals([y, X])

Return residuals of time series forecasts.

predict_var([fh, X, cov])

Compute/return variance forecasts.

pretrain(y[, X, fh])

Pre-train forecaster on panel (global) data.

reset()

Reset the object to a clean post-init state.

save([path, serialization_format])

Save serialized self to bytes-like object or to (.zip) file.

score(y[, X, fh])

Scores forecast against ground truth, using MAPE (non-symmetric).

set_config(**config_dict)

Set config flags to given values.

set_params(**params)

Set the parameters of this object.

set_random_state([random_state, deep, ...])

Set random_state pseudo-random seed parameters for self.

set_tags(**tag_dict)

Set instance level tag overrides to given values.

summary()

Return a summary DataFrame of all evaluated blueprints.

update(y[, X, update_params])

Update cutoff value and, optionally, fitted parameters.

update_predict(y[, cv, X, update_params, ...])

Make predictions and update model iteratively over the test set.

update_predict_single([y, fh, X, update_params])

Update model with new data and make forecasts.