AutoResearchForecaster
AutoResearchForecaster
- class AutoResearchForecaster(cv, model='openai/gpt-4o-mini', n_iterations=3, n_blueprints=5, n_fix_attempts=0, api_params=None, system_prompt=None, refinement_prompt=None, llm_func=None, description_method='basic', estimator_info='names')[source]
Forecaster that uses an LLM to generate and refine sktime pipeline blueprints.
Inspired by Karpathy’s autoresearch project, this forecaster:
Asks an LLM to propose diverse forecasting pipeline blueprints
Evaluates each blueprint on a validation split of the training data
Feeds results back to the LLM for iterative refinement
Selects the best-performing blueprint as the final forecaster
- Parameters:
- modelstr, default=”openai/gpt-4o-mini”
LLM model identifier compatible with litellm (e.g., “openai/gpt-4o-mini”, “anthropic/claude-sonnet-4-20250514”).
- n_iterationsint, default=3
Number of generate-evaluate-refine iterations.
- n_blueprintsint, default=5
Number of blueprints to generate per iteration.
- n_fix_attemptsint, default=0
Number of additional LLM calls to attempt fixing each failed blueprint. For each blueprint that fails evaluation, the LLM is asked to correct the spec up to
n_fix_attemptstimes. Set to 0 to disable.- api_paramsdict or None, default=None
Additional keyword arguments passed to litellm.completion (e.g., temperature, max_tokens, api_key).
- system_promptstr or None, default=None
Custom system prompt for the LLM. If None, uses the default prompt. The prompt should contain placeholders for {n_blueprints}, {forecaster_names}, and {transformer_names} which are filled in automatically.
- refinement_promptstr or None, default=None
Custom refinement prompt template for the LLM. If None, uses the default prompt. The prompt should contain placeholders for {results_summary}, {all_results_ranked}, {best_name}, {best_score}, and {n_blueprints} which are filled in automatically.
- llm_funccallable or None, default=None
Custom callable to invoke the LLM. If None, uses litellm.completion via the internal
_call_llmfunction. The callable must have the signaturellm_func(messages, model, api_params) -> str, wheremessagesis a list of chat message dicts,modelis the model identifier string,api_paramsis a dict of extra kwargs, and the return value is the raw response text. Primarily useful for testing without an API key.- description_methodstr, default=”basic”
Method for generating dataset description for the LLM. Options:
“basic”: Text-only statistics (length, frequency, mean, std, etc.)
“described_plot”: Generates a plot and uses a vision LLM to describe it, combined with basic statistics.
“image”: Generates a plot and provides it as an image to the blueprint generation LLM (requires vision-capable model).
Vision-based methods (“described_plot”, “image”) require a model that supports image input, which is checked during initialization.
- Attributes:
- best_blueprint_dict
The best-performing blueprint found during fitting.
- best_score_float
The validation score of the best blueprint.
- best_forecaster_BaseForecaster
The fitted forecaster from the best blueprint.
- blueprint_history_list of dict
History of all evaluated blueprints and their scores.
- llm_conversation_list of dict
The full LLM conversation history.
Examples
>>> from sktime_autoresearch import AutoResearchForecaster >>> from sktime.datasets import load_airline >>> from sktime.split import SingleWindowSplitter >>> y = load_airline() >>> forecaster = AutoResearchForecaster( ... cv=SingleWindowSplitter(fh=[1, 2, 3]), ... model="openai/gpt-4o-mini", ... n_iterations=2, ... n_blueprints=3, ... ) >>> forecaster.fit(y, fh=[1, 2, 3]) >>> y_pred = forecaster.predict(fh=[1, 2, 3])
Methods
check_is_fitted([method_name])Check if the estimator has been fitted.
clone()Obtain a clone of the object with same hyper-parameters and config.
clone_tags(estimator[, tag_names])Clone tags from another object as dynamic override.
create_test_instance([parameter_set])Construct an instance of the class, using first test parameter set.
create_test_instances_and_names([parameter_set])Create list of all test instances and a list of names for them.
fit(y[, X, fh])Fit forecaster to training data.
fit_predict(y[, X, fh, X_pred])Fit and forecast time series at future horizon.
get_class_tag(tag_name[, tag_value_default])Get class tag value from class, with tag level inheritance from parents.
get_class_tags()Get class tags from class, with tag level inheritance from parent classes.
get_config()Get config flags for self.
get_fitted_params([deep])Get fitted parameters.
get_param_defaults()Get object's parameter defaults.
get_param_names([sort])Get object's parameter names.
get_params([deep])Get a dict of parameters values for this object.
get_pretrained_params([deep])Get pretrained parameters of this estimator.
get_tag(tag_name[, tag_value_default, ...])Get tag value from instance, with tag level inheritance and overrides.
get_tags()Get tags from instance, with tag level inheritance and overrides.
get_test_params([parameter_set])Return testing parameter settings for the estimator.
is_composite()Check if the object is composed of other BaseObjects.
load_from_path(serial)Load object from file location.
load_from_serial(serial)Load object from serialized memory container.
predict([fh, X])Forecast time series at future horizon.
predict_interval([fh, X, coverage])Compute/return prediction interval forecasts.
predict_proba([fh, X, marginal])Compute/return fully probabilistic forecasts.
predict_quantiles([fh, X, alpha])Compute/return quantile forecasts.
predict_residuals([y, X])Return residuals of time series forecasts.
predict_var([fh, X, cov])Compute/return variance forecasts.
pretrain(y[, X, fh])Pre-train forecaster on panel (global) data.
reset()Reset the object to a clean post-init state.
save([path, serialization_format])Save serialized self to bytes-like object or to (.zip) file.
score(y[, X, fh])Scores forecast against ground truth, using MAPE (non-symmetric).
set_config(**config_dict)Set config flags to given values.
set_params(**params)Set the parameters of this object.
set_random_state([random_state, deep, ...])Set random_state pseudo-random seed parameters for self.
set_tags(**tag_dict)Set instance level tag overrides to given values.
summary()Return a summary DataFrame of all evaluated blueprints.
update(y[, X, update_params])Update cutoff value and, optionally, fitted parameters.
update_predict(y[, cv, X, update_params, ...])Make predictions and update model iteratively over the test set.
update_predict_single([y, fh, X, update_params])Update model with new data and make forecasts.

