ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction

August 24, 2026 ยท View on GitHub

ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction

arXiv Project Code

๐Ÿš€ Overview

Open-web future event prediction requires agents to identify reliable signals in noisy, redundant, incomplete, and sometimes conflicting evidence. ForeDreamer treats this as an evidence-to-memory transformation problem: raw search results are converted into structured, question-specific factual memory before the forecasting agent reasons over them.

ForeDreamer framework overview

๐Ÿ“– Method

ForeDreamer separates two forms of memory:

  • Factual memory is the processed evidence state for one forecasting question.
  • Experiential memory persists across forecasting episodes and guides future search, evidence processing, and prediction.

The framework combines:

  • a main agent that plans cutoff-aware web searches and produces forecasts;
  • a memory-processing subagent that follows a MemGuide and executes sandboxed MemTools to transform search results into factual memory;
  • textual experience evolution that updates an Experience Bank for search planning, evidence integration, and calibration;
  • procedural experience evolution that updates MemGuides and executable MemTools;
  • Compositional Tool Reuse and Diversity-Guided Exploration for less redundant and more diverse procedural evolution.

This repository focuses on the minimum code path required to evolve and evaluate ForeDreamer. Baseline implementations, ablation runners, large batch schedulers, and historical experiment outputs are not included.

๐Ÿ—‚๏ธ Repository Structure

.
โ”œโ”€โ”€ data/
โ”‚   โ”œโ”€โ”€ FutureX/
โ”‚   โ””โ”€โ”€ Prophet-arena/
โ”œโ”€โ”€ scripts/
โ”‚   โ”œโ”€โ”€ evolve.sh
โ”‚   โ””โ”€โ”€ test.sh
โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ SelfEvolving/        # textual and procedural evolution
โ”‚   โ”œโ”€โ”€ prediction/          # datasets, runners, and metrics
โ”‚   โ”œโ”€โ”€ DefaultTool/         # search-result processing runtime
โ”‚   โ”œโ”€โ”€ MemGuide/            # initial evidence-processing guide
โ”‚   โ”œโ”€โ”€ MemTool/             # initial tools and sandbox runtime
โ”‚   โ”œโ”€โ”€ experience_bank.py
โ”‚   โ”œโ”€โ”€ run_eval.py
โ”‚   โ””โ”€โ”€ summarize_test_results.py
โ””โ”€โ”€ requirements.txt

The two main entry points are:

  • scripts/evolve.sh: evolve the Experience Bank, MemGuides, and MemTools, then select the best validated assets.
  • scripts/test.sh: evaluate the selected assets and produce both per-example predictions and aggregate scores.

โš™๏ธ Getting Started

Requirements

  • Linux
  • Python 3.10 or newer
  • Bubblewrap (bwrap) for sandboxed MemTool execution
  • an OpenAI-compatible LLM API
  • a Tavily or Firecrawl search API

Environment Setup

Install the system dependency on Ubuntu/Debian:

sudo apt-get update
sudo apt-get install -y bubblewrap

Create an environment and install the Python dependencies:

conda create -n foredreamer python=3.10 -y
conda activate foredreamer

python -m pip install --upgrade pip
python -m pip install -r requirements.txt

export PYTHON_BIN="$(which python)"

pm-rank is used to compute the Prophet Arena Brier score and average return.

๐Ÿ”‘ API Configuration

ForeDreamer reads all credentials from environment variables. Do not write API keys into scripts or commit them to the repository.

LLM API

For OpenAI or another OpenAI-compatible provider:

export OPENAI_API_KEY="your-llm-api-key"
export OPENAI_MODEL="your-model-name"
export OPENAI_BASE_URL="https://api.openai.com/v1"

OPENAI_BASE_URL defaults to https://api.openai.com/v1 and may be omitted when using OpenAI directly.

OpenRouter example:

export OPENROUTER_API_KEY="your-openrouter-key"
export OPENAI_BASE_URL="https://openrouter.ai/api/v1"
export OPENAI_MODEL="provider/model-name"

Search API

Tavily is the default search provider:

export SEARCH_PROVIDER="tavily"
export TAVILY_API_KEY="your-tavily-key"

To use Firecrawl:

export SEARCH_PROVIDER="firecrawl"
export FIRECRAWL_API_KEY="your-firecrawl-key"

Both evolution and testing call the LLM and search APIs and may incur usage charges.

๐Ÿ“ฆ Data

FutureX

FilePurpose
data/FutureX/train-20of208.parquet20 examples used for evolution and validation
data/FutureX/train.parquetall 208 resolved examples used as the test input

Following the paper protocol, the 20 evolution/validation examples are excluded when reporting the held-out score. The remaining evaluation set contains 188 examples.

Prophet Arena

The repository contains the eight Prophet Arena categories used in the paper:

CategoryEvolution/validation fileEvaluation input
Climate and Weathersubset_data_Climate_and_Weather_train_5.csvsubset_data_Climate_and_Weather_13.csv
Companiessubset_data_Companies_train_5.csvsubset_data_Companies_27.csv
Economicssubset_data_Economics_train_5.csvsubset_data_Economics_19.csv
Entertainmentsubset_data_Entertainment_train_5.csvsubset_data_Entertainment_93.csv
Mentionssubset_data_Mentions_train_5.csvsubset_data_Mentions_26.csv
Othersubset_data_Other_train_5.csvsubset_data_Other_37.csv
Politicssubset_data_Politics_train_5.csvsubset_data_Politics_91.csv
Sportssubset_data_Sports_train_5.csvsubset_data_Sports_200.csv

data/Prophet-arena/subset_data_1200.csv contains the aggregate 1,200-example set.

๐Ÿงฌ Evolve and Evaluate: FutureX

The default evolution command uses the 20-example FutureX evolution/validation set and performs three update iterations as a lower-cost functional run:

bash scripts/evolve.sh

Generated assets are written under runs/futurex/. The selected MemGuide and Experience Bank are recorded in runs/futurex/best_assets.json.

Test all 208 FutureX examples:

bash scripts/test.sh

The test script uses:

TRAIN_DATA_PATH=data/FutureX/train-20of208.parquet
INPUT_PATH=data/FutureX/train.parquet

It predicts all 208 inputs once and reports both:

  • the score on all 208 examples;
  • the held-out score on the 188 examples remaining after removing the 20 evolution/validation IDs.

To test one or more comma-separated FutureX sample IDs:

RUN_SPECIFIC="sample-id-1,sample-id-2" bash scripts/test.sh

To manually select a MemGuide:

MEM_GUIDE="guide_2.json" bash scripts/test.sh

๐Ÿ”ฌ Paper-scale FutureX Configuration

The paper uses 20 evolution/validation examples, 60 update iterations, four search results per request, and a 1:4 exploration-to-expansion ratio:

DATASET_TYPE=futurex \
TRAIN_DATA_PATH="./data/FutureX/train-20of208.parquet" \
VAL_DATA_PATH="./data/FutureX/train-20of208.parquet" \
RUN_DIR="./runs/futurex-paper" \
NUM_ITERATIONS=60 \
MEMGUIDE_ROUNDS_PER_CYCLE=10 \
EXPERIENCE_ROUNDS_PER_CYCLE=5 \
EXPLORATION_OVER_EXPANSION=1:4 \
SEARCH_MAX_RESULTS=4 \
PARALLELISM=2 \
bash scripts/evolve.sh

Evaluate the resulting assets:

DATASET_TYPE=futurex \
RUN_DIR="./runs/futurex-paper" \
TRAIN_DATA_PATH="./data/FutureX/train-20of208.parquet" \
INPUT_PATH="./data/FutureX/train.parquet" \
SEARCH_MAX_RESULTS=4 \
bash scripts/test.sh

๐Ÿ”ฎ Prophet Arena Example

The following example evolves and evaluates ForeDreamer on the Companies category:

DATASET_TYPE=prophet_arena \
TRAIN_DATA_PATH="./data/Prophet-arena/subset_data_Companies_train_5.csv" \
VAL_DATA_PATH="./data/Prophet-arena/subset_data_Companies_train_5.csv" \
RUN_DIR="./runs/prophet-companies" \
bash scripts/evolve.sh

DATASET_TYPE=prophet_arena \
TRAIN_DATA_PATH="./data/Prophet-arena/subset_data_Companies_train_5.csv" \
INPUT_PATH="./data/Prophet-arena/subset_data_Companies_27.csv" \
RUN_DIR="./runs/prophet-companies" \
bash scripts/test.sh

Use the corresponding files from the data table to run another category.

๐Ÿ“ˆ Evaluation Outputs

By default, testing creates:

runs/EXPERIMENT_NAME/predictions.csv
runs/EXPERIMENT_NAME/test_summary.json

predictions.csv contains per-example model outputs and metrics. test_summary.json contains two evaluation scopes:

{
  "all_test": {
    "counts": {},
    "scores": {}
  },
  "train_overlap": {
    "key_field": "id or submission_id",
    "test_rows_in_train": 0,
    "test_rows_without_train": 0
  },
  "test_without_train": {
    "counts": {},
    "scores": {}
  }
}
  • FutureX overlap is determined with id and reports exact-match accuracy.
  • Prophet Arena overlap is determined with submission_id and reports mean Brier score and mean average return.
  • If the test input is already disjoint from the training set, all_test and test_without_train contain the same examples.

๐Ÿ› ๏ธ Configuration

Environment variableDefaultDescription
PYTHON_BINpython3.10Python executable
DATASET_TYPEfuturexfuturex or prophet_arena
RUN_DIRruns/$DATASET_TYPEIsolated evolution and test output directory
TRAIN_DATA_PATHdataset-specificEvolution data; during testing, IDs from this file are removed from the held-out summary
VAL_DATA_PATHTRAIN_DATA_PATHEvolution validation data
INPUT_PATHdataset-specificComplete test/evaluation input
OUTPUT_CSV$RUN_DIR/predictions.csvPer-example test output
TEST_SUMMARY_JSON$RUN_DIR/test_summary.jsonAggregate score summary
NUM_ITERATIONS3Number of evolution updates
PARALLELISM1Parallel evolution attempts
MEMGUIDE_ROUNDS_PER_CYCLE2Procedural updates per cycle
EXPERIENCE_ROUNDS_PER_CYCLE1Textual experience updates per cycle
EXPLORATION_OVER_EXPANSION1:1Exploration-to-rollout-expansion ratio
SEARCH_MAX_RESULTS1Search results per request
MAX_TURNS2Maximum main-agent turns
SUBAGENT_MAX_TURNS10Maximum memory-processing subagent turns
ENABLE_API_CACHE1Cache LLM and search requests under RUN_DIR/cache
USE_RAW_CONTEXT1Use Tavily raw_content when available
RESUME1Resume failed or missing test examples from an existing CSV
RUN_SPECIFICunsetComma-separated event tickers or FutureX sample IDs
MEM_GUIDEautomatically selectedOverride the selected MemGuide

Show entry-point help:

bash scripts/evolve.sh --help
bash scripts/test.sh --help

๐Ÿ“ Run Directory

All generated state is isolated under RUN_DIR:

runs/EXPERIMENT_NAME/
โ”œโ”€โ”€ best_assets.json
โ”œโ”€โ”€ predictions.csv
โ”œโ”€โ”€ test_summary.json
โ”œโ”€โ”€ ExperienceBank/
โ”œโ”€โ”€ FactualMemory/
โ”œโ”€โ”€ HistoryEvolution/
โ”œโ”€โ”€ HistoryRollout/
โ”œโ”€โ”€ MemGuide/
โ”œโ”€โ”€ MemTool/
โ””โ”€โ”€ cache/

The scripts do not modify the initial assets under src/MemGuide/ or src/MemTool/.

๐Ÿ–Š๏ธ Citation

If you find ForeDreamer useful in your research, please cite:

@article{zhong2026foredreamer,
  title={ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction},
  author={Zhong, Linhao and Du, Zongze and Wu, Linyu and Bo, Yu and Li, Hourong and Jing, Chenchen and Chen, Hao and Xi, Yuling and Shen, Chunhua},
  journal={arXiv preprint arXiv:2608.20920},
  year={2026}
}

๐Ÿ™ Acknowledgements

ForeDreamer is evaluated on Prophet Arena and FutureX. We thank the authors of these benchmarks and the open-source projects used in this repository.