LLM Council Plus
July 1, 2026 · View on GitHub

Inspired by Andrej Karpathy's LLM Council - see his original tweet about the concept.
The idea of this repo is that instead of asking a question to your favorite LLM provider (e.g. OpenAI GPT-5.5, Google Gemini 3.1 Pro, Anthropic Claude Sonnet 5, xAI Grok 4.3, etc.), you can group them into your "LLM Council Plus" council. This is a containerized web app with a Setup Wizard that guides you through configuration. It uses OpenRouter to send your query to multiple LLMs, asks them to review and rank each other's work, and finally a Chairman LLM produces the final response.
In a bit more detail, here is what happens when you submit a query:
- Stage 1: First opinions. The user query is given to all LLMs individually, and the responses are collected. The individual responses are shown in a "tab view", so that the user can inspect them all one by one.
- Stage 2: Review. Each individual LLM is given the responses of the other LLMs. Under the hood, the LLM identities are anonymized so that the LLM can't play favorites when judging their outputs. The LLM is asked to rank them in accuracy and insight.
- Stage 3: Final response. The designated Chairman takes all of the model's responses and compiles them into a single final answer that is presented to the user.
What's Different from the Original
This fork extends Andrej Karpathy's llm-council with production-ready features:
| Feature | Original | LLM Council Plus |
|---|---|---|
| Deployment | Manual Python/npm setup | Docker Compose (one command) |
| Configuration | Edit config.py manually | Visual Setup Wizard |
| Models | 4 hardcoded models | Full OpenRouter catalog (100+ models) |
| Local Models | ❌ | ✅ Ollama support |
| Authentication | ❌ | ✅ JWT multi-user system |
| Storage | JSON files only | JSON, PostgreSQL, MySQL |
| Token Optimization | ❌ | ✅ TOON format (20-60% savings) |
| Web Search | ❌ | ✅ DuckDuckGo + Tavily + Exa + Brave |
| File Attachments | ❌ | ✅ PDF, TXT, MD, MDX, images (drag-and-drop) |
| System Prompts | ❌ | ✅ Per-conversation custom system prompts |
| Tools | ❌ | ✅ Calculator, Wikipedia, ArXiv, Yahoo Finance |
| Real-time Streaming | Basic | SSE with heartbeats & state persistence |
| Error Handling | Silent failures | Toast notifications + visual indicators |
| Hot Reload | ❌ | ✅ Config changes without restart |
| Conversation Search | ❌ | ✅ Filter by title with relative dates |
| Stage Timeouts | ❌ | ✅ 90s stage-level timeout (first-N-complete) |
Recent Updates (v1.4.0)
| New Feature | What it adds | Where |
|---|---|---|
| Custom System Prompts | Per-conversation custom system prompts | Model Selector |
| Drag-and-Drop Upload | Drop files directly onto the chat area with visual overlay feedback | Chat Interface |
| MDX File Support | Upload .mdx files alongside PDF, TXT, MD, and images | File Upload |
Recent Updates (v1.3.1)
| New Feature | What it adds | Where |
|---|---|---|
| Per-conversation router | Choose OpenRouter or Ollama per conversation (no fallback) | Model Selector → Router |
| Model presets + council sizing | Built-in presets + saved presets, configurable council size, “I’m Feeling Lucky” | Model Selector |
| Runtime Settings panel | Edit prompts + temperatures; persist / reset / export / import (no secrets) | Settings → Prompts/Temps/Backup |
| Stop / Abort streaming | Cancel an in-flight request cleanly (frontend abort + backend task cancel) | Composer “Stop” |
| Web search provider selection | DuckDuckGo (free) + Tavily + Exa + Brave, per-message selection | Composer search pill |
| Full-article fetch (optional) | Fetch top-N pages via Jina Reader for richer context (DDG/Brave) | Settings → Web Search |
| Search context panel | Compact view (title + domain) + expandable content per result | Assistant message UI |
| UI polish pass | Dark dropdown menus, cleaner sidebar actions, compact composer layout | App UI |
Recent Updates (v1.3.0)
- Conversation Search — Filter conversations by title in real-time
- Relative Dates — Sidebar shows "Today 14:30", "Yesterday", "Mon", "Jan 2" etc.
- Stage-level Timeouts — Stage 2/3 proceed after 90s with partial results (first-N-complete pattern)
- Toast Notifications — Visual alerts for model timeouts and errors
- Improved Error Handling — Graceful fallback when models fail or timeout
Quick Start
Easiest way — use the Setup Wizard:
# 1. Create your .env file (REQUIRED — backend will crash without it)
cp .env.example .env
# 2. Create the data directory for conversation storage
mkdir -p data/conversations
# 3. Build and start (first time — full rebuild recommended)
docker compose build --no-cache && docker compose up -d
# 4. Open http://localhost
# The Setup Wizard will guide you through configuration
Important: Steps 1 and 2 are required before the first run. Without the
.envfile anddata/directory, the backend container will crash-loop.
The Setup Wizard lets you:
- Choose LLM provider (OpenRouter or Ollama)
- Enter API keys
- Optionally enable authentication
- Optionally configure web search (Tavily or Exa)
Alternative: Manual configuration
cp .env.example .env
mkdir -p data/conversations
# Edit .env with your API keys (see examples below)
docker compose up --build -d
# Open http://localhost
Note: Authentication is disabled by default for easy setup. For production deployment with user authentication, see SECURITY.md.
Vibe Code Alert
This project was 99% vibe coded as a fun Saturday hack because I wanted to explore and evaluate a number of LLMs side by side in the process of reading books together with LLMs. It's nice and useful to see multiple responses side by side, and also the cross-opinions of all LLMs on each other's outputs. I'm not going to support it in any way, it's provided here as is for other people's inspiration and I don't intend to improve it. Code is ephemeral now and libraries are over, ask your LLM to change it in whatever way you like.
Detailed Setup
1. Install Dependencies
The project uses uv for project management.
Backend:
uv sync
Frontend:
cd frontend
npm install
cd ..
2. Configure API Key
Create a .env file in the project root:
OPENROUTER_API_KEY=sk-or-v1-...
Get your API key at openrouter.ai. Make sure to purchase the credits you need, or sign up for automatic top up.
Alternatively, you can use Ollama for local models by setting ROUTER_TYPE=ollama in your .env file (see below).
Models are selected via the Setup Wizard or dynamically in the UI. Browse all available models at openrouter.ai/models.
Running the Application
Docker (Recommended)
Prerequisite: Docker Desktop must be running on your machine before executing any
docker composecommands.
# First-time setup or after dependency changes — full rebuild (no cache)
docker compose build --no-cache && docker compose up -d
# Regular start (uses cached layers, faster)
docker compose up --build -d
# Access the application at http://localhost
Backend API is available at http://localhost:8001
Stopping and managing containers:
# Stop and remove all containers
docker compose down
# View running containers
docker compose ps
# View backend logs (useful for debugging)
docker compose logs backend
# Follow logs in real-time
docker compose logs -f backend
Troubleshooting
Backend crash-looping (Restarting (1) in docker compose ps):
# Check the error logs first
docker compose logs backend --tail 50
# Most common cause: missing .env file or data directory
cp .env.example .env # if .env doesn't exist
mkdir -p data/conversations # if data/ doesn't exist
docker compose down
docker compose up -d
ModuleNotFoundError or import errors:
Docker's build cache is likely stale. Force a full rebuild:
docker compose down
docker compose build --no-cache
docker compose up -d
Development Mode (without Docker)
Terminal 1 (Backend):
uv run python -m backend.main
Terminal 2 (Frontend):
cd frontend
npm run dev
Then open http://localhost:5173 in your browser.
Common Configuration
OpenRouter (cloud)
ROUTER_TYPE=openrouter
OPENROUTER_API_KEY=sk-or-v1-...
Ollama (local)
ROUTER_TYPE=ollama
If you're running the backend in Docker and Ollama is running on your host machine, set:
OLLAMA_HOST=host.docker.internal:11434
Web Search (Optional)
The UI can run web search and inject results into Stage 1 as context. You can select the provider per message (search pill) or set a default in Settings.
Providers:
- DuckDuckGo (free): enabled by default, no API key required
- Tavily / Exa / Brave: require API keys (optional)
Enable any of the optional providers:
ENABLE_TAVILY=true
TAVILY_API_KEY=tvly-...
or:
ENABLE_EXA=true
EXA_API_KEY=...
or:
ENABLE_BRAVE=true
BRAVE_API_KEY=...
Storage
Default storage is local JSON files under data/conversations/ (mounted into the backend container).
Database storage is supported via DATABASE_TYPE=postgresql or DATABASE_TYPE=mysql (see backend/storage.py).
Tech Stack
- Backend: FastAPI (Python 3.10+), async httpx, OpenRouter API or Ollama
- Frontend: React + Vite, react-markdown for rendering
- Storage: JSON files in
data/conversations/(optional DB storage) - Package Management: uv for Python, npm for JavaScript
Contributing
See CONTRIBUTING.md for development setup and guidelines.
Security
See SECURITY.md for security considerations and reporting vulnerabilities.
License
MIT License - see LICENSE for details.