LLM Provider Setup

July 4, 2026 · View on GitHub

tinyRAG is provider-agnostic: anything that speaks the OpenAI /v1/chat/completions and /v1/embeddings API works, local or cloud. This page collects setup notes for every provider preset in the web UI's provider switcher (#llmSwitcher in index.html), plus how to wire up something that isn't in the list.

Two independent endpoints can be configured — chat_base and embed_base (Settings → LLM Backend, or -url on first run). Most local runners serve both from the same base URL; for cloud chat-only providers (e.g. Anthropic, Groq) you typically keep a separate embedding backend (a local model, or OpenAI) since not every provider offers an embeddings endpoint.

Local runners

LM Studio

  • Download from lmstudio.ai.
  • Load a chat model and an embedding model (e.g. nomic-embed-text) in the "Local Server" tab, then start the server.
  • Default base URL: http://localhost:1234.
  • Fully OpenAI-compatible, including /v1/models for auto-discovery.

Ollama

  • Install: curl -fsSL https://ollama.ai/install.sh | sh (Linux/macOS) or the installer from ollama.ai (Windows/macOS).
  • Pull models: ollama pull llama3.1 and ollama pull nomic-embed-text.
  • Default base URL: http://localhost:11434.
  • Ollama's OpenAI-compatibility layer lives under /v1 — tinyRAG accounts for this automatically.

llama.cpp (llama-server)

  • Build or download llama-server from ggml-org/llama.cpp.
  • Run: llama-server -m model.gguf --embedding -m embed-model.gguf --port 8080 (embeddings require --embedding and a model that supports pooling).
  • Default base URL: http://localhost:8080.
  • Note: this port is shared with LocalAI and some llmster setups — the auto-detected provider hint is a best-effort guess; pick "llama.cpp" explicitly in the switcher if the guess is wrong.

vLLM

  • Install: pip install vllm.
  • Run: vllm serve <model> --port 8000 (vLLM exposes an OpenAI-compatible server by default).
  • Default base URL: http://localhost:8000.
  • Best for GPU-backed high-throughput serving; embeddings support depends on the model and vLLM version.

text-generation-webui (oobabooga)

  • Install per the project's instructions.
  • Enable the OpenAI-compatible extension (--extensions openai or via the UI).
  • Default base URL: http://localhost:5000.

KoboldCpp

  • Download from KoboldCpp releases.
  • Run with --port 5001; KoboldCpp exposes an OpenAI-compatible endpoint alongside its native API.
  • Default base URL: http://localhost:5001.

Jan

  • Download from jan.ai.
  • Enable the local API server in Jan's settings (Settings → Advanced → Local API Server).
  • Default base URL: http://localhost:1337.

LocalAI

  • Run via Docker: docker run -p 8080:8080 localai/localai.
  • Fully OpenAI-compatible, supports chat, embeddings, and more.
  • Default base URL: http://localhost:8080 (shares the port convention with llama.cpp — pick "LocalAI" explicitly in the switcher).

Demo mode: embedded GopherLLM (no external tool at all)

For trying tinyRAG with zero setup — no LM Studio, no Ollama, no llama.cpp, nothing to install or start separately — it can run GopherLLM, a pure-Go GGUF inference runtime, in-process. This is a demo/evaluation convenience, not a production inference path: no chat UI, no tool calling, no request concurrency tuning — whatever tiny model is loaded answers every request, on CPU, at whatever speed pure-Go inference gets you (noticeably slower than llama.cpp/LM Studio's hand-tuned C++ kernels, especially for anything bigger than a ~1B-parameter model).

Not part of the default build — it pulls in GopherLLM's full GGUF/tokenizer/ SIMD stack, which most deployments don't need. Build with:

go build -tags demo_llm ./...

Then either point at a specific .gguf file, or let it auto-discover one:

# Explicit model file
./tinyRAG -demo-llm-model /path/to/model.Q4_K_M.gguf

# Auto-discover: scans GopherLLM's default model directory
# ($RUSTY_LLM_MODEL_DIR, or LM Studio's community model cache under $HOME)
# for the first supported, non-projector GGUF file it finds.
./tinyRAG -demo-llm-model auto

# Change the local port the embedded server listens on (default 127.0.0.1:8091)
./tinyRAG -demo-llm-model auto -demo-llm-addr 127.0.0.1:9000

On startup tinyRAG loads the model, starts GopherLLM's OpenAI-compatible server as a background goroutine on -demo-llm-addr, waits for it to answer, then points chat_base/embed_base at it automatically — no manual Settings step needed. GopherLLM does not auto-download models; you need a .gguf file already on disk. Good demo-sized picks (small enough to run acceptably on CPU):

ModelApprox. size (Q4_K_M)Notes
Qwen2.5-0.5B-Instruct~400 MBGood instruction-following for its size
SmolLM2-360M-Instruct~250 MBSmaller/faster, weaker reasoning
SmolLM2-135M-Instruct~100 MBFastest, best for quick UI demos

Supported architectures: llama, llama2, llama3, mistral, mistral3, qwen2, gemma, gemma2, gpt-oss. See demo_llm.go (startEmbeddedDemoLLM/resolveDemoLLMModel) for the integration; the important gotcha: go mod tidy without -tags=demo_llm will drop the GopherLLM requirement from go.mod since it can't see the build-tag-gated import — run go mod tidy -tags=demo_llm instead when tidying this module.

Cloud providers

ProviderBase URLNotes
OpenAIhttps://api.openai.comSet the API key in Settings or OPENAI_API_KEY env var.
Anthropichttps://api.anthropic.comChat only — pair with a local or OpenAI embedding backend.
Google Geminihttps://generativelanguage.googleapis.comUses the OpenAI-compatibility endpoint Google provides under this host.
Mistral AIhttps://api.mistral.aiOffers both chat and embedding models.
Groqhttps://api.groq.com/openaiChat only, very low latency.
DeepSeekhttps://api.deepseek.comChat only.
Together AIhttps://api.together.xyzHosts many open-weight models; check model-specific embedding support.
xAI (Grok)https://api.x.aiChat only.
Coherehttps://api.cohere.aiOffers a dedicated embeddings API alongside chat.
Perplexityhttps://api.perplexity.aiChat only, includes web-grounded models.
OpenRouterhttps://openrouter.ai/apiRoutes to many upstream models through one key.

All cloud providers need an API key — set it in Settings → LLM Backend ("OpenAI API Key" field applies to whichever base URL is active) or via the OPENAI_API_KEY environment variable, which is used as a fallback whenever no key is stored in settings.json.

Using a provider that isn't listed

Pick "Custom..." in the provider switcher (or just fill in Settings → LLM Backend directly): any server implementing /v1/chat/completions and /v1/embeddings works, including self-hosted proxies, Azure OpenAI deployments, or a provider added after this document was written. Enter the base URL (without a trailing /v1 — tinyRAG normalizes that), click "Test & Load Models" to populate the model dropdowns via /v1/models, then save.

How auto-detection works

  • providerHintFromURL (llm_discovery.go) maps a base URL to a human-readable provider name, first by hostname (cloud providers) and then by port (local runners) — used to pre-select the provider switcher and label the workspace pill.
  • localLLMCandidates lists the local ports above; maybePreferOfflineLLM probes them in order on startup if the configured endpoint doesn't respond, and switches to the first reachable one automatically.
  • POST /api/llm/list-models probes a given base URL live from the web UI (used by the provider switcher and Settings' "Test & Load Models" button).