InferrLM REST API Documentation
April 1, 2026 · View on GitHub
Complete API reference for InferrLM's local HTTP server that exposes AI inference capabilities over your local network.
Getting Started
Quick Start
- Start the server — Open InferrLM, go to the Server tab, and toggle it on. Your URL will appear (e.g.
http://192.168.1.10:8889). - Download a model — Make sure at least one GGUF model is downloaded in the Models tab. The model name (without
.gguf) is used in API requests. - Configure your client — Point any OpenAI-compatible client to
http://YOUR_DEVICE_IP:8889/v1. No API key is required — use any placeholder if the client requires one. - Send a request:
curl -X POST http://YOUR_DEVICE_IP:8889/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "llama-3.2-1b", "messages": [{"role": "user", "content": "Hello!"}]}'
This works with any application or library that supports the OpenAI API — just point it to
http://YOUR_DEVICE_IP:8889/v1. Both devices must be on the same local network. The.ggufextension is optional in the model name.
Starting the Server
- Open the InferrLM app on your device
- Navigate to the Server tab
- Toggle the server switch to start it
- Your server URL will be displayed (typically
http://YOUR_DEVICE_IP:8889) - You can share this URL via QR code or copy it to access from other devices
Configuration Options
- Auto-start: Automatically start the server when the app launches
- Port: Default port is 8889 (configurable in settings)
Base Configuration
Base URL: http://YOUR_DEVICE_IP:8889
Content-Type: application/json
CORS: Enabled for all origins
Selecting a Model Target
Every request that generates text includes a model string that determines which execution backend handles the workload:
| Model value | Routed backend | Notes |
|---|---|---|
Stored model name (e.g. llama-3.2-1b) | Local GGUF running on-device | Download the GGUF via the InferrLM app first. |
apple-foundation | Apple Intelligence Foundation model | iOS only. Enable in app settings and verify via GET /api/models/apple-foundation. |
Chat & Completion APIs
POST /api/chat
Stream or complete a chat with full conversation history. Accepts local GGUF model names or apple-foundation.
Request Body:
{
"model": "llama-3.2-1b",
"messages": [
{"role": "system", "content": "You are a helpful assistant"},
{"role": "user", "content": "Hello!"}
],
"stream": true,
"temperature": 0.7,
"max_tokens": 512
}
Streaming Response (NDJSON):
{"model":"llama-3.2-1b.gguf","created_at":"...","message":{"role":"assistant","content":"Hi"},"done":false}
{"model":"llama-3.2-1b.gguf","created_at":"...","message":{"role":"assistant","content":" there"},"done":false}
{"model":"llama-3.2-1b.gguf","created_at":"...","message":{"role":"assistant","content":""},"done":true}
Non-streaming Response:
{
"model": "llama-3.2-1b.gguf",
"created_at": "2026-03-20T10:00:00.000Z",
"message": {"role": "assistant", "content": "Hi there!"},
"done": true
}
Parameters:
model(string, required): Target backendmessages(array, required): Conversation history — each entry hasrole(system|user|assistant) andcontentstream(boolean, optional): Enable streaming NDJSON responses (default:true)temperature(number, optional): Sampling temperature 0.0–2.0max_tokens(number, optional): Maximum tokens to generatetop_p(number, optional): Top-p nucleus samplingtop_k(number, optional): Top-k sampling
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/chat \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.2-1b",
"messages": [{"role": "user", "content": "Explain AI"}],
"stream": false
}'
POST /api/generate
Generate a completion from a single prompt (no conversation context).
Request Body:
{
"model": "llama-3.2-1b",
"prompt": "Explain quantum computing in simple terms",
"stream": false,
"max_tokens": 500
}
Response:
{
"model": "llama-3.2-1b.gguf",
"created_at": "2026-03-20T10:00:00.000Z",
"response": "Quantum computing uses quantum mechanics principles...",
"done": true
}
Parameters:
model(string, required): Target backendprompt(string, required): Input promptstream(boolean, optional): Enable streaming NDJSON responsesmax_tokens(number, optional): Maximum tokens to generatetemperature(number, optional): Sampling temperature
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/generate \
-H "Content-Type: application/json" \
-d '{"model": "llama-3.2-1b", "prompt": "Hello world", "stream": false}'
OpenAI-Compatible API
GET /v1/models
List available models in OpenAI format. Compatible with any OpenAI client library.
Response:
{
"object": "list",
"data": [
{
"id": "llama-3.2-1b.gguf",
"object": "model",
"created": 1700000000,
"owned_by": "local"
}
]
}
Example:
curl http://YOUR_DEVICE_IP:8889/v1/models
POST /v1/chat/completions
OpenAI-compatible chat completions endpoint. Drop-in replacement for apps built against the OpenAI API. No API key is required — set any non-empty placeholder in the Authorization header if your client requires one.
Request Body:
{
"model": "llama-3.2-1b",
"messages": [
{"role": "user", "content": "Hello!"}
],
"stream": false,
"max_tokens": 100
}
Non-streaming Response:
{
"id": "chatcmpl-...",
"object": "chat.completion",
"created": 1700000000,
"model": "llama-3.2-1b.gguf",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "Hi there!"},
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0}
}
Streaming Response (SSE):
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"llama-3.2-1b.gguf","choices":[{"index":0,"delta":{"content":"Hi"},"finish_reason":null}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":...,"model":"llama-3.2-1b.gguf","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "llama-3.2-1b", "messages": [{"role": "user", "content": "Hi"}], "stream": false}'
Chat History
GET /api/chats
List all saved chat conversations (without messages).
Response:
{
"chats": [
{
"id": "chat-abc123",
"title": "Quantum Physics Discussion",
"timestamp": 1700000000000,
"modelPath": "/path/to/model.gguf",
"messageCount": 12
}
]
}
Example:
curl http://YOUR_DEVICE_IP:8889/api/chats
POST /api/chats
Create a new chat conversation.
Request Body:
{
"title": "My Conversation",
"messages": [
{"role": "user", "content": "Hello"},
{"role": "assistant", "content": "Hi there!"}
]
}
Response (201):
{
"chat": {
"id": "chat-abc123",
"title": "My Conversation",
"timestamp": 1700000000000,
"modelPath": null,
"messageCount": 2,
"messages": [...]
}
}
Parameters:
title(string, optional): Chat titlemessages(array, optional): Initial messages to seed the conversation
GET /api/chats/:id
Get a specific chat including all messages.
Response:
{
"chat": {
"id": "chat-abc123",
"title": "My Conversation",
"timestamp": 1700000000000,
"modelPath": null,
"messageCount": 4,
"messages": [...]
}
}
Example:
curl http://YOUR_DEVICE_IP:8889/api/chats/chat-abc123
DELETE /api/chats/:id
Delete a chat conversation.
Response:
{
"status": "deleted",
"chatId": "chat-abc123"
}
Example:
curl -X DELETE http://YOUR_DEVICE_IP:8889/api/chats/chat-abc123
GET /api/chats/:id/messages
Get only the messages for a specific chat.
Response:
{
"messages": [
{"id": "msg-1", "role": "user", "content": "Hello"},
{"id": "msg-2", "role": "assistant", "content": "Hi there!"}
]
}
Example:
curl http://YOUR_DEVICE_IP:8889/api/chats/chat-abc123/messages
POST /api/chats/:id/messages
Append one or more messages to an existing chat.
Request Body:
{
"messages": [
{"role": "user", "content": "Follow-up question"}
]
}
Response (201):
{
"messages": [
{"id": "msg-3", "role": "user", "content": "Follow-up question"}
]
}
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/chats/chat-abc123/messages \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Follow-up"}]}'
Model Management
GET /api/tags
List all models stored on the device.
Response:
{
"models": [
{
"name": "llama-3.2-1b.gguf",
"modified_at": "2026-03-20T10:00:00.000Z",
"size": 1234567890,
"digest": null,
"model_type": "llama",
"is_external": false
}
]
}
Example:
curl http://YOUR_DEVICE_IP:8889/api/tags
GET /api/ps
List currently loaded models (models in memory).
Response:
{
"models": [
{
"name": "llama-3.2-1b.gguf",
"model": "/path/to/model.gguf",
"size": 1234567890,
"loaded_at": "2026-03-20T10:00:00.000Z",
"is_external": false,
"model_type": "llama"
}
]
}
Returns an empty models array when no model is loaded.
Example:
curl http://YOUR_DEVICE_IP:8889/api/ps
POST /api/show
Get detailed information about a specific model including GGUF metadata and current settings.
Request Body (use name, model, or path):
{
"model": "llama-3.2-1b"
}
Response:
{
"name": "llama-3.2-1b.gguf",
"path": "/path/to/model.gguf",
"size": 1234567890,
"modified_at": "2026-03-20T10:00:00.000Z",
"is_external": false,
"model_type": "llama",
"capabilities": ["completion"],
"multimodal": false,
"default_projection_model": null,
"settings": {
"temperature": 0.7,
"topP": 0.9,
"maxTokens": 2048
},
"info": {
"general.architecture": "llama",
"general.parameter_count": 1000000000
}
}
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/show \
-H "Content-Type: application/json" \
-d '{"model": "llama-3.2-1b"}'
POST /api/pull
Download a model from a URL directly to the device.
Request Body:
{
"url": "https://huggingface.co/model.gguf",
"model": "my-custom-model"
}
Response:
{
"status": "downloading",
"model": "my-custom-model",
"downloadId": "download-abc123"
}
The download runs in the background. Use GET /api/tags to check when the model appears.
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/pull \
-H "Content-Type: application/json" \
-d '{"url": "https://huggingface.co/model.gguf", "model": "my-model"}'
POST /api/copy
Copy an existing model file under a new name.
Request Body:
{
"source": "llama-3.2-1b",
"destination": "llama-3.2-1b-backup"
}
Response:
{
"status": "copied",
"source": "llama-3.2-1b.gguf",
"destination": "llama-3.2-1b-backup.gguf"
}
Returns 409 if the destination name already exists. External models cannot be copied.
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/copy \
-H "Content-Type: application/json" \
-d '{"source": "llama-3.2-1b", "destination": "llama-backup"}'
DELETE /api/delete
Delete a model from local storage.
Request Body (use name or path):
{
"name": "llama-3.2-1b"
}
Response:
{
"success": true
}
Example:
curl -X DELETE http://YOUR_DEVICE_IP:8889/api/delete \
-H "Content-Type: application/json" \
-d '{"name": "old-model"}'
POST /api/models
Perform model lifecycle operations.
Request Body:
{
"action": "load",
"model": "llama-3.2-1b"
}
Available Actions:
| Action | Description | model field |
|---|---|---|
load | Load a model into memory | Required — model name or path |
unload | Release the currently loaded model | Not used |
reload | Reinitialise the currently loaded model | Not used |
refresh | Rescan storage and reload the model list | Not used |
Response (load):
{
"status": "loaded",
"model": {
"name": "llama-3.2-1b.gguf",
"path": "/path/to/model.gguf",
"projector": null
}
}
Response (refresh):
{
"status": "refreshed",
"count": 3,
"models": [...]
}
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/models \
-H "Content-Type: application/json" \
-d '{"action": "load", "model": "llama-3.2-1b"}'
GET /api/models/apple-foundation
Check Apple Foundation model availability and readiness (iOS only).
Response:
{
"available": true,
"requirementsMet": true,
"enabled": true,
"status": "ready",
"message": "Apple Foundation is ready to use."
}
status | Meaning |
|---|---|
ready | Available and enabled — use model: "apple-foundation" in requests |
configure | Not available or not enabled — see message for details |
Example:
curl http://YOUR_DEVICE_IP:8889/api/models/apple-foundation
POST /api/models/apple-foundation
Verify that Apple Foundation is ready to process requests. Returns an error if not available, requirements are not met, or the feature is not enabled in app settings.
Response (ready):
{
"status": "ready"
}
Error responses:
501—apple_foundation_unavailable: device does not support Apple Intelligence428—requirements_not_met: device needs to be updated409—apple_foundation_disabled: enable it in app settings first
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/models/apple-foundation
GET /api/version
Get the current app version.
Response:
{
"version": "0.8.3"
}
Example:
curl http://YOUR_DEVICE_IP:8889/api/version
RAG & Embeddings
POST /api/embeddings
Generate embeddings for one or more texts using a local model.
Request Body:
{
"model": "llama-3.2-1b",
"input": "The quick brown fox jumps over the lazy dog"
}
Pass an array to embed multiple texts in one request:
{
"model": "llama-3.2-1b",
"input": ["First text", "Second text"]
}
Response:
{
"embeddings": [
[0.123, -0.456, 0.789, "..."]
],
"model": "llama-3.2-1b.gguf"
}
Parameters:
model(string, required): Local model to use for embeddinginput(string or array, required): Text(s) to embed. Also accepted aspromptortext.
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/embeddings \
-H "Content-Type: application/json" \
-d '{"model": "llama-3.2-1b", "input": "Sample text"}'
POST /api/files/ingest
Ingest content into the RAG system. Accepts raw text, a file path on the device, or multiple file paths.
Request Body (raw text):
{
"content": "Document content to store for RAG...",
"fileName": "my-doc.txt"
}
Request Body (single file path):
{
"filePath": "/path/to/doc.txt"
}
Request Body (multiple file paths):
{
"files": ["/path/to/doc1.txt", "/path/to/doc2.txt"]
}
Response:
{
"status": "stored",
"documentId": "1700000000000-abc123",
"fileName": "my-doc.txt",
"model": null
}
Parameters:
content(string): Raw text content (required iffilePathandfilesare omitted)filePath(string, optional): Absolute path to a file on the devicefiles(array, optional): Array of absolute file pathsfileName(string, optional): Display name for the document (default:"uploaded.txt")chatId(string, optional): Associate the document with a specific chatprovider(string, optional): RAG embedding providerrag(boolean, optional): Set tofalseto skip RAG indexing (default:true)
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/files/ingest \
-H "Content-Type: application/json" \
-d '{"content": "Machine learning is a subset of AI...", "fileName": "ml-intro.txt"}'
GET /api/rag
Get the current RAG system status.
Response:
{
"enabled": true,
"ready": true,
"storage": "persistent",
"documentCount": 3
}
Example:
curl http://YOUR_DEVICE_IP:8889/api/rag
POST /api/rag
Configure the RAG system (enable/disable, set storage type, or initialise).
Request Body:
{
"enabled": true,
"storage": "persistent",
"initialize": true
}
Response:
{
"enabled": true,
"ready": true,
"storage": "persistent",
"documentCount": 0
}
Parameters:
enabled(boolean, optional): Enable or disable RAGstorage(string, optional):"memory"or"persistent"initialize(boolean, optional): Trigger RAG initialisationprovider(string, optional): Embedding provider to use on initialisation
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/rag \
-H "Content-Type: application/json" \
-d '{"enabled": true, "storage": "persistent"}'
POST /api/rag/reset
Clear all ingested documents from the RAG system.
Response:
{
"status": "cleared",
"enabled": true,
"ready": false,
"documentCount": 0
}
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/rag/reset
Server & Settings
GET /api/status
Get server status, active model, and RAG state.
Response:
{
"server": {
"isRunning": true,
"url": "http://192.168.1.110:8889",
"port": 8889,
"clientCount": 1
},
"model": {
"loaded": true,
"path": "/path/to/model.gguf"
},
"rag": {
"ready": false
}
}
Example:
curl http://YOUR_DEVICE_IP:8889/api/status
POST /api/settings/thinking
Enable or disable thinking mode (extended reasoning) for the currently loaded model.
Request Body:
{
"enabled": true
}
Response:
{
"status": "updated",
"enabled": true
}
Parameters:
enabled(boolean, required): Enable or disable thinking mode
Example:
curl -X POST http://YOUR_DEVICE_IP:8889/api/settings/thinking \
-H "Content-Type: application/json" \
-d '{"enabled": true}'
Error Handling
All endpoints return standard HTTP status codes with a JSON error body.
Success codes:
200 OK201 Created
Error codes:
400 Bad Request— missing or invalid parameters404 Not Found— resource does not exist405 Method Not Allowed409 Conflict— precondition not met (e.g. remote models disabled)422 Unprocessable Entity— valid request but action cannot be performed (e.g. API key missing)500 Internal Server Error503 Service Unavailable— model not loaded
Error response format:
{
"error": "error_code"
}
Security Considerations
- The server is designed for local network use only
- No authentication is required (secured by network isolation)
- CORS is enabled for all origins
- Consider a VPN or firewall before exposing the server beyond your local network
Rate Limiting
No rate limits are enforced. Performance depends on device CPU/RAM, model size, and number of concurrent connections.
Common Use Cases
Chat with a local model
curl -X POST http://YOUR_DEVICE_IP:8889/api/chat \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.2-1b",
"messages": [
{"role": "system", "content": "You are a helpful coding assistant"},
{"role": "user", "content": "Write a Python function to calculate fibonacci"}
],
"stream": false
}'
Ingest a document and check RAG status
curl -X POST http://YOUR_DEVICE_IP:8889/api/files/ingest \
-H "Content-Type: application/json" \
-d '{"content": "Your document content here...", "fileName": "doc.txt"}'
curl http://YOUR_DEVICE_IP:8889/api/rag
Model management
# List available models
curl http://YOUR_DEVICE_IP:8889/api/tags
# Load a model
curl -X POST http://YOUR_DEVICE_IP:8889/api/models \
-H "Content-Type: application/json" \
-d '{"action": "load", "model": "llama-3.2-1b"}'
# Check what is loaded
curl http://YOUR_DEVICE_IP:8889/api/ps
Example Applications
InferrLM CLI
The InferrLM CLI is a command-line interface tool built with React, Ink, and TypeScript. It connects to your InferrLM server and provides a fully functional terminal-based chat interface with streaming support and conversation history.
Source code: github.com/sbhjt-gr/inferra-cli
Additional Resources
Last Updated: March 20, 2026
API Version: 0.8.3