Lightspeed Core Stack

August 3, 2026 ยท View on GitHub


๐Ÿ“‹ Configuration schema

A2AStateConfiguration

A2A protocol persistent state configuration.

Configures how A2A task state and context-to-conversation mappings are stored. For multi-worker deployments, use SQLite or PostgreSQL to ensure state is shared across all workers.

If no configuration is provided, in-memory storage is used (default). This is suitable for single-worker deployments but state will be lost on restarts and not shared across workers.

Attributes: sqlite: SQLite database configuration for A2A state storage. postgres: PostgreSQL database configuration for A2A state storage.

FieldTypeDescription
sqliteSQLite database configuration for A2A state storage.
postgresPostgreSQL database configuration for A2A state storage.

APIKeyTokenConfiguration

API Key Token configuration.

FieldTypeDescription
api_keystring

AccessRule

Rule defining what actions a role can perform.

FieldTypeDescription
rolestringName of the role
actionsarrayAllowed actions for this role

Action

Available actions in the system.

Note: this is not a real model, just an enumeration of all action names.

ApprovalFilter

Granular approval control for specific MCP tools.

Attributes: always: Tool names that always require human approval before execution. never: Tool names that never require approval (pre-approved).

FieldTypeDescription
alwaysarrayList of tool names that always require human approval
neverarrayList of tool names that never require approval

ApprovalsConfiguration

Configuration for human-in-the-loop approvals.

Attributes: approval_timeout_seconds: How long approval requests remain pending before expiring. approval_retention_days: How long to retain decided approvals for audit purposes before cleanup.

FieldTypeDescription
approval_timeout_secondsintegerSeconds before pending approval requests expire
approval_retention_daysintegerDays to retain decided approvals before cleanup

AuthenticationConfiguration

Authentication configuration.

FieldTypeDescription
modulestring
skip_tls_verificationboolean
skip_for_health_probesbooleanSkip authorization for readiness and liveness probes
skip_for_metricsbooleanSkip authorization for the /metrics endpoint
k8s_cluster_apistring
k8s_ca_cert_pathstring
jwk_config
api_key_config
rh_identity_config
trusted_proxy_config

AuthorizationConfiguration

Authorization configuration.

FieldTypeDescription
access_rulesarrayRules for role-based access control

AzureEntraIdConfiguration

Microsoft Entra ID authentication attributes for Azure.

FieldTypeDescription
tenant_idstring
client_idstring
client_secretstring
scopestringAzure Cognitive Services scope for token requests. Override only if using a different Azure service.

ByokRag

BYOK (Bring Your Own Knowledge) RAG configuration.

FieldTypeDescription
rag_idstringUnique RAG ID
rag_typestringType of RAG database (e.g. 'inline::faiss', 'remote::pgvector').
embedding_modelstringEmbedding model identification
embedding_dimensionintegerDimensionality of embedding vectors.
vector_db_idstringVector database identification.
db_pathstringPath to RAG database. Required for inline::faiss.
score_multipliernumberMultiplier applied to relevance scores from this vector store. Used to weight results when querying multiple knowledge sources. Values > 1 boost this store's results; values <; 1 reduce them.
hoststringPostgreSQL host for remote::pgvector. Defaults to ${env.POSTGRES_HOST} when rag_type is remote::pgvector.
portstringPostgreSQL port for remote::pgvector. Defaults to ${env.POSTGRES_PORT} when rag_type is remote::pgvector.
dbstringPostgreSQL database name for remote::pgvector. Defaults to ${env.POSTGRES_DATABASE} when rag_type is remote::pgvector.
userstringPostgreSQL user for remote::pgvector. Defaults to ${env.POSTGRES_USER} when rag_type is remote::pgvector.
passwordstringPostgreSQL password for remote::pgvector. Defaults to ${env.POSTGRES_PASSWORD} when rag_type is remote::pgvector.

CORSConfiguration

CORS configuration.

CORS or 'Cross-Origin Resource Sharing' refers to the situations when a frontend running in a browser has JavaScript code that communicates with a backend, and the backend is in a different 'origin' than the frontend.

Useful resources:

FieldTypeDescription
allow_originsarrayA list of origins allowed for cross-origin requests. An origin is the combination of protocol (http, https), domain (myapp.com, localhost, localhost.tiangolo.com), and port (80, 443, 8080). Use ['*'] to allow all origins.
allow_credentialsbooleanIndicate that cookies should be supported for cross-origin requests
allow_methodsarrayA list of HTTP methods that should be allowed for cross-origin requests. You can use ['*'] to allow all standard methods.
allow_headersarrayA list of HTTP request headers that should be supported for cross-origin requests. You can use ['*'] to allow all headers. The Accept, Accept-Language, Content-Language and Content-Type headers are always allowed for simple CORS requests.

CompactionConfiguration

Configuration for conversation history compaction.

Compaction summarizes older conversation turns when their estimated token count approaches the context window limit, keeping the conversation usable instead of failing with HTTP 413. The configuration here controls when compaction triggers and how much recent context is preserved verbatim.

Attributes: enabled: Master switch. When False, compaction never triggers and other fields are inert. threshold_ratio: Trigger compaction when estimated input tokens exceed this fraction of the model's context window (clamped to 0.0..1.0). token_floor: Minimum estimated token count before compaction can trigger, regardless of threshold_ratio. Prevents triggering on very small context windows. buffer_turns: Initial number of recent turns to keep verbatim. The runtime applies a degrading guard โ€” if these turns exceed the available budget, it reduces buffer_turns by one repeatedly until the budget fits, down to zero. buffer_max_ratio: Hard cap on the fraction of the context window the buffer zone may occupy, regardless of buffer_turns.

FieldTypeDescription
enabledbooleanWhen true, older conversation turns are summarized when estimated tokens approach the context window limit.
threshold_rationumberTrigger compaction when estimated tokens exceed this fraction of the model's context window (0.0-1.0).
token_floorintegerMinimum token count before compaction can trigger. Prevents triggering on very small context windows.
buffer_turnsintegerNumber of recent turns to keep verbatim.
buffer_max_rationumberMaximum fraction of context window the buffer zone can occupy, regardless of buffer_turns.

Configuration

Global service configuration.

FieldTypeDescription
namestringName of the service. That value will be used in REST API endpoints.
serviceThis section contains Lightspeed Core Stack service configuration.
llama_stackThis section contains Llama Stack configuration. Lightspeed Core Stack service can call Llama Stack in library mode or in server mode.
user_data_collectionThis section contains configuration for subsystem that collects user data(transcription history and feedbacks).
databaseConfiguration for database to store conversation IDs and other runtime data
mcp_serversarrayMCP (Model Context Protocol) servers provide tools and capabilities to the AI agents. These are configured in this section. Only MCP servers defined in the lightspeed-stack.yaml configuration are available to the agents. Tools configured in the llama-stack run.yaml are not accessible to lightspeed-core agents.
authenticationAuthentication configuration
authorizationLightspeed Core Stack implements a modular authentication and authorization system with multiple authentication methods. Authorization is configurable through role-based access control. Authentication is handled through selectable modules configured via the module field in the authentication configuration.
customizationIt is possible to customize Lightspeed Core Stack via this section. System prompt can be customized and also different parts of the service can be replaced by custom Python modules.
inferenceOne LLM provider and one its model might be selected as default ones. When no provider+model pair is specified in REST API calls (query endpoints), the default provider and model are used.
conversation_cache
compactionControls when conversation history is summarized to keep the model's input below the context window limit. Disabled by default โ€” when disabled, requests that exceed the window continue to surface as HTTP 413.
approvalsSettings for human-in-the-loop approval of MCP tool invocations
byok_ragarrayBYOK RAG configuration. This configuration can be used to reconfigure Llama Stack through its run.yaml configuration file
vector_storeDynamic vector-store provider capacity for runtime POST /v1/vector-stores creates. Not the same as byok_rag (static registered corpora). When providers is non-empty, default_provider is required and must match one of providers[].id. Applied in unified synthesis only.
a2a_stateConfiguration for A2A protocol persistent state storage.
quota_handlersQuota handlers configuration
azure_entra_id
rlsapi_v1Configuration for the rlsapi v1 /infer endpoint used by the RHEL Lightspeed Command Line Assistant (CLA).
splunkSplunk HEC configuration for sending telemetry events.
deployment_environmentstringDeployment environment name (e.g., 'development', 'staging', 'production'). Used in telemetry events.
ragConfiguration for all RAG strategies (inline and tool-based).
okpOKP provider settings. Only used when 'okp' is listed in rag.inline or rag.tool.
rerankerConfiguration for neural reranking of RAG chunks using cross-encoder.
skillsAgent skills configuration. Specifies paths to skill directories.
shieldsarrayConfiguration for a single named guardrail shield (question validity or redaction).

ConversationHistoryConfiguration

Conversation history configuration.

FieldTypeDescription
typestringType of database where the conversation history is to be stored.
memoryIn-memory cache configuration
sqliteSQLite database configuration
postgresPostgreSQL database configuration

CustomProfile

Custom profile customization for prompts and validation.

FieldTypeDescription
pathstringPath to Python modules containing custom profile.
promptsobjectDictionary containing map of system prompts

Customization

Service customization.

FieldTypeDescription
profile_pathstring
disable_query_system_promptboolean
disable_shield_ids_overrideboolean
system_prompt_pathstring
system_promptstring
agent_card_pathstring
agent_card_configobject
custom_profile

DatabaseConfiguration

Database configuration.

FieldTypeDescription
sqliteSQLite database configuration
postgresPostgreSQL database configuration

FaissVectorStoreProvider

Dynamic FAISS vector-store provider (runtime create capacity).

FieldTypeDescription
idstringLlama Stack vector_io provider_id. Surrounding whitespace is stripped before validation and emission. Must match [a-z0-9_-]+ and must not start with byok_.
typestringProduct type for this dynamic vector-store provider. Must be faiss.
embedding_modelstringEmbedding model identification used for stores created against this provider. Required.
embedding_dimensionintegerDimensionality of embedding vectors for this provider. Required.
configFAISS storage settings for this provider.

FaissVectorStoreProviderConfig

Storage config for a FAISS dynamic vector-store provider.

FieldTypeDescription
pathstringOn-disk FAISS/SQLite path for this provider.

InMemoryCacheConfig

In-memory cache configuration.

FieldTypeDescription
max_entriesintegerMaximum number of entries stored in the in-memory cache

InferenceConfiguration

Inference configuration.

FieldTypeDescription
default_modelstringIdentification of default model used when no other model is specified.
default_providerstringIdentification of default provider used when no other model is specified.
context_windowsobjectMap of fully-qualified model identifier (e.g., "openai/gpt-4o-mini") to context window size in tokens. Used by the conversation compaction trigger to decide when older turns must be summarized before the input exceeds the window. Models absent from this map have no registered window โ€” callers fall back to their own default or skip the token-based trigger.
providersarrayUnified-mode synthesis input (Decision S5): a high-level, backend-agnostic list of inference providers the synthesizer expands into Llama Stack provider entries. Lives at the configuration root so it survives a future backend change. A non-empty list signals unified mode. Empty (the default) leaves legacy/remote modes unaffected. The sibling default_model / default_provider keep their query-time routing meaning and are independent of this list.
max_infer_itersintegerServer-side default for the maximum number of inference iterations a model can perform in a single request. Prevents small models from looping indefinitely on tool calls. Per-request values take precedence over this default. Set to None to disable the limit.
max_tool_callsintegerServer-side default for the maximum number of tool calls allowed in a single response. Prevents small models from exhausting the context window with repeated tool calls. Per-request values take precedence over this default. Set to None to disable the limit.

JsonPathOperator

Supported operators for JSONPath evaluation.

Note: this is not a real model, just an enumeration of all supported JSONPath operators.

JwkConfiguration

JWK (JSON Web Key) configuration.

A JSON Web Key (JWK) is a JavaScript Object Notation (JSON) data structure that represents a cryptographic key.

Useful resources:

FieldTypeDescription
urlstringHTTPS URL of the JWK (JSON Web Key) set used to validate JWTs.
jwt_configurationJWT (JSON Web Token) configuration

JwtConfiguration

JWT (JSON Web Token) configuration.

JSON Web Token (JWT) is a compact, URL-safe means of representing claims to be transferred between two parties. The claims in a JWT are encoded as a JSON object that is used as the payload of a JSON Web Signature (JWS) structure or as the plaintext of a JSON Web Encryption (JWE) structure, enabling the claims to be digitally signed or integrity protected with a Message Authentication Code (MAC) and/or encrypted.

Useful resources:

FieldTypeDescription
user_id_claimstringJWT claim name that uniquely identifies the user (subject ID).
username_claimstringJWT claim name that provides the human-readable username.
role_rulesarrayRules for extracting roles from JWT claims

JwtRoleRule

Rule for extracting roles from JWT claims.

FieldTypeDescription
jsonpathstringJSONPath expression to evaluate against the JWT payload
operatorJSON path comparison operator
negatebooleanIf set to true, the meaning of the rule is negated
valueValue to compare against
rolesarrayRoles to be assigned if the rule matches

LlamaStackConfiguration

Llama stack configuration.

Llama Stack is a comprehensive system that provides a uniform set of tools for building, scaling, and deploying generative AI applications, enabling developers to create, integrate, and orchestrate multiple AI services and capabilities into an adaptable setup.

Useful resources:

FieldTypeDescription
urlstringURL to Llama Stack service; used when library mode is disabled. Must be a valid HTTP or HTTPS URL.
api_keystringAPI key to access Llama Stack service
use_as_library_clientbooleanWhen set to true Llama Stack will be used in library mode, not in server mode (default)
library_client_config_pathstringPath to configuration file used when Llama Stack is run in library mode
timeoutintegerTimeout in seconds for requests to Llama Stack service. Default is 180 seconds (3 minutes) to accommodate long-running RAG queries.
max_retriesintegerMaximum number of connection attempts before giving up. Used on startup to connect to Llama Stack and retrieve its version. Connection attempts are retried with a fixed delay to handle the case where Llama Stack is still starting up (e.g., when running as a sidecar in the same pod).
retry_delayintegerDelay in seconds between retry attempts. Used on startup to connect to Llama Stack and retrieve its version. Connection attempts are retried with a fixed delay to handle the case where Llama Stack is still starting up (e.g., when running as a sidecar in the same pod).
allow_degraded_modebooleanIf enabled, Lightspeed Core can be started even when Llama Stack is not accessible (valid for server mode only)
configBackend-specific knobs for unified mode, where LCORE synthesizes the Llama Stack run.yaml instead of reading an external file. Holds the baseline selector, an optional profile path, and a raw native_override escape hatch. Backend-agnostic high-level sections (e.g. inference.providers) live at the configuration root, not here. Mutually exclusive with library_client_config_path; that cross-field check lives on the root Configuration model. When set in library mode, library_client_config_path is not required.

ModelContextProtocolServer

Model context protocol server configuration.

MCP (Model Context Protocol) servers provide tools and capabilities to the AI agents. These are configured by this structure. Only MCP servers defined in the lightspeed-stack.yaml configuration are available to the agents. Tools configured in the llama-stack run.yaml are not accessible to lightspeed-core agents.

Useful resources:

FieldTypeDescription
namestringMCP server name that must be unique
provider_idstringMCP provider identification
urlstringURL of the MCP server
authorization_headersobjectHeaders to send to the MCP server. The map contains the header name and the path to a file containing the header value (secret). There are 3 special cases: 1. Usage of the kubernetes token in the header. To specify this use a string 'kubernetes' instead of the file path. 2. Usage of the client-provided token in the header. To specify this use a string 'client' instead of the file path. 3. Usage of the oauth token in the header. To specify this use a string 'oauth' instead of the file path.
headersarrayList of HTTP header names to automatically forward from the incoming request to this MCP server. Headers listed here are extracted from the original client request and included when calling the MCP server. This is useful when infrastructure components (e.g. API gateways) inject headers that MCP servers need, such as x-rh-identity in HCC. Header matching is case-insensitive. These headers are additive with authorization_headers and MCP-HEADERS.
require_approvalWhen to require human approval for tool invocations. 'always' requires approval for all tools, 'never' auto-approves, or use ApprovalFilter for granular control.
timeoutintegerTimeout in seconds for requests to the MCP server. If not specified, the default timeout from Llama Stack will be used. Note: This field is reserved for future use when Llama Stack adds timeout support.

OkpConfiguration

OKP (Offline Knowledge Portal) provider configuration.

Controls provider-specific behaviour for the OKP vector store. Only relevant when "okp" is listed in rag.inline or rag.tool.

FieldTypeDescription
rhokp_urlstringBase URL for the OKP server (http or https). Set to ${env.RH_SERVER_OKP} in YAML to use the environment variable. When unset, the default from constants is used.
offlinebooleanWhen True, use parent_id for OKP chunk source URLs. When False, use reference_url for chunk source URLs.
chunk_filter_querystringAdditional OKP filter query applied to every OKP search request. Use Solr boolean syntax, e.g. 'product:ansible AND product:openshift'.

PgvectorVectorStoreProvider

Dynamic pgvector vector-store provider (runtime create capacity).

FieldTypeDescription
idstringLlama Stack vector_io provider_id. Surrounding whitespace is stripped before validation and emission. Must match [a-z0-9_-]+ and must not start with byok_.
typestringProduct type for this dynamic vector-store provider. Must be pgvector.
embedding_modelstringEmbedding model identification used for stores created against this provider. Required.
embedding_dimensionintegerDimensionality of embedding vectors for this provider. Required.
configpgvector connection settings for this provider.

PgvectorVectorStoreProviderConfig

Storage config for a pgvector dynamic vector-store provider.

FieldTypeDescription
hoststringPostgreSQL host. Defaults to ${env.POSTGRES_HOST}.
portstringPostgreSQL port. Defaults to ${env.POSTGRES_PORT}.
dbstringPostgreSQL database name. Defaults to ${env.POSTGRES_DATABASE}.
userstringPostgreSQL user. Defaults to ${env.POSTGRES_USER}.
passwordstringPostgreSQL password. Defaults to ${env.POSTGRES_PASSWORD}.

PostgreSQLDatabaseConfiguration

PostgreSQL database configuration.

PostgreSQL database is used by Lightspeed Core Stack service for storing information about conversation IDs. It can also be leveraged to store conversation history and information about quota usage.

Useful resources:

FieldTypeDescription
hoststringDatabase server host or socket directory
portintegerDatabase server port
dbstringDatabase name to connect to
userstringDatabase user name used to authenticate
passwordstringPassword used to authenticate
namespacestringDatabase namespace
ssl_modestringSSL mode
gss_encmodestringThis option determines whether or with what priority a secure GSS TCP/IP connection will be negotiated with the server.
ca_cert_pathstringPath to CA certificate

QuotaHandlersConfiguration

Quota limiter configuration.

It is possible to limit quota usage per user or per service or services (that typically run in one cluster). Each limit is configured as a separate quota limiter. It can be of type user_limiter or cluster_limiter (which is name that makes sense in OpenShift deployment).

FieldTypeDescription
sqliteSQLite database configuration
postgresPostgreSQL database configuration
limitersarrayQuota limiters configuration
schedulerQuota scheduler configuration
enable_token_historybooleanEnables storing information about token usage history

QuotaLimiterConfiguration

Configuration for one quota limiter.

There are three configuration options for each limiter:

  1. period is specified in a human-readable form, see https://www.postgresql.org/docs/current/datatype-datetime.html#DATATYPE-INTERVAL-INPUT for all possible options. When the end of the period is reached, the quota is reset or increased.
  2. initial_quota is the value set at the beginning of the period.
  3. quota_increase is the value (if specified) used to increase the quota when the period is reached.

There are two basic use cases:

  1. When the quota needs to be reset to a specific value periodically (for example on a weekly or monthly basis), set initial_quota to the required value.
  2. When the quota needs to be increased by a specific value periodically (for example on a daily basis), set quota_increase.
FieldTypeDescription
typestringQuota limiter type, either user_limiter or cluster_limiter
namestringHuman readable quota limiter name
initial_quotaintegerQuota set at beginning of the period
quota_increaseintegerDelta value used to increase quota when period is reached
periodstringPeriod specified in human readable form

QuotaSchedulerConfiguration

Quota scheduler configuration.

FieldTypeDescription
periodintegerQuota scheduler period specified in seconds
database_reconnection_countintegerDatabase reconnection count on startup. When database for quota is not available on startup, the service tries to reconnect N times with specified delay.
database_reconnection_delayintegerDatabase reconnection delay specified in seconds. When database for quota is not available on startup, the service tries to reconnect N times with specified delay.

RHIdentityConfiguration

Red Hat Identity authentication configuration.

FieldTypeDescription
required_entitlementsarrayList of all required entitlements.
max_header_sizeintegerMaximum allowed size in bytes for the base64-encoded x-rh-identity header. Headers exceeding this size are rejected before decoding.

RagConfiguration

RAG strategy configuration.

Controls which RAG sources are used for inline and tool-based retrieval.

Each strategy lists RAG IDs to include. The special ID "okp" defined in constants, activates the OKP provider; all other IDs refer to entries in byok_rag.

Both inline and tool default to [] (disabled). Each must be explicitly configured to activate its respective RAG strategy.

FieldTypeDescription
inlinearrayRAG IDs whose sources are injected as context before the LLM call. Use 'okp' to enable OKP inline RAG. Empty by default (no inline RAG).
toolarrayRAG IDs made available to the LLM as a file_search tool. Use 'okp' to include the OKP vector store. When omitted, tool RAG is disabled.

RerankerConfiguration

Reranker configuration for RAG chunk reranking.

FieldTypeDescription
enabledbooleanWhen True, reranking applied to RAG chunks. When False, reranking is disabled and original scoring used.
modelstringCross-encoder model name for reranking RAG chunks. Defaults to 'cross-encoder/ms-marco-MiniLM-L6-v2' from sentence-transformers.

RlsapiV1Configuration

Configuration for the rlsapi v1 /infer endpoint.

Settings specific to the RHEL Lightspeed Command Line Assistant (CLA) stateless inference endpoint. Kept separate from shared configuration sections so that CLA-specific options do not affect other endpoints.

FieldTypeDescription
allow_verbose_inferbooleanAllow /v1/infer to return extended metadata (tool_calls, rag_chunks, token_usage) when the client sends "include_metadata": true. Should NOT be enabled in production. If production use is needed, consider RBAC-based access control via an Action.RLSAPI_V1_INFER authorization rule.
quota_subjectstringIdentity field used as the quota subject for /v1/infer. When set, token quota enforcement is enabled for this endpoint. Requires quota_handlers to be configured. "org_id" and "system_id" require rh-identity authentication; falls back to user_id when rh-identity data is unavailable.

SQLiteDatabaseConfiguration

SQLite database configuration.

FieldTypeDescription
db_pathstringPath to file where SQLite database is stored

ServiceConfiguration

Service configuration.

Lightspeed Core Stack is a REST API service that accepts requests on a specified hostname and port. It is also possible to enable authentication and specify the number of Uvicorn workers. When more workers are specified, the service can handle requests concurrently.

FieldTypeDescription
hoststringService hostname
portintegerService port
base_urlstringExternally reachable base URL for the service; needed for A2A support.
auth_enabledbooleanEnables the authentication subsystem
workersintegerNumber of Uvicorn worker processes to start
color_logbooleanEnables colorized logging
access_logbooleanEnables logging of all access information
tls_configTransport Layer Security configuration for HTTPS support
root_pathstringASGI root path for serving behind a reverse proxy on a subpath
corsCross-Origin Resource Sharing configuration for cross-domain requests

SkillsConfiguration

Agent skills configuration.

Specifies paths to skill directories. Skill metadata (name, description) is read from SKILL.md frontmatter at startup.

Each path can point to either:

  • A directory containing a SKILL.md file (single skill)
  • A directory containing subdirectories with SKILL.md files (multiple skills)

Paths are validated at startup to ensure they exist and contain valid SKILL.md files.

FieldTypeDescription
pathsarrayPaths to skill directories or directories containing skill subdirectories.

QuestionValidityConfig

Configuration for the question validity guardrail.

FieldTypeDescription
model_idstringThe model_id to use for the guard
model_promptstringPrompt sent to the LLM used to validate the user's question
invalid_question_responsestringResponse when the user's question is determined to be invalid

QuestionValidityShieldConfiguration

Configuration for a named question-validity guardrail shield.

FieldTypeDescription
namestringUnique, user-facing name identifying this shield instance
provider_idstringDiscriminator identifying this as a question-validity shield
configQuestion-validity-specific configuration for this shield

RedactionRule

A single regex-based redaction rule.

FieldTypeDescription
patternstringRegex pattern to match sensitive data
replacementstringReplacement string for matched text
case_sensitivebooleanPer-rule override; when null, the global RedactionConfig flag applies

RedactionConfig

Configuration for PII redaction with regex-based rules.

FieldTypeDescription
rulesarrayOrdered list of PII redaction rules
case_sensitivebooleanWhen false, patterns are compiled with re.IGNORECASE

RedactionShieldConfiguration

Configuration for a named PII-redaction guardrail shield.

FieldTypeDescription
namestringUnique, user-facing name identifying this shield instance
provider_idstringDiscriminator identifying this as a redaction shield
configRedaction-specific configuration for this shield

SplunkConfiguration

Splunk HEC (HTTP Event Collector) configuration.

Splunk HEC allows sending events directly to Splunk over HTTP/HTTPS. This configuration is used to send telemetry events for inference requests to the corporate Splunk deployment.

Useful resources:

FieldTypeDescription
enabledbooleanEnable or disable Splunk HEC integration.
urlstringSplunk HEC endpoint URL.
token_pathstringPath to file containing the Splunk HEC authentication token.
indexstringTarget Splunk index for events.
sourcestringEvent source identifier.
timeoutintegerHTTP timeout in seconds for HEC requests.
verify_sslbooleanWhether to verify SSL certificates for HEC endpoint.

TLSConfiguration

TLS configuration.

Transport Layer Security (TLS) is a cryptographic protocol designed to provide communications security over a computer network, such as the Internet. The protocol is widely used in applications such as email, instant messaging, and voice over IP, but its use in securing HTTPS remains the most publicly visible.

Useful resources:

FieldTypeDescription
tls_certificate_pathstringSSL/TLS certificate file path for HTTPS support.
tls_key_pathstringSSL/TLS private key file path for HTTPS support.
tls_key_passwordstringPath to file containing the password to decrypt the SSL/TLS private key.

TrustedProxyConfiguration

Configuration for trusted-proxy auth module.

FieldTypeDescription
user_headerstringHTTP header containing the forwarded user identity.
allowed_service_accountsarrayOptional allowlist of Kubernetes ServiceAccount identities permitted to act as trusted proxies. When set to null/omitted, any ServiceAccount with a valid token is accepted. When set to a non-empty list, only the listed ServiceAccounts are allowed. An empty list behaves the same as null (no restriction).

TrustedProxyServiceAccount

A Kubernetes ServiceAccount identity for trusted-proxy allowlist.

FieldTypeDescription
namespacestringKubernetes namespace of the ServiceAccount.
namestringName of the Kubernetes ServiceAccount.

UnifiedInferenceProvider

A high-level inference provider entry for unified-mode synthesis.

Operators describe inference providers at this high level (backend-agnostic vocabulary) instead of authoring raw Llama Stack provider blocks. The synthesizer (apply_high_level_inference) expands each entry into a Llama Stack providers.inference entry, mapping type to a provider_type and emitting ${env.<VAR>} references for secrets (never literal values).

Attributes: type: Canonical provider identifier. Vendor-neutral so it survives a future backend change; each backend-specific synthesizer maps it to its own provider vocabulary. id: Optional identifier emitted as the Llama Stack provider_id. When omitted, synthesized as type with underscores hyphenated. If set, must be non-empty after stripping whitespace and may contain only lowercase letters, digits, underscores, and hyphens. api_key_env: Name of the environment variable holding the provider API key. Emitted verbatim as ${env.<name>} so the secret never lands on disk resolved. allowed_models: Optional allow-list of model identifiers passed through to the synthesized provider config. extra: Additional provider-config keys merged verbatim into the synthesized provider's config block โ€” an escape hatch for provider-specific knobs not modeled here.

FieldTypeDescription
typestringCanonical, backend-agnostic provider identifier mapped to a Llama Stack provider_type by the synthesizer.
idstringOptional identifier emitted as the Llama Stack provider_id. When omitted, synthesized as type with underscores hyphenated. If set, must be non-empty after stripping whitespace and may contain only lowercase letters, digits, underscores, and hyphens.
api_key_envstringName of the environment variable holding the provider API key. Emitted as a ${env.} reference so the secret is never written to disk in resolved form.
allowed_modelsarrayOptional allow-list of model identifiers for this provider.
extraobjectAdditional provider-config keys merged verbatim into the synthesized provider's config block.

UnifiedLlamaStackConfig

Backend-specific knobs for unified-mode Llama Stack synthesis.

Per Decision S5 of the design spike, backend-agnostic high-level sections (inference, ...) live at the configuration root, not here. This block holds only the Llama-Stack-specific synthesis controls: which baseline to start from, an optional profile file, and a raw native_override escape hatch.

During synthesis from the default baseline or a profile, LCORE ensures the Llama Stack MCP tool_runtime provider (provider_id: model-context-protocol, provider_type: remote::model-context-protocol) is present so static mcp_servers and dynamic MCP registration work. That ensure is skipped when baseline: empty (migration / blank-slate); supply MCP via native_override in that case.

Attributes: baseline: Synthesis starting point. "default" begins from LCORE's built-in baseline (src/data/default_run.yaml); "empty" begins from an empty dict (used by the migration tool for an exact round-trip). Ignored when profile is set. profile: Optional path to a user-authored run.yaml-shaped file used as the synthesis baseline. Relative paths resolve against the directory of the loaded lightspeed-stack.yaml. native_override: Raw Llama Stack schema deep-merged last (maps merge recursively, lists and scalars replace). The escape hatch for anything the high-level sections do not express.

FieldTypeDescription
baselinestringSynthesis starting point: 'default' uses LCORE's built-in baseline, 'empty' starts from {}. Ignored when 'profile' is set.
profilestringPath to a run.yaml-shaped baseline file. Relative paths resolve against the directory of the loaded lightspeed-stack.yaml.
native_overrideobjectRaw Llama Stack schema deep-merged last (maps merge recursively; lists and scalars replace).

UserDataCollection

User data collection configuration.

FieldTypeDescription
feedback_enabledbooleanWhen set to true the user feedback is stored and later sent for analysis.
feedback_storagestringPath to directory where feedback will be saved for further processing.
transcripts_enabledbooleanWhen set to true the conversation history is stored and later sent for analysis.
transcripts_storagestringPath to directory where conversation history will be saved for further processing.

VectorStoreConfiguration

Configuration for dynamic vector-store providers.

Mirrors InferenceConfiguration: a providers list plus a sibling default_provider pointer, rather than a per-entry default flag.

FieldTypeDescription
default_providerstringProvider id used for vector_stores.default_* in the synthesized Llama Stack config. Required when providers is non-empty; must match one of providers[].id. Must be omitted when providers is empty.
providersarrayDynamic vector-store provider capacity for runtime POST /v1/vector-stores creates. Not the same as byok_rag (static registered corpora).