Named AI Deployment Profiles
September 18, 2026 ยท View on GitHub
Ax selects deployment behavior with the name passed to ai(). The model ID is
data inside that deployment; it does not select another provider's request
rules.
import { ai } from '@ax-llm/ax';
const together = ai({
name: 'together',
apiKey: process.env.TOGETHER_API_KEY!,
config: { model: 'deepseek-ai/DeepSeek-V4-Pro' },
});
This request uses Together's endpoint, authentication, reasoning-effort
mapping, and response fields. It does not use DeepSeek's native thinking
request object merely because the model ID contains DeepSeek.
Reasoning replay is profile-owned too. Together replays its reasoning field,
Fireworks replays reasoning_content, and OpenRouter can replay both plaintext
reasoning and structured reasoning_details; Ax does not translate these by
sniffing the model name.
Discovery
Use axAIProfiles() to build selectors or inspect the complete catalog, and
axGetAIProfile(id) to inspect one profile. The returned summary includes the
transport, endpoint requirements, authentication, operations, default
capabilities, deployment-scoped model rules, official sources, and review date.
import { axAIProfiles, axGetAIProfile } from '@ax-llm/ax';
const available = axAIProfiles();
const fireworks = axGetAIProfile('fireworks');
See the generated deployment profile matrix for all profiles and their official source links.
OpenAI And OpenAI-Compatible Endpoints
openai means the official OpenAI Chat Completions deployment and carries
OpenAI's verified behavior. openai-compatible is the conservative profile for
a custom endpoint:
const custom = ai({
name: 'openai-compatible',
apiURL: 'https://gateway.example/v1',
apiKey: process.env.GATEWAY_API_KEY,
config: { model: 'organization/model-id' },
});
Unknown profile names fail with the known profile IDs instead of silently falling back to the compatibility profile.
Meta Model API
Meta has three explicit hosted profiles with shared bearer authentication and
the https://api.meta.ai/v1 base URL:
metauses Responses and is the recommended default.meta-chatuses OpenAI-compatible Chat Completions.meta-messagesuses Anthropic-compatible Messages.
All default to muse-spark-1.3. The Responses profile also runs
muse-image-1.0 through chat(), returning generated or edited images in the
normal results[].images field, and routes muse-voice-transcribe-1.0 through
transcribe() or realtime PCM chat(). Ax deliberately does not add direct
Images, Files, Models, or Responses resource-management methods.
The same profile descriptors, request/stream mappings, image results, transcription fields, and realtime event folding are generated for Python, Go, Rust, Java, and C++ from AxIR.
Contributor model variants are cataloged with provider-training data-use
metadata and are never defaults. Muse Glimmer remains self-hosted: use the
existing vllm, llama-cpp, ollama, or lm-studio profile with a server you
operate. Ax does not download or manage model weights.
Typesafe / Jev
The typesafe deployment uses the native System One transport at
https://api.typesafe.ai. Jev is its model family; both Ax interfaces default
to jev-latest.
ai({ name: 'typesafe', apiKey, trueThreshold: 0.9 })supports ordinary Ax programs with required boolean and class outputs. Value descriptions become native criteria and remain readable prompt/schema descriptions with other providers. Boolean conversion isnoul >= trueThreshold(default0.5).typesafe({ apiKey }).systemOne({ state, questions })exposes native Noul, Choice, and Score answers. Structured criteria and fractional scoring remain native; numeric signature bounds are validation constraints. NativelistModels()is separate from Ax's configured model aliases.
Unsupported schemas and tools are excluded before requests, including router fallbacks with degradation enabled. Mixed pools require an already-supported schema for Typesafe eligibility; Typesafe-only pools propagate its schema requirement. Use an explicit generative second step for prose.
The Typesafe/Jev skill is the detailed usage guide, including transport settings, examples, and provider-specific limits. The adapter, native client, value descriptions, and routing validation are also generated for Python, Java, C++, Go, and Rust. Their packages include a dedicated Typesafe/Jev skill and provider-backed signature, native, and hybrid examples.
Capability Resolution
Capabilities resolve in this order:
- caller
modelInfooverride; - exact model rule in the selected profile;
- profile-scoped prefix or substring rule;
- profile default;
- conservative transport default.
Ax never sniffs endpoint URLs or applies another profile's rules. Explicitly
requesting thinking or structured output for an unverified profile/model fails
before the request. A profile may demote Ax's synthetic __axOutput forced tool
choice when the deployment supports tools but not tool_choice; a caller's
explicitly forced tool choice fails instead of being silently changed.
Structured-Output Modes
AxStructuredOutputRung is native | function | json_object, and
AxStructuredOutputMode adds auto. Profiles and model rules advertise an
ordered structuredOutputModes list. auto uses that order, except that a
required singleton string or code field can take the optimized json_object
path when native schema is unavailable. An explicit mode must be advertised or
Ax fails before network I/O. Custom clients that do not expose the new list keep
the legacy capability heuristic.
structuredOutputs remains the compatibility alias for native JSON Schema
support. It does not imply json_object: direct json_schema and
json_object chat requests validate those capabilities independently.
Vertex rules are deliberately model-scoped. Documented Gemini MaaS IDs prefer
native, then function, then json_object. The exact
google/gemma-4-26b-a4b-it-maas rule prefers json_object, supports
function, and excludes native schema. Unknown Vertex models remain at the
conservative profile default.
Renewable Credentials
Use credentialProvider when a deployment token can expire. Ax calls it for
each request attempt with { profile, operation, method, url }; its returned
headers override static authentication headers. The callback covers chat,
streaming, embeddings, Responses, transcription, speech, and retries. Callback
errors stop before transport. Ax does not automatically replay a completed 401
or 403 because generation requests are not assumed idempotent.
const vertex = ai({
name: 'vertex-ai',
apiURL: process.env.VERTEX_AI_API_URL!,
config: { model: 'google/gemma-4-26b-a4b-it-maas' },
credentialProvider: async ({ operation, url }) => ({
Authorization: `Bearer ${await accessTokens.getFreshToken({ operation, url })}`,
}),
});
Authentication-required profiles accept either apiKey or a credential
provider. Keep cloud SDK and ADC dependencies in the host application; the
provider callback is the dependency-free Ax boundary.
Thinking defaults are deployment- and model-scoped request rules. Omitting
thinkingTokenBudget selects Ax's logical max level for verified reasoning
models and maps it to the strongest effort documented by that deployment:
| Profile | Verified model rule | Logical max wire value | Explicit none |
|---|---|---|---|
deepseek, together, fireworks, openrouter | Deployment-specific DeepSeek rules | Deployment-specific | Supported only where documented |
grok | Grok 4.6 / 4.5 / 4.3 | xhigh / high / high | Rejected for 4.6 and 4.5 |
groq | GPT-OSS 20B/120B; Qwen 3.6 27B | high; default (reasoning enabled) | Rejected for GPT-OSS; supported for Qwen |
cerebras | GPT-OSS 120B; Gemma 4 31B | high | Rejected for GPT-OSS; supported for Gemma |
deepinfra | DeepSeek R1 family | high | Supported |
The native google-gemini profile and its gemini / google_gemini aliases
resolve the effective model before mapping an explicit logical budget. Gemini 3
uses model-supported thinkingLevel values, while Gemini 2.5 uses numeric
thinkingBudget; none always hides thoughts and clamps to the model's minimum
when thinking cannot be disabled. The OpenAI-compatible vertex-ai profile
does not inherit this native request shape from Gemini-looking model IDs.
An explicitly unsupported level fails before the request. Hugging Face Router
stays conservative because its :fastest, :cheapest, and :preferred
policies can choose different underlying providers for the same model ID; use
an exact caller modelInfo override only after verifying the selected route.
No model inherits behavior merely because its ID resembles another provider's
model.
Major-Version Migration
Profile-only provider classes were removed. Keep model enum/catalog imports if
they are useful, but construct the deployment through ai():
| Removed construction | Replacement |
|---|---|
new AxAIAzureOpenAI(args) | ai({ name: 'azure-openai', ...args }) |
new AxAICohere(args) | ai({ name: 'cohere', ...args }) |
new AxAIDeepSeek(args) | ai({ name: 'deepseek', ...args }) |
new AxAIDeepSeekResponses(args) | ai({ name: 'deepseek-responses', ...args }) |
new AxAIMistral(args) | ai({ name: 'mistral', ...args }) |
new AxAIReka(args) | ai({ name: 'reka', ...args }) |
new AxAIGrok(args) | ai({ name: 'grok', ...args }) |
The genuine transport/runtime classes remain: OpenAI-compatible Chat Completions, OpenAI Responses, Anthropic Messages, Gemini GenerateContent, and WebLLM. Generated Go, Python, Rust, Java, and C++ packages follow the same boundary: their named factories resolve a profile and construct a retained transport client; branded client constructors no longer exist.
GraphJin Vertex migration
Remove the request-rewriting responseFormatClient. Select Vertex explicitly,
let the profile resolve the model's output modes, and refresh credentials at the
request boundary:
client := ax.NewAI("vertex-ai", map[string]ax.Value{
"api_url": vertexOpenAIURL,
"model": "google/gemma-4-26b-a4b-it-maas",
"credential_provider": ax.AxCredentialProviderFunc(
func(ctx context.Context, request ax.AxCredentialRequest) (map[string]string, error) {
token, err := vertexAccessToken(ctx) // host-owned ADC or token source
if err != nil { return nil, err }
return map[string]string{"Authorization": "Bearer " + token}, nil
},
),
})
Ax's Gemma rule sends response_format: {type: "json_object"}, defaults
thinking to max, writes chat_template_kwargs.enable_thinking: true, and
extracts/replays reasoning_content. AxGen already supplies the exact-shape
prompt and client-side validation, so do not append a second schema prompt. To
make the rich-output choice explicit, pass
structured_output_mode: "json_object" in the generated Go forward options.
This mapping is scoped from the
GraphJin compatibility workaround,
Google's official structured-output
and thinking
documentation, and the official Vertex OpenAI-compatible endpoint guide.
Maintaining The Catalog
ir/axcore/data/provider-profiles.json is the source of truth. Run:
npm run profiles:generate
npm run axir:conformance:write
npm run axir:generate-packages
The first command validates IDs, aliases, transports, operation dialects, endpoint/auth requirements, model-rule precedence, and source metadata, then regenerates the TypeScript registry, AxIR registry/descriptors, and profile matrix. Live credentialed provider smoke tests are opt-in and are not mandatory CI.