Iris on CoreWeave

August 30, 2026 · View on GitHub

Iris runs CoreWeave jobs directly as Kubernetes Pods. The controller does not start Iris worker daemons on CoreWeave. This page describes the current architecture and configuration boundary.

Use the cloud GPU tutorial to run a job, the federation reference to understand cross-cluster routing, and the Iris operations guide for live diagnosis and approved operational changes.

Quickstart

The marin federation config lists three CoreWeave research peers:

Iris clusterAcceleratorGPUs per node
cw-rno2aH1008
cw-us-east-02aH1008
cw-us-east-08aGB2004

The table identifies configured hardware, not live free capacity. The canonical inventory is lib/iris/config/marin.yaml; each peer's hardware and Kubernetes settings live in lib/iris/config/<cluster>.yaml.

Use the use-iris skill for a short development session and the cloud GPU tutorial for a normal Iris job. These commands give a read-only view of a candidate cluster:

CLUSTER=cw-rno2a

uv run iris --cluster="$CLUSTER" cluster status
uv run iris --cluster="$CLUSTER" rpc controller list-backends
uv run iris --cluster="$CLUSTER" cluster dashboard

list-backends reports accelerator groups, nodes, and current availability. CoreWeave has no Iris worker daemon, so the worker count in cluster status is not a capacity signal. cluster dashboard holds a port-forward open until Ctrl+C.

Controller lifecycle commands and direct Kubernetes writes change a shared cluster. Use the use-iris skill and obtain explicit approval before running them.

Connecting

Operator-side Kubernetes access uses the kubeconfig and context in the selected cluster file. Current research configs use ~/.kube/coreweave-iris and pin a context, so do not switch the file's current context.

For first-time access, follow CoreWeave's token and kubeconfig guide for each required CKS cluster. Merge the downloaded cluster, user, and context entries into ~/.kube/coreweave-iris; preserve existing entries and set the file mode to 600.

If cluster status or cluster dashboard says the context does not exist, compare the configured values with the file Iris is actually reading:

CLUSTER=cw-rno2a

printf 'KUBECONFIG=%s\n' "${KUBECONFIG:-<unset>}"
rg -n 'kubeconfig_path|kube_context|namespace' "lib/iris/config/${CLUSTER}.yaml"
kubectl --kubeconfig ~/.kube/coreweave-iris config get-contexts -o name

env -u KUBECONFIG uv run iris --cluster="$CLUSTER" cluster status

An exported KUBECONFIG selects the file, while the cluster config still pins the context. Unset a stale override instead of copying contexts between files. Iris commands that target CoreWeave use a Kubernetes port-forward and loopback trust; they do not need iris login.

For a live access incident, continue in CoreWeave GPU Operations. It separates the port-forward, controller Service, Traefik ingress, and public LoadBalancer layers.

Architecture

The control path is:

  1. The Iris controller runs as a Kubernetes Deployment in the cluster.
  2. K8sTaskProvider applies one Pod for each task attempt and reconciles Pod state back into Iris.
  3. Kueue admits every task Pod. It gang-admits coscheduled jobs and applies the configured topology constraints.
  4. Kubernetes places Pods on CoreWeave NodePools. CoreWeave, not Iris, manages node provisioning.

There is no Iris worker daemon, synthetic worker row, or Iris-managed slice on this path. An Iris node-agent DaemonSet publishes same-node host and GPU measurements to telemetry_v1. The cluster view retains Kubernetes readiness, allocatable capacity, scheduling, and Pod state; hardware history remains in telemetry_v1. K8sTaskProvider.schedule() and autoscale() are no-ops. Capacity and task state come from Kubernetes nodes, Pods, and Kueue workloads.

GPU Pods request nvidia.com/gpu. On clusters with host_network: true, they also request rdma/ib and use ClusterFirstWithHostNet. Coscheduled jobs add Kueue topology annotations derived from kubernetes_provider.kueue.topologies, or from the CoreWeave defaults when the map is empty.

Task logs are shipped from the node's container log file to Finelog by a sidecar. Process inspection and profiling run against the task container with kubectl exec; they do not call a worker service.

Federation ingress, authentication, and DNS

The public controller hostname is only a federation route. DNS points it at the CoreWeave Traefik LoadBalancer, cert-manager issues its TLS certificate, and the iris-federation-ipallowlist middleware admits only the configured federation parent egress addresses. The controller then applies a second gate:

  • in-cluster private addresses and loopback are trusted by auth.trusted_cidrs;
  • off-cluster requests must carry a bearer token;
  • federation tokens are verified with the public keys in auth.federation_peers.

The allowlist and controller authentication are independent. Do not weaken one because the other is working. Pulumi owns Traefik, cert-manager, the Middleware, and Cloudflare DNS; the cluster config supplies their inputs. See the network manifest builders, the Pulumi Traefik component, and the federation reference for the handoff protocol.

GPU topology and Kueue safety

H100 gangs use preferred InfiniBand leaf-group colocation. GB200 and GB300 NVL72 gangs use stricter rack-aware placement:

  • an 18-node rack has 16 nodes guaranteed schedulable at once;
  • up to 16 replicas bind to one NVLink domain;
  • larger gangs require one node-saturating Pod with all four GPUs per tray, then split evenly over the fewest racks with 10 to 16 replicas per rack slice. Valid examples include 20, 24, 32, and 48 replicas.

Invalid multi-rack shapes fail before submission rather than silently losing the one-slice-per-rack guarantee. The canonical arithmetic is in coreweave_topology.py.

Kueue's admission webhooks must remain scoped to the Iris task namespace. An unscoped fail-closed webhook can block CNI and system Pods and deadlock node delivery before Kueue itself starts. Pulumi pins the chart and applies the namespace scope in KueueAddon; do not replace that release with an ad hoc cluster-wide install.

Pulumi pins cks-kueue 1.5.0, which contains Kueue 0.18.3. Iris explicitly enables TASRecomputeAssignmentWithinSchedulingCycle, added in Kueue 0.18.2, so a workload's hostname assignment is recomputed when an earlier admission in the same scheduling cycle consumes that domain.

TAS preemption and CPU spillover

All Iris Pods use the topology-aware cw-tas ResourceFlavor. Its iris.kueue=true selector covers every Iris-managed NodePool. Accelerator-free Pods request unconstrained TAS, so Kueue records their per-node CPU reservations in the same flavor as GPU gangs. The ClusterQueue's binding GPU and RDMA quotas are derived from the configured maximum GPU node count and per-node device count; CPU, memory, disk, and Pods retain non-binding sentinels. Accelerator quota pressure activates preemption.withinClusterQueue: LowerPriority, and TAS then checks whether removing lower-priority candidates creates a compatible topology before admitting the waiting Workload.

Each Iris band has three Kueue admission tiers. The controller reconciles the twelve WorkloadPriorityClass objects at startup.

Iris bandOrdinary CPUStandalone acceleratorCo-scheduled group
batch012
interactive101112
production100010001000
system100001000010000

Batch and interactive workloads are ordered CPU < accelerator < co-scheduled group within their band. Kueue can reclaim same-band CPU reservations for one accelerator Pod, or both lower tiers for a co-scheduled GPU group. SYSTEM and PRODUCTION workloads share one Kueue priority within their respective bands, so request shape cannot cause same-band preemption. A higher band still outranks every tier in the band below it. Pod priorityClassName remains the ordinary Iris band, so this ordering affects Kueue admission and preemption but does not change kube-scheduler priority within an admitted workload.

For example, a batch CPU coordinator uses tier 0, its separately admitted accelerator child uses tier 1, and a co-scheduled CPU/GPU group uses tier 2. Choose the Iris band for operational importance; Iris derives the tier from the request shape.

The lowest batch Workload priority is 0. CoreWeave node-health-check Pods use Kubernetes priority -1, and Iris batch Pods retain Kubernetes priority 0. Kueue Workload priority and Kubernetes Pod priority are separate scheduling domains, but neither representation places Iris work below the health checker.

flowchart TD
    request[RunTaskRequest] --> protected{System or production band?}
    protected -- Yes --> fixed[Use the band's fixed priority]
    protected -- No --> gang
    gang -- Yes --> native[Use band plus 2<br/>co-scheduled tier]
    gang -- No --> accelerator{Accelerator requested?}
    accelerator -- Yes --> gpu[Use band plus 1<br/>accelerator tier]
    accelerator -- No --> cpu[Use native Iris band<br/>CPU tier]
    fixed --> queue[Kueue LocalQueue and shared ClusterQueue]
    native --> queue[Kueue LocalQueue and shared ClusterQueue]
    gpu --> queue
    cpu --> queue
    queue --> quota{Accelerator quota available?}
    quota -- No --> preempt[withinClusterQueue: LowerPriority<br/>select compatible victims]
    preempt --> fit{Quota and TAS fit after victims?}
    quota -- Yes --> fit
    fit -- Yes --> admit[Admit workload and release scheduling gate]
    fit -- No --> wait[Remain SchedulingGated]
    admit --> schedule[Kubernetes schedules Pods]

Accelerator-free jobs use any compatible node by default. Iris does not expose a CPU-only placement constraint; rare jobs that require hard CPU-node placement must use Kubernetes-native scheduling outside Iris. GPU and RDMA resource requests exclude CPU nodes without an additional selector.

Kueue requires every node in a TAS flavor to carry every level in the referenced Topology. CoreWeave supplies the physical hierarchy on accelerator nodes. Iris labels CPU NodePools with a synthetic iris-cpu-only fabric, superpod, leafgroup, and NVLink domain so unconstrained TAS can assign them at the hostname level. The synthetic values do not advertise GPU or RDMA resources.

This layout replaces the selectorless, non-TAS cw-cpu flavor that caused #7916. Kueue v0.18 could not reclaim those Pods during cw-ib topology fit; the general upstream case is tracked by kubernetes-sigs/kueue#9992. When migrating from the split flavors, apply the NodePool labels before switching the ClusterQueue to cw-tas, then verify CPU nodes appear in Kueue's topology cache.

TAS admission pins each Pod to a hostname from Kueue's capacity snapshot. The same-cycle recomputation gate prevents two newly admitted Workloads from keeping the same hostname when only one has capacity. If capacity becomes unavailable after admission for another reason, kube-scheduler reports an affinity plus resource failure and Iris returns the attempt to PENDING for fresh admission.

Resource ownership

ResourceOwner
CKS cluster and operator kubeconfigCoreWeave and the cluster operator
Namespace, RBAC, NodePools, Kueue operator, ClusterQueue, ResourceFlavor, ingress, and DNSinfra/pulumi
Pod and Workload priority classes, Kueue LocalQueue, controller ConfigMap, Deployment, Service, PDB, state volume, and SecretsK8sControllerProvider
Task Pods and their lifecycleK8sTaskProvider
Node scheduling and provisioningKubernetes and CoreWeave

K8sControllerProvider.verify_prerequisites() only checks that the Pulumi-owned resources exist. It does not create or repair them. Read the Pulumi guide before changing that substrate; a destructive NodePool plan can deprovision reserved hardware.

Configuration

lib/iris/src/iris/cluster/config.py defines the schema. The checked-in cluster files are the source of truth for deployed values.

Platform and controller

FieldMeaning
platform.label_prefixPrefix for managed labels and NodePool names.
platform.coreweave.regionCoreWeave region for this CKS cluster.
platform.coreweave.namespaceNamespace used by the controller lifecycle; defaults to iris.
platform.coreweave.kubeconfig_pathOperator-side kubeconfig. Current clusters use ~/.kube/coreweave-iris.
platform.coreweave.kube_contextContext bound to every operator-side Kubernetes call. Do not rely on the kubeconfig's current context.
platform.coreweave.object_storage_endpointS3 endpoint seen by Pods inside CoreWeave.
platform.coreweave.external_object_storage_endpointEndpoint for the same store when Iris runs outside CoreWeave. It falls back to the internal endpoint when empty.
controller.coreweave.portController RPC port; defaults to 10000.
controller.coreweave.service_nameIn-cluster Service name; defaults to iris-controller-svc.
controller.coreweave.scale_groupCPU scale group that hosts the controller. This is required.
controller.coreweave.ingress_classIngress class used by the external federation route.

The operator kubeconfig path and context are removed before the cluster config is written to the in-cluster ConfigMap. The controller and task provider use their Kubernetes service account inside the cluster.

Kubernetes task provider

FieldMeaning
kubernetes_provider.namespaceNamespace for task Pods and the LocalQueue.
kubernetes_provider.service_accountOptional service account assigned to task Pods.
kubernetes_provider.host_networkEnables host networking and RDMA requests for GPU Pods.
kubernetes_provider.cache_dirNode-local cache root. CoreWeave configs use /mnt/local/iris-cache.
kubernetes_provider.cache_max_ageMaximum time since the last file write or recorded access before the node-agent reclaims a top-level cache entry. Omit it to disable reclamation.
kubernetes_provider.controller_addressIn-cluster controller address injected into task Pods.
kubernetes_provider.kueue.cluster_queuePulumi-owned ClusterQueue to which Iris binds its LocalQueue. This is required.
kubernetes_provider.kueue.topologiesOptional group_by to CoreWeave node-label mappings.
kubernetes_provider.preempt_namespacesNamespaces containing provider health-check Pods that Iris may clear when they block an admitted GPU job.

An empty kubernetes_provider.kubeconfig means in-cluster authentication. That is the normal CoreWeave controller configuration.

Scale groups and NodePools

Each CoreWeave scale group becomes one NodePool:

FieldEffect
resources.device_variant and resources.device_countAccelerator identity and per-node count advertised to Iris.
resources.cpu, resources.ram, and resources.diskPer-node capacity recorded in the scale-group config.
buffer_slicesMinimum node count for node-based pools.
max_slicesMaximum node count for node-based pools.
slice_template.num_vmsMultiplier applied to the node counts. Current CoreWeave groups use one VM per slice.
slice_template.coreweave.instance_typeCoreWeave node SKU.

CoreWeave Console display names do not always match the Kubernetes spec.instanceType. Use the value accepted by the live NodePool API.

For rack-based NVL72 SKUs, the NodePool uses targetRacks instead of the node autoscaler fields. max_slices * num_vms must be divisible by the 18-node rack size. infra/pulumi/src/iac/nodepools.py and lib/iris/src/iris/cluster/platforms/k8s/nodepool_manifests.py contain the projection rules.

CoreWeave AI Object Storage access

The research clusters set MARIN_PREFIX to CoreWeave object storage. Use it for durable inputs, outputs, and caches. Use marin_temp_bucket for disposable data with a lifecycle deadline. Do not read or copy GCS data from CoreWeave without explicit approval because the transfer incurs egress cost.

The cluster config carries two endpoints for the same S3-compatible store:

CallerConfig fieldCurrent endpoint
Task or controller Pod inside CoreWeaveplatform.coreweave.object_storage_endpointhttp://cwlota.com
Operator process outside CoreWeaveplatform.coreweave.external_object_storage_endpointhttps://cwobject.com

Both domains require virtual-hosted bucket addressing. Let Iris, Rigging, or fsutil derive it; do not build endpoint URLs by hand. For operator-side access, create a CoreWeave object-storage access key and expose only its expected names:

export CW_KEY_ID=<key-id>
export CW_KEY_SECRET=<key-secret>
uv run fsutil buckets

During a controller deployment, Iris maps these values to the S3 variables in iris-task-env. Normal task submissions then receive cluster-managed storage access without carrying the operator's shell environment.

Task working directories and caches are node-local:

Container pathKubernetes volumeLifetime
/app, /tmpemptyDirPod
/uv/cache, /hf/cache, /cargo, /cachehostPath below kubernetes_provider.cache_dirNode
/dev/shmmemory-backed emptyDirPod

Keep cache_dir on /mnt/local, the node's NVMe storage. The shared Hugging Face path is HF_HUB_CACHE; Iris deliberately leaves HF_HOME private because it may contain the submitter's token. The CoreWeave configs set cache_max_age to seven days. Every five minutes, the node-agent removes top-level entries whose files have not been modified or accessed within that age. The sweep ignores directory access times because walking a directory updates them. File access times are best-effort: noatime, O_NOATIME, and some memory-mapped reads do not refresh them. The node-agent atomically renames an expired entry before recursive deletion so another task can refill the original path. Durable outputs belong in object storage.

/cache is unclaimed node-local scratch: a task that needs a real directory on the node instead of a bucket picks its own subdirectory there. The node-agent reclaims those subdirectories under the same policy, so treat anything written there as recoverable. iris.runtime.jax_init uses /cache/xla for XLA's per-fusion autotune results on GPU tasks, because XLA opens that directory from C++ and cannot read an object-store URL. Iris warms one local leader per node from a FineStore file-set snapshot and has global rank 0 publish newly created files in bounded transactions. The remote root is keyed by the launch tree hash under 30-day temporary storage; a storage failure starts cold and does not block JAX initialization. JAX's own compilation cache stays on object storage under the Marin prefix: JAX writes it only from process 0, so a node-local copy would leave every other node permanently cold.

storage.local_state_dir controls controller SQLite storage. When it is empty, Iris creates a controller state PVC. storage.remote_state_dir stores durable controller checkpoints in object storage.

Credentials Summary

Credentials have separate owners and scopes:

LocationScope
~/.kube/coreweave-irisOperator access to CKS. It is not copied into Pods.
Checkout-local .marin.yaml or explicit job environmentSubmitter-provided values such as W&B or Hugging Face credentials.
Operator CW_KEY_ID and CW_KEY_SECRETInput used when deploying the cluster-managed object-storage Secret.
iris-task-env SecretObject-storage credentials and names listed in defaults.inject_env; mounted by the controller and task Pods.
iris-controller-env SecretController-only credentials such as the cluster signing key; task Pods do not mount it.
defaults.task_envNon-secret cluster defaults written into task Pod environments.

The Iris CLI loads .marin.yaml only when run from that checkout. SDK submissions and commands from another directory must pass their task environment explicitly. The deploy path rejects literal secrets before writing iris-cluster-config: references remain in the ConfigMap, while resolved values go through the two Secrets above.

Do not dump Pod environments or use kubectl describe pod on a task Pod; literal job environment values can appear in the output. Inspect named keys and scheduling fields instead, without printing Secret values.

Troubleshooting

Dashboard lists and capacity

Worker and task lists are paginated. Use all pages or the corresponding paginated API when counting workers or tasks. The dashboard command blocks while its port-forward is active; stopping the command closes the tunnel.

Configured or registered accelerator inventory is distinct from free capacity. A healthy controller and a large advertised GB200 inventory do not establish that a requested topology can be admitted now. Confirm current slices, Kueue Workloads, and Pending Pods with the read-only checks below.

Start with Iris's retained task view. It already joins Pod, scheduler, and Kueue state and remains available after Kubernetes garbage collection:

CLUSTER=cw-rno2a
TASK=/<user>/<job>/0

uv run iris --cluster="$CLUSTER" task describe "$TASK"
uv run iris --cluster="$CLUSTER" task events "$TASK"
uv run iris --cluster="$CLUSTER" rpc controller list-backends

task events records ImagePullBackOff and CrashLoopBackOff warnings. Pair it with iris job logs /<user>/<job> for task output. An Evicted task can indicate ephemeral-storage pressure. Raise the task's disk request only when its /app, /tmp, or writable container layer is full; escalate node DiskPressure instead of editing the Pod or NodePool.

For a Pending task, use these read-only Kubernetes checks when the Iris message is not enough. Read the kubeconfig, context, and namespace from the cluster config rather than copying the example values:

KUBECONFIG=~/.kube/coreweave-iris
CONTEXT=<platform.coreweave.kube_context>
NAMESPACE=<kubernetes_provider.namespace>

kubectl --kubeconfig "$KUBECONFIG" --context "$CONTEXT" get nodepool -o wide
kubectl --kubeconfig "$KUBECONFIG" --context "$CONTEXT" -n "$NAMESPACE" \
  get pods,workloads.kueue.x-k8s.io -o wide
kubectl --kubeconfig "$KUBECONFIG" --context "$CONTEXT" -n "$NAMESPACE" \
  get workload <workload-name> \
  -o jsonpath='{range .status.conditions[*]}{.type}{"\t"}{.status}{"\t"}{.reason}{"\t"}{.message}{"\n"}{end}'

SchedulingGated means Kueue has not admitted the Workload. QuotaReserved=False explains quota or topology-fit failures. A NodePool with Valid=False has a rejected configuration; a target/current mismatch usually means nodes are still provisioning or unhealthy. Do not work around either by editing the live objects. Pulumi owns NodePools and cluster-scoped Kueue resources.

For actor-level Kubernetes API history after ordinary Events expire, use the Kubernetes Audit Logs dashboard in CoreWeave Observe. iris task events is the Iris-side record of backend observations and controller decisions; it is not an API audit log.

For context errors, return to Connecting. For public federation timeouts, controller restarts, kernel-deadlock recovery, or other live faults, use CoreWeave GPU Operations. Kubernetes apply, delete, scale, drain, cordon, and uncordon require explicit operator approval.

Onboarding and source routing

For a new CoreWeave cluster, follow the ownership boundary in order:

  1. Obtain the CKS cluster, operator kubeconfig, object-storage bucket, and access keys. CKS and object storage are not Pulumi-managed today.
  2. Add lib/iris/config/<cluster>.yaml from live cluster facts and the schema in config.py. Do not copy fleet sizes from a document.
  3. Provision the controller signing key with iris cluster init-keys and reference its private half from auth.signing_key. Add each parent controller's public key under auth.federation_peers so the CoreWeave controller can verify handoffs. Signing keys stay out of Pulumi state. See federation credentials.
  4. Follow the Pulumi guide to create or adopt RBAC, NodePools, Kueue, Traefik, TLS, and DNS. Stop on any NodePool replacement or deletion in the preview.
  5. If the cluster names a Finelog config, follow the Finelog operations guide for its separate forwarding key and deploy that service before the Iris controller.
  6. Register the peer in the parent federation config and deploy the controller through the use-iris skill.
  7. Verify cluster status, list-backends, the public federation route, and one representative topology-aware smoke before sending normal jobs.

Use this page for CoreWeave architecture and config, the federation reference for cross-cluster job semantics, the Pulumi guide for infrastructure, and OPS.md for live diagnosis.

Source map