FlagQuantum Capabilities

September 22, 2026 · View on GitHub

Choose a supported workflow by user goal, runtime, hardware, and evidence level. This catalog is generated from the machine-validated capability-maturity.toml source of truth.

Maturity applies only to the scope stated in each row. A local, replicated, sliced, or planned execution path is not distributed scalability evidence.

How to read maturity

LevelMeaning
Release certifiedRelease-gated with audited, reproducible evidence and no unresolved release blocker.
Production supportedSupported path with compatibility, operational guidance, and target-hardware evidence.
Development evidenceExecutable and tested development result; not a production or general scalability claim.
ExperimentalResearch surface without compatibility or production guarantees.

Find a capability by goal

GoalCapabilityMaturityStart here
Build portable quantum circuitsUnified circuit API and FlagQuantum IRRelease certifiedRun example
Compile or export a circuitUnified circuit API and FlagQuantum IRRelease certifiedRun example
Inspect a stable circuit representationUnified circuit API and FlagQuantum IRRelease certifiedRun example
Simulate a small or medium circuit exactlyLocal statevector simulation and trainingProduction supportedRun example
Train a parameterized quantum circuitLocal statevector simulation and trainingProduction supportedRun example
Run local VQE and quantum machine learningLocal statevector simulation and trainingProduction supportedRun example
Train one statevector workload across multiple ranksSharded statevector trainingProduction supportedRun example
Plan distributed statevector ownershipSharded statevector trainingProduction supportedRun example
Inspect communication and sharding semanticsSharded statevector trainingProduction supportedRun example
Check FlagQuantum and Torch-FL integrationFlagOS local statevector CUDA referenceDevelopment evidenceRun example
Audit the statevector operator profileFlagOS local statevector CUDA referenceDevelopment evidenceRun example
Compare complex numerical behavior with a CPU complex128 referenceFlagOS local statevector CUDA referenceDevelopment evidenceRun example
Audit FlagOS distributed statevector workload supportFlagOS distributed statevector workloadsDevelopment evidenceRun example
Run sharded FlagOS statevector forward workloadsFlagOS distributed statevector workloadsDevelopment evidenceRun example
Run bounded sharded FlagOS training trajectoriesFlagOS distributed statevector workloadsDevelopment evidenceRun example
Audit matched FlagOS statevector capacity expansionFlagOS statevector capacity expansionDevelopment evidenceRun example
Run one full-width complex128 statevector across eight ranksFlagOS statevector capacity expansionDevelopment evidenceRun example
Inspect fail-closed single, replicated, and sharded capacity evidenceFlagOS statevector capacity expansionDevelopment evidenceRun example
Audit observable FlagOS collective behaviorFlagOS distributed transport observabilityDevelopment evidenceRun example
Inspect bounded profiler evidence for complex collectivesFlagOS distributed transport observabilityDevelopment evidenceRun example
Distinguish FlagOS validation from unverified FlagCX route attributionFlagOS distributed transport observabilityDevelopment evidenceRun example
Train a large low-entanglement systemDifferentiable and sharded MPS trainingDevelopment evidenceRun example
Distribute one MPS across several GPUsDifferentiable and sharded MPS trainingDevelopment evidenceRun example
Inspect variable-bond MPS capacityDifferentiable and sharded MPS trainingDevelopment evidenceRun example
Evaluate software-extended precision on FP32 hardwareDouble-Single FP32 numerical primitivesExperimentalRun example
Measure cancellation error against a float64 referenceDouble-Single FP32 numerical primitivesExperimentalRun example
Prepare a precision provider without changing the runtimeDouble-Single FP32 numerical primitivesExperimentalRun example
Evaluate statevector forward execution on FP32-only PyTorch devicesSplit real/imag FP32 local statevector P0ExperimentalRun example
Validate Torch-FL flagos logical-device residencySplit real/imag FP32 local statevector P0ExperimentalRun example
Compare split FP32 numerical error with a CPU complex128 referenceSplit real/imag FP32 local statevector P0ExperimentalRun example
Evaluate Pauli expectation values on FP32-only PyTorch devicesSplit real/imag FP32 observable and parameter-shift P1ExperimentalRun example
Compute explicit parameter-shift gradients for named rotation parametersSplit real/imag FP32 observable and parameter-shift P1ExperimentalRun example
Compare training primitives with a CPU complex128 referenceSplit real/imag FP32 observable and parameter-shift P1ExperimentalRun example
Survive cancellation-sensitive Hamiltonian reductions on FP32-only devicesSelective Double-Single split statevector precision P2ExperimentalRun example
Request an auditable selective precision planSelective Double-Single split statevector precision P2ExperimentalRun example
Fail closed when requested accuracy exceeds certified evidenceSelective Double-Single split statevector precision P2ExperimentalRun example
Evaluate full Double-Single state evolution on FP32-only PyTorch devicesFull Double-Single split statevector P3ExperimentalRun example
Break the FP32 state-evolution error floorFull Double-Single split statevector P3ExperimentalRun example
Measure state, norm, expectation, and gradient errors against complex128Full Double-Single split statevector P3ExperimentalRun example
Avoid CPU complex128 gate encoding on FP32-only PyTorch devicesDevice-generated Double-Single split statevector P4ExperimentalRun example
Evaluate bounded Double-Single trigonometry and state evolutionDevice-generated Double-Single split statevector P4ExperimentalRun example
Measure state, norm, expectation, and gradient errors against complex128 diagnosticsDevice-generated Double-Single split statevector P4ExperimentalRun example
Train through a conventional PyTorch scalar loss on CPUCPU PyTorch autograd bridge over Double-Single P5ExperimentalRun example
Retain Double-Single arithmetic inside forward and parameter-shift evaluationCPU PyTorch autograd bridge over Double-Single P5ExperimentalRun example
Audit the exact boundary where PyTorch receives one FP32 gradient wordCPU PyTorch autograd bridge over Double-Single P5ExperimentalRun example
Preserve parameter updates below one FP32 ULPSingle-device precision-preserving Double-Single SGD P5ExperimentalRun example
Compare Double-Single and FP32-master optimization against complex128 trajectoriesSingle-device precision-preserving Double-Single SGD P5ExperimentalRun example
Run auditable precision SGD without claiming torch.optim compatibilitySingle-device precision-preserving Double-Single SGD P5ExperimentalRun example
Reproduce small open-chain imaginary-time TEBD studiesConstrained local MPS TEBDExperimentalRun example
Cross-check MPS evolution against an independent oracleConstrained local MPS TEBDExperimentalRun example
Inspect truncation and normalization evidenceConstrained local MPS TEBDExperimentalRun example
Evaluate a circuit with tensor-network contractionTensor-network execution and trainingExperimentalRun example
Compare statevector, MPS, and tensor-network modesTensor-network execution and trainingExperimentalRun example
Validate small noisy circuits exactlyExact and trajectory-based noisy simulationExperimentalRun example
Evaluate low-entanglement noisy circuits with MPS trajectoriesExact and trajectory-based noisy simulationExperimentalRun example
Resume reproducible trajectory ensemblesExact and trajectory-based noisy simulationExperimentalRun example
Package a trained parameterized circuitCircuit packaging and cloud deploymentDevelopment evidenceRun example
Export a circuit for a providerCircuit packaging and cloud deploymentDevelopment evidenceRun example
Run a circuit through a deployment abstractionCircuit packaging and cloud deploymentDevelopment evidenceRun example
Predict a mapped QPU measurement distributionEvidence-qualified QPU digital twinsDevelopment evidenceRun example
Compare a frozen prediction with later QPU countsEvidence-qualified QPU digital twinsDevelopment evidenceRun example
Reject circuits outside validated topology and depthEvidence-qualified QPU digital twinsDevelopment evidenceRun example
Compose compatible local Twin cells into a connected regional prediction modelEvidence-qualified QPU digital twinsDevelopment evidenceRun example
Prospectively validate one exact regional circuit with repeated QPU tasksEvidence-qualified QPU digital twinsDevelopment evidenceRun example
Validate a fixed regional circuit suite under one simultaneous confidence levelEvidence-qualified QPU digital twinsDevelopment evidenceRun example
Compare predeclared reference and holdout regional circuitsEvidence-qualified QPU digital twinsDevelopment evidenceRun example
Track the same fixed regional suite across calibration snapshotsEvidence-qualified QPU digital twinsDevelopment evidenceRun example
Discover registered interoperability adaptersInteroperability adapter contractExperimentalRun example
Implement a framework adapter without changing FlagQuantum coreInteroperability adapter contractExperimentalRun example
Certify round-trip and fail-closed adapter behaviorInteroperability adapter contractExperimentalRun example
Handle conversion diagnostics consistentlyInteroperability adapter contractExperimentalRun example
Import a supported PennyLane QuantumScriptPennyLane QuantumScript interoperabilityExperimentalRun example
Export static FlagQuantum IR to PennyLanePennyLane QuantumScript interoperabilityExperimentalRun example
Audit semantic loss at the boundaryPennyLane QuantumScript interoperabilityExperimentalRun example
Import a supported Qiskit circuitQiskit IR interoperabilityExperimentalRun example
Export FlagQuantum IR to QiskitQiskit IR interoperabilityExperimentalRun example
Audit semantic loss at a framework boundaryQiskit IR interoperabilityExperimentalRun example
Prototype mid-circuit measurement and feed-forwardDynamic circuits and backend assessmentExperimentalRun example
Assess backend support before executionDynamic circuits and backend assessmentExperimentalRun example
Exercise a fixed-round QEC control workflowRepetition-code memory experimentDevelopment evidenceRun example
Inspect syndrome and detection-event recordsRepetition-code memory experimentDevelopment evidenceRun example
Prototype a decoder against a typed contractRepetition-code memory experimentDevelopment evidenceRun example
Prototype a FlagQuantum extensionExtension SDKExperimentalRun example
Register custom framework behaviorExtension SDKExperimentalRun example
Install an external circuit compilerExtension SDKExperimentalRun example
Convert a QUBO problem into a HamiltonianQUBO to Ising mappingExperimentalRun example
Recover the QUBO form from a HamiltonianQUBO to Ising mappingExperimentalRun example
Evaluate a QUBO objective on a candidate assignmentQUBO to Ising mappingExperimentalRun example
Prepare a uniform superpositionQuantum state preparationExperimentalRun example
Prepare a state matching a given amplitude vectorQuantum state preparationExperimentalRun example
Cross-check a prepared state against the target amplitudesQuantum state preparationExperimentalRun example
Flip a target qubit only when every control is setOracle building blocksExperimentalRun example
Mark the states where one bit string is greater than anotherOracle building blocksExperimentalRun example
Compose reversible classical logic into a circuitOracle building blocksExperimentalRun example
Mark the states satisfying a predicate with a phaseTruth-table oracle synthesisExperimentalRun example
Write a predicate's value onto an output qubitTruth-table oracle synthesisExperimentalRun example
List the states a predicate marksTruth-table oracle synthesisExperimentalRun example
Search for the states a predicate marksGrover searchExperimentalRun example
Amplify the marked states' amplitudesGrover searchExperimentalRun example
Read the most likely marked state from samplesGrover searchExperimentalRun example
Estimate the amplitude a marking operator selectsQuantum amplitude estimationExperimentalRun example
Read an amplitude off the counting registerQuantum amplitude estimationExperimentalRun example
Compare an estimate against a known amplitudeQuantum amplitude estimationExperimentalRun example
Estimate the spectrum of a density matrixQuantum principal component analysisExperimentalRun example
Read an eigenvalue off a counting registerQuantum principal component analysisExperimentalRun example
Compare a read-out eigenvalue against the density matrix's own spectrumQuantum principal component analysisExperimentalRun example
Assign each point to its nearest centroidQuantum k-mediansExperimentalRun example
Read an assignment off sampled searchesQuantum k-mediansExperimentalRun example
Update centroid positions to the medians of the points assigned to themQuantum k-mediansExperimentalRun example
Estimate the kernel matrix of a set of feature vectorsQuantum kernel estimation and kernel ridge classificationExperimentalRun example
Read a kernel entry off sampled swap testsQuantum kernel estimation and kernel ridge classificationExperimentalRun example
Classify held-out rows with a kernel ridge classifierQuantum kernel estimation and kernel ridge classificationExperimentalRun example
Build the binary objective of a feature-selection instanceFeature selection as a QUBOExperimentalRun example
Read the objective's value off an assignmentFeature selection as a QUBOExperimentalRun example
Map the objective to an Ising Hamiltonian for a solver to consumeFeature selection as a QUBOExperimentalRun example
Estimate the fraction of a database's items whose support meets a thresholdFrequent-item fractions by amplitude estimationExperimentalRun example
Read that fraction off an amplitude estimation counting registerFrequent-item fractions by amplitude estimationExperimentalRun example
Compare the estimate against an enumerated frequent fractionFrequent-item fractions by amplitude estimationExperimentalRun example
Estimate a matrix's singular values from a counting registerSingular values by phase estimation over the Hermitian embeddingExperimentalRun example
Read one singular value off the register's modeSingular values by phase estimation over the Hermitian embeddingExperimentalRun example
Compare a read-out value against the matrix's own decompositionSingular values by phase estimation over the Hermitian embeddingExperimentalRun example

Build and compile

Unified circuit API and FlagQuantum IR

Build, validate, serialize, compile, and inspect quantum circuits through the stable FlagQuantum interface.

  • Maturity: Release certified
  • Public API: fq.Circuit, fq.CircuitIR, flagquantum.compiler.compile
  • Runtime modes: not_applicable
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: not_applicable
  • Start: quick example
  • Documentation: guide
  • Known boundary: IR v1; incompatible schema changes require an explicit migration.

Simulation and training

Local statevector simulation and training

Run exact circuits and differentiable quantum workloads on a CPU or one GPU.

  • Maturity: Production supported
  • Public API: fq.Circuit, fq.run, fq.Module, fq.train
  • Runtime modes: statevector
  • Hardware: cpu, single_gpu
  • Gradient support: exact
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Capacity is bounded by one device; distributed capacity claims use the sharded capability.

FlagOS local statevector CUDA reference

Exercise the local differentiable statevector path through Torch-FL's logical flagos device on a locked CUDA reference environment.

  • Maturity: Development evidence
  • Public API: flagquantum.runtime.resolve_device, fq.run
  • Runtime modes: statevector
  • Hardware: nvidia_a100_cuda_reference
  • Gradient support: development_evidence
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: CUDA-backed development reference only. It does not certify a domestic accelerator, prove absence of Torch-FL host fallback, establish production performance, or authorize a scalability claim.

Double-Single FP32 numerical primitives

Use residual-preserving pairs of float32 tensors for bounded real and split-complex arithmetic experiments on PyTorch devices.

  • Maturity: Experimental
  • Public API: Not exposed; internal development evidence only
  • Runtime modes: numerical_primitive_conformance
  • Hardware: cpu, device_generic_pytorch
  • Gradient support: experimental_composed_primitives
  • Distribution semantics: single_device_primitive_only
  • Start: quick example
  • Documentation: guide
  • Known boundary: Pure FP32 real and split-complex eager primitives, selective split-statevector P2 reductions, full-state P3, and bounded device-generated-gate P4 experiments are available. P4 removes CPU float64/complex128 gate encoding for its certified gate and angle scope, but optimized kernels, compiled execution, distributed collectives, decompositions, optimizer state, provider-owned Torch-FL route auditing, domestic accelerators, performance, convergence, and production use remain uncertified. Double-Single retains FP32 exponent range and is not generally equivalent to FP64 or complex128.

Split real/imag FP32 local statevector P0

Execute a bounded forward-only statevector using two device-resident float32 tensors without requiring accelerator complex dtypes.

  • Maturity: Experimental
  • Public API: fq.experimental.numerics.execute_split_real_imag_statevector
  • Runtime modes: split_real_imag_statevector_p0
  • Hardware: cpu, single_cuda_reference, device_generic_pytorch
  • Gradient support: unsupported
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Explicit experimental forward-only executor for a bounded built-in gate set using separate FP32 real and imaginary tensors. It is not selected by the default runtime. Custom matrices, gradients, optimizer steps, sampling and observables APIs, compiled execution, distributed execution, Double-Single storage, provider-internal route auditing, domestic-hardware certification, performance, and production use remain unsupported. CUDA or CUDA-backed flagos evidence is portability evidence only.

Split real/imag FP32 observable and parameter-shift P1

Evaluate bounded Pauli Hamiltonians and explicit parameter-shift gradients using two device-resident float32 state tensors without accelerator complex dtypes.

  • Maturity: Experimental
  • Public API: fq.experimental.numerics.execute_split_real_imag_expectation, fq.experimental.numerics.parameter_shift_split_real_imag_gradient
  • Runtime modes: split_real_imag_statevector_p1
  • Hardware: cpu, single_cuda_reference, device_generic_pytorch
  • Gradient support: parameter_shift_experimental
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Explicit experimental batch-one Pauli expectation and occurrence-wise two-term parameter-shift gradients for direct named scalar parameters on RX, RY, RZ, RXX, RYY, and RZZ. It is not native autograd and provides no optimizer integration. Parameter expressions, trainable coefficients, custom matrices or states, sampling, compilation, distributed execution, automatic runtime selection, performance, convergence, provider-internal route auditing, and domestic-hardware certification remain unsupported. CUDA or CUDA-backed flagos evidence is portability evidence only.

Selective Double-Single split statevector precision P2

Retain FP32 state and gates while upgrading Pauli inner products, Hamiltonian sums, and parameter-shift accumulation to explicit Double-Single high/low reductions.

  • Maturity: Experimental
  • Public API: Not exposed; internal development evidence only
  • Runtime modes: split_real_imag_statevector_p2_precision
  • Hardware: cpu, single_cuda_reference, device_generic_pytorch
  • Gradient support: parameter_shift_selective_double_single_experimental
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Explicit experimental selective precision path only. State storage, gate generation, and gate application remain split FP32; only Pauli inner products, Hamiltonian term sums, and parameter-shift accumulation retain Double-Single high/low words. Full Double-Single statevectors, native autograd, optimizer integration, residual checkpoints, decomposition, compilation, distributed execution, automatic selection, convergence certification, provider-internal route auditing, domestic-hardware certification, performance, and production use remain unsupported.

Full Double-Single split statevector P3

Store every complex amplitude as four FP32 high/low words and retain residuals through gate application, periodic normalization, observables, and parameter-shift gradients.

  • Maturity: Experimental
  • Public API: Not exposed; internal development evidence only
  • Runtime modes: split_real_imag_statevector_p3_double_single
  • Hardware: cpu, single_cuda_reference, device_generic_pytorch
  • Gradient support: parameter_shift_full_double_single_experimental
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Explicit correctness-first full Double-Single state experiment. CPU float64/complex128 parameter and gate encoding is required before four FP32 words are transferred to the execution device; state evolution itself has no host fallback. Device-only Double-Single trigonometry, optimized/fused kernels, native autograd, optimizer integration, compilation, distributed execution, automatic selection, algorithmic convergence certification, provider-internal route auditing, domestic-hardware certification, performance, and production use remain unsupported.

Device-generated Double-Single split statevector P4

Generate bounded fixed and rotation gates with device-resident FP32 Double-Single arithmetic, then retain four FP32 words through state evolution, observables, and parameter-shift gradients.

  • Maturity: Experimental
  • Public API: Not exposed; internal development evidence only
  • Runtime modes: split_real_imag_statevector_p4_device_double_single
  • Hardware: cpu, single_cuda_reference, device_generic_pytorch
  • Gradient support: parameter_shift_device_double_single_experimental
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Explicit correctness-first P4 path for built-in gates and direct scalar parameters within |angle| <= 1024. Python scalars are host-ingested as FP32 values; device-resident FP32 or Double-Single parameters remain on device. Parameter expressions, float64 parameter tensors, custom matrices, unbounded angles, native autograd, optimizer integration, compilation, distributed execution, automatic selection, algorithmic convergence certification, provider-internal route auditing, domestic-hardware certification, performance, and production use remain unsupported.

CPU PyTorch autograd bridge over Double-Single P5

Expose a bounded first-order PyTorch autograd path whose forward and parameter-shift backward use P4 Double-Single arithmetic before an explicit FP32 tensor delivery boundary.

  • Maturity: Experimental
  • Public API: Not exposed; internal development evidence only
  • Runtime modes: split_real_imag_statevector_p5_autograd_bridge
  • Hardware: cpu
  • Gradient support: first_order_parameter_shift_internal_double_single_float32_delivery_experimental
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Explicit CPU-only first-order PyTorch autograd bridge for direct named scalar FP32 parameters and bounded Pauli expectations. Forward and backward use P4 Double-Single arithmetic internally, but the scalar loss and Tensor.grad are one-word FP32 delivery boundaries and no end-to-end Double-Single gradient claim is made. A separate explicit Double-Single SGD lane is available; higher-order and compiled autograd, CUDA, Torch-FL flagos, FlagCX, distributed execution, automatic selection, convergence, hardware certification, performance, and production use remain unsupported.

Single-device precision-preserving Double-Single SGD P5

Consume explicit P4 high/low parameter-shift gradients and return updated high/low master parameters without crossing the one-word PyTorch Tensor.grad boundary.

  • Maturity: Experimental
  • Public API: Not exposed; internal development evidence only
  • Runtime modes: split_real_imag_statevector_p5_double_single_sgd
  • Hardware: cpu, single_cuda_reference, torch_fl_flagos_cuda_reference
  • Gradient support: explicit_parameter_shift_double_single_optimizer_experimental
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Explicit single-device functional high/low master-parameter SGD using P4 high/low parameter-shift gradients. CPU, native CUDA on A800, and CUDA-backed Torch-FL flagos:0 portability trajectories are recorded; the A800 routes do not certify a domestic accelerator or provider internals. It is not torch.optim compatible and has no momentum, weight decay, loss scaling, Adam-family algorithm, checkpoint/state-dict compatibility, higher-order autograd, FlagCX, distributed execution, automatic selection, convergence certification, hardware certification, performance, or production claim.

Constrained local MPS TEBD

Evolve open-chain local Pauli Hamiltonians with fail-closed second-order imaginary-time TEBD.

  • Maturity: Experimental
  • Public API: fq.experimental.simulation.run_tebd
  • Runtime modes: mps_tebd
  • Hardware: cpu, single_gpu
  • Gradient support: unsupported
  • Distribution semantics: single_device_fast_path
  • Start: quick example
  • Documentation: guide
  • Known boundary: Static real one-site and adjacent two-site Pauli terms on an open chain, batch one, second-order imaginary-time evolution, and product initial states only. Real-time evolution, periodic and nonlocal terms, gradients, TDVP, distributed execution, and production or scalability claims are unsupported.

Tensor-network execution and training

Execute tensor-network circuit paths and evaluate experimental contraction and gradient workflows.

  • Maturity: Experimental
  • Public API: flagquantum.simulation.tensor_network.run_tensor_network
  • Runtime modes: tensor_network
  • Hardware: cpu, single_gpu
  • Gradient support: experimental
  • Distribution semantics: manual_sliced_tensor_contraction
  • Start: quick example
  • Documentation: guide
  • Known boundary: General reverse contraction and production distributed transport are not certified.

Exact and trajectory-based noisy simulation

Lower validated Kraus noise models into FlagQuantum IR and execute exact density-matrix or MPS quantum-trajectory paths.

  • Maturity: Experimental
  • Public API: flagquantum.noise.NoiseModel, flagquantum.noise.noisy_density_matrix, flagquantum.runtime.run_noisy_mps, flagquantum.runtime.run_noisy_statevector
  • Runtime modes: density_matrix, noisy_mps
  • Hardware: cpu, single_gpu
  • Gradient support: unsupported
  • Distribution semantics: single_device_fast_path_or_rank_local_trajectory_partition
  • Start: quick example
  • Documentation: guide
  • Known boundary: Validated Markovian Kraus channels, timestamped DeviceNoiseProfile input, ASAP gate/idle thermal lowering, classical readout confusion, exact density execution, and reproducible MPS trajectories with single-rank adaptive stopping are available. Pulse overlap, crosstalk, leakage, provider calibration adapters, distributed adaptive stopping, batched statevector trajectories, production multi-GPU scheduling, and noisy gradients remain unsupported. Multi-wire MPS channels use an explicitly dense correctness fallback.

Repetition-code memory experiment

Run a bounded three-data-qubit memory experiment with timed errors or circuit-location bit-flip/readout noise, compiled feedback, per-round Runtime decoding, or Pauli-frame correction, including a two-round temporal reference decoder.

  • Maturity: Development evidence
  • Public API: flagquantum.qec.run_repetition_memory_experiment, flagquantum.qec.run_repetition_memory_noise_sweep, flagquantum.qec.RepetitionNoiseProfile, flagquantum.qec.ErrorSchedule, flagquantum.qec.Decoder, flagquantum.qec.StreamingDecoder, flagquantum.qec.RepetitionTemporalDecoder
  • Runtime modes: local_statevector_compiled_feedback, local_statevector_runtime_decoder, local_statevector_runtime_pauli_frame, local_statevector_runtime_temporal_decoder, local_statevector_runtime_temporal_pauli_frame, local_statevector_offline_pauli_frame, local_statevector_noisy_trajectory
  • Hardware: cpu
  • Gradient support: unsupported
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: A synchronous local reference for one fixed three-data-qubit repetition-code profile. It supports bounded deterministic X-error schedules, replaceable per-round trajectory decoding with physical-X or Pauli-frame-X actions, and a two-round temporal rule that rejects an isolated readout excursion. Confirmable data errors require a following round; terminal-round onsets remain unconfirmed. The circuit-location stochastic profile contains independent bit flips after parity-check CNOTs and independent syndrome/final-readout confusion. The middle data wire has two CNOT noise opportunities per round while edge wires have one. Feedback traces separate true and observed bits, actions, and frame evolution. Sweeps report finite-shot observations only, not logical suppression or thresholds. The temporal rule is not maximum-likelihood decoding and repeated readout faults may mimic data errors. Batched decoder feedback, general channels/codes, correlated or timing noise, hard-real-time/provider control, gradients, distributed execution, capacity, performance, and fault-tolerance claims remain unsupported. The feedback records are private subinterfaces and the namespace is not exported from the stable package root.

QUBO to Ising mapping

Express a quadratic unconstrained binary optimization problem as an Ising Hamiltonian for the existing variational workflows.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.qubo
  • Runtime modes: not_applicable
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: not_applicable
  • Start: quick example
  • Documentation: guide
  • Known boundary: A polynomial classical transformation with no advantage of its own: any advantage a caller observes belongs to the solver that consumes the Hamiltonian. Only Z-basis objectives are representable, so a Hamiltonian outside the Z basis is rejected. The mapping carries the constant as an identity term and recovers it on the way back, so a caller comparing the two forms sees identical values; a caller who strips the identity term loses that constant. It certifies no solver, convergence, performance, or hardware behavior.

Quantum state preparation

Prepare a uniform or arbitrary quantum state from a classical amplitude vector using uniformly controlled rotations.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.primitives.state_preparation
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: Demonstration scale. Computing the rotation angles requires a classical pass over all 2n amplitudes and a 2n by 2**n linear solve, so the input is already exponential in size: this unit shows that a state can be prepared efficiently given its amplitudes, not that preparing a state is cheaper than its classical description. It makes no advantage, performance, convergence, or hardware claim, and is not selected by the default runtime. The preparation is exact only to the working precision of the solve.

Oracle building blocks

Multi-controlled X and a reversible bit-string comparator: the reversible classical logic the oracle units are built from.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.primitives.oracle.append_multi_controlled_x, flagquantum.algorithms.primitives.oracle.append_comparator
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: Reversible classical logic of O(n) Toffoli-style cost with no advantage premise of its own. A multi-controlled X above two controls needs len(controls) - 2 caller-supplied ancillas, each of which must be in |0> on entry: measured, a dirty ancilla makes the target silently wrong on a large fraction of inputs (8 of 16 at three controls, 32 of 96 at four) while never corrupting the ancilla itself, so the failure is invisible from the ancilla. The comparator restores every wire it is given except the target; its len(lhs) + 1 equality flags and its scratch wire must all be in |0> on entry too, and a dirty equality[0] makes the target silently wrong on a large fraction of the operand patterns (6 of 16 at two bits) before the ladder restores the flag. The comparator's published record is semi-verified: its venue is not indexed by Crossref, DBLP or INSPIRE, so its volume and page numbers are reported by citing works rather than index-confirmed. It makes no advantage, performance, or hardware claim.

Truth-table oracle synthesis

Turn a classical predicate into a phase oracle or a bit oracle by enumerating its truth table.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.primitives.oracle.marked_states, flagquantum.algorithms.primitives.oracle.phase_oracle, flagquantum.algorithms.primitives.oracle.append_phase_oracle, flagquantum.algorithms.primitives.oracle.bit_oracle, flagquantum.algorithms.primitives.oracle.append_bit_oracle
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: Synthesis enumerates all 2**n inputs of the truth table classically, so it carries no advantage of its own at any scale beyond demonstration and its cost is exponential in the register width. Only a truth table is accepted: there is no boolean-expression parser and no other predicate form. A phase oracle is capped at three wires, because a multi-controlled Z above that needs ladder ancillas a standalone circuit does not have; the in-place append form takes them from the caller. The bit oracle's output is XORed rather than assigned, and it restores every wire it allocates. No performance, convergence, or hardware claim is made.

Search a classical predicate's truth table by amplitude amplification over an evaluation register.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.grover
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: The improvement is in query complexity, against an oracle this unit synthesizes from a truth table at a classical cost of 2**n. No end-to-end advantage follows at any scale beyond demonstration, and no qRAM, block encoding or amplitude encoding is assumed. The search is bounded at three evaluation wires, because a phase oracle above that needs ladder ancillas a register of exactly that width does not have. The register width makes the classical truth-table enumeration exponential, which is the honest limit of the unit. It makes no performance, convergence, or hardware claim.

Quantum amplitude estimation

Estimate the probability a state-preparation unitary's marked subspace carries, by phase estimation over the Grover operator.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.amplitude_estimation
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: The quadratic speedup over classical sampling is real only given that the state-preparation unitary A is free: a real distribution needs QRAM, and this unit does not supply one, so it does not show that any Monte Carlo integral is estimated faster than classically. The estimate lies on the amplitude grid sin^2(pi j / 2**(m+1)) and is accurate to about one grid step; that is a resolution, not a coverage-calibrated confidence interval, and no confidence interval is reported. The operator must be able to apply A, its adjoint, the marking operator and the zero-state reflection each under control, which excludes state-preparation circuits that cannot be controlled. It makes no performance, convergence, or hardware claim.

Quantum principal component analysis

Estimate the eigenvalues of a data matrix's density matrix by phase estimation over its exponential, reading them off a counting register.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.pca
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: Demonstration scale. The density matrix is materialized classically and its exponential is built with a dense matrix exponential, so the unit does not reproduce quantum PCA's input model: it never avoids forming rho and never uses the O(1/eps^3) state copies the algorithm is built on. The purified input is supplied as all 2**n amplitudes. Meaningful only for low effective rank.

Quantum k-medians

Assign points to their nearest centroids with a Grover-style minimum search over a centroid index register, then move each centroid to the classical coordinate-wise median of its cluster.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.kmedians
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: The advantage premise is Grover's oracle model, and it is not met: the oracle is not free here. The search's oracle is synthesized from the predicate's truth table at O(2**n) cost, so no end-to-end advantage follows at this scale, and the distance table the predicate compares is computed classically, one point at a time, before any circuit is built -- the register is capped at three wires, which is also what keeps the distances out of it. Nothing here reads a qRAM or runs an adiabatic evolution, so no conclusion that rests on either applies to this unit. Demonstration scale: the search is bounded at three evaluation wires, so at most eight centroids, and the assignment is sampled rather than read out, which is why a small sample can stop a point's search short. It makes no performance, convergence, or hardware claim.

Quantum kernel estimation and kernel ridge classification

Estimate a kernel matrix by swap test over an angle-encoded feature map, and fit a classical kernel ridge classifier on the estimated entries.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.quantum_kernel
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: The advantage premise is the data-access model, and this unit does not meet it. The kernel-matrix circuit's cost statement assumes the two feature states are available, reached through a qRAM or an amplitude-encoding unitary whose cost the estimate does not count; here each feature state is built gate by gate from the classical feature vector on every run, so that cost is paid rather than assumed away and no end-to-end advantage follows. The sampling cost is the paper's own -- O(eps**-2) shots per kernel entry and O(m2 / eps2) for an m by m kernel matrix -- and no error bound, confidence interval or repetition scheme is computed or reported anywhere in this unit. Every entry is a sampled estimate, so the classifier's coefficients and its predictions inherit the sample, and an estimate of a near-zero overlap can come back slightly negative because the readout is 1 - 2 * share and is not clamped. The classical hardness of estimating these kernel entries is a conjecture in Havlicek et al. and not a theorem, and the rigorous speed-up results for quantum kernel methods require a fault-tolerant quantum computer (Liu, Arunachalam and Temme 2021). Demonstration scale: the feature map is a three-feature angle encoding of this package's own, and the classifier is classical kernel ridge regression whose only quantum part is the kernel. It makes no performance, convergence, or hardware claim.

Feature selection as a QUBO

Build the binary objective of a feature-selection instance -- a subset's relevance and pairwise redundancy, scored with a penalty on the subset's size -- and evaluate or map it for a solver.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.feature_selection
  • Runtime modes: not_applicable
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: not_applicable
  • Start: quick example
  • Documentation: guide
  • Known boundary: This unit does not solve, and the repository has no annealer: it builds the objective of one feature-selection instance and evaluates that objective at an assignment the caller supplies. Which subset a solver returns, and at what cost, belongs to the solver the problem is handed to, so any advantage such a solver observes is the solver's and this construction carries none of its own -- the QUBO form and the Ising form are a polynomial classical transformation with no advantage of their own. Both scores are the caller's data: this unit defines no relevance measure and no redundancy measure and puts no interpretation on either. The penalty weight is the caller's too, with no default here and no weight at which the target size starts to bind computed or predicted, so a weight small enough against the scores can leave a subset of another size cheapest. Demonstration scale: the instance is a set of feature scores the caller brings, and the objective is an ordinary quadratic binary form whose quadratic terms are the pairwise scores folded together with the size penalty. It certifies no solver, convergence, performance, or hardware behavior, and it selects no runtime.

Frequent-item fractions by amplitude estimation

Estimate the fraction of a binary incidence matrix's items whose support meets a threshold, with a support register the circuit fills one controlled increment per transaction-item membership.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.qarm
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: The advantage premise is coherent entry-wise database access, and this unit does not meet it. The cited paper's speed-up is measured in calls to an oracle that returns one database entry per call, and it reaches the candidate itemset superpositions it prepares through a qRAM; neither is exercised here: the transactions are iterated over classically, one controlled increment of the support register per transaction-item membership, emitted from the incidence matrix as Python reads it, so the access cost is paid explicitly rather than assumed away and no end-to-end advantage follows. The improvement the paper claims is quadratic in the number of database queries and is stated conditionally, for the case M_f^(k) << M_c^(k); it is not exponential, and that wording appears only in an earlier arXiv listing of the same work. The support register must be wide enough to hold the largest support any database of that transaction count could produce: the increment is a permutation of the register's own values, so a narrower register wraps a support into another value and the readout can come back wrong with nothing raised, and a width that cannot hold every support is therefore refused rather than left to wrap. The item register is addressed by one wire per item-index bit, so the item count is a power of two. Demonstration scale: at most eight items and seven transactions, and the marking operator enumerates the support values at or above the threshold. It makes no performance, convergence, or hardware claim.

Singular values by phase estimation over the Hermitian embedding

Estimate a matrix's singular values from the phase of the Hermitian matrix that carries it as its off-diagonal block, with a private block encoding of that embedding under a stated subnormalisation.

  • Maturity: Experimental
  • Public API: flagquantum.algorithms.svd
  • Runtime modes: local_statevector
  • Hardware: cpu
  • Gradient support: not_applicable
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: The advantage premise is the input model, and this unit does not meet it. The cited algorithm's cost is counted in queries to a structure that returns the matrix's entries, and against that count the state the estimation is applied to is assumed to be preparable; neither is present here. The matrix is held as an ordinary tensor, its embedding is formed and exponentiated as a dense matrix, and the input state is built from the singular vectors a classical torch.linalg.svd returns -- the very decomposition the readout estimates -- so the access and the preparation are paid explicitly rather than assumed away and no end-to-end advantage follows. Dequantization is recorded rather than glossed over: Tang's classical algorithm for the recommendation problem removes the exponential speed-up and is only polynomially slower, its bound containing eps**-12, which the author calls a large slowdown in some exponents; it is not a classical algorithm that matches the quantum runtime. Arrazola et al. record the practical conditions the dequantized algorithms need, and Gharibian-Le Gall dequantize the quantum singular value transformation for sparse matrices at constant precision; their hardness result is for a different task, estimating a local Hamiltonian's ground-state energy at inverse-polynomial precision given a state close to the ground state. The block encoding is not free to read: a readout that post-selects the ancilla succeeds with probability ||(A/alpha)|psi>||**2 on a normalised input, whose greatest value over inputs is (||A||/alpha)**2, and where alpha is much larger than ||A|| that probability is exponentially small. The readout is the counting register's mode, and the unit does not claim that the mode's value is the largest singular value: the readout is a grid value at the register's own resolution, and within() is the whole of the accuracy contract. The counting width is the caller's and is at least one; at one wire the register can only return alpha itself, and a run whose mode falls in the register's lower half is refused rather than returning a value above alpha -- an outcome of the sample rather than a precondition on the arguments. The input state is prepared from the singular vectors, so the classical work includes the very decomposition the unit estimates. Demonstration scale: at most four rows and four columns, with the embedding, its exponential and the state preparation all classical. It makes no performance, convergence, or hardware claim.

Distributed execution

Sharded statevector training

Partition one logical statevector workload across ranks while preserving differentiable training semantics.

  • Maturity: Production supported
  • Public API: fq.plan, flagquantum.experimental.distributed.train_distributed_statevector
  • Runtime modes: distributed_statevector
  • Hardware: multi_gpu, multi_node
  • Gradient support: exact
  • Distribution semantics: sharded_across_ranks
  • Start: quick example
  • Documentation: guide
  • Known boundary: Multi-node release certification remains dependent on promoted audited hardware evidence.

FlagOS distributed statevector workloads

Run sharded statevector forward and bounded training workloads through the public FlagOS boundary on the locked CUDA development reference.

  • Maturity: Development evidence
  • Public API: flagquantum.experimental.distributed.train_distributed_statevector
  • Runtime modes: distributed_statevector
  • Hardware: nvidia_a800_cuda_reference, single_node_2_4_8_gpu
  • Gradient support: development_evidence_exact_autograd
  • Distribution semantics: sharded_across_ranks
  • Start: quick example
  • Documentation: guide
  • Known boundary: Development evidence only for complex64 and complex128 on one CUDA-backed A800 node at 2, 4, and 8 cards. The inner communication route and host staging remain unattributed; complex reduce_scatter_tensor is unsupported in the tested full collective matrix; multi-node behavior, single-device capacity failure, performance, convergence, production support, scalability, and release certification are not established.

FlagOS statevector capacity expansion

Demonstrate one matched complex128 statevector that fails on a single device and completes when sharded across eight devices through the public FlagOS boundary.

  • Maturity: Development evidence
  • Public API: flagquantum.experimental.distributed.train_distributed_statevector
  • Runtime modes: distributed_statevector
  • Hardware: nvidia_a800_cuda_reference, single_node_8_gpu
  • Gradient support: forward_only_development_evidence
  • Distribution semantics: sharded_across_ranks
  • Start: quick example
  • Documentation: guide
  • Known boundary: One exact 32-qubit complex128 forward workload on one CUDA-backed eight-A800 node. The single-device and replicated paths measured OOM while the eight-rank sharded path completed. The inner communication route and host staging remain unattributed; determinism replay, backward and optimizer capacity, performance, convergence, multi-node behavior, production support, general scalability, and release certification are not established.

FlagOS distributed transport observability

Observe a fixed multi-rank complex collective matrix through the public FlagOS boundary while keeping inner-route and host-staging claims fail-closed.

  • Maturity: Development evidence
  • Public API: flagquantum.experimental.distributed.train_distributed_statevector
  • Runtime modes: distributed_transport_observation
  • Hardware: nvidia_a800_cuda_reference, single_node_2_4_8_gpu
  • Gradient support: not_applicable
  • Distribution semantics: rank_local_collective_observation
  • Start: quick example
  • Documentation: guide
  • Known boundary: One CUDA-backed A800 node at 2, 4, and 8 ranks for four collectives with complex64 and complex128. Correctness and logical FlagOS residency passed, but CUPTI device activity capture was incomplete. The inner communication route, FlagCX use, absence of host staging, performance, multi-node behavior, scalability, production support, and release certification are not established.

Differentiable and sharded MPS training

Train low-entanglement quantum systems with local or rank-owned matrix product states.

  • Maturity: Development evidence
  • Public API: flagquantum.simulation.mps.run_mps, flagquantum.experimental.distributed.train_distributed_mps
  • Runtime modes: mps, distributed_mps
  • Hardware: cpu, single_gpu, multi_gpu, multi_node
  • Gradient support: exact
  • Distribution semantics: sharded_across_ranks
  • Start: quick example
  • Documentation: guide
  • Known boundary: Single-node and dual-node execution plus matched checkpoint/restart have development evidence. The only public capacity measurement is emitted from the validated claim below; it is one exact-workload result, not general scalability or release evidence. Boundary instructions still execute serially by owner, and layer-parallel contraction/SVD, capacity multi-step soak, a sealed fault matrix, repeated evidence, and the release payload remain incomplete.

Deployment and extension

Circuit packaging and cloud deployment

Package trained circuits, export provider formats, and route them through deployment provider abstractions.

  • Maturity: Development evidence
  • Public API: flagquantum.deployment.create_deployment_package, flagquantum.deployment.deploy_circuit
  • Runtime modes: provider
  • Hardware: provider_dependent
  • Gradient support: not_applicable
  • Distribution semantics: provider_dependent
  • Start: quick example
  • Documentation: guide
  • Known boundary: Provider support and credential/runtime behavior vary; no provider is release-certified by this matrix.

Evidence-qualified QPU digital twins

Predict calibration-conditioned measurement distributions, prospectively validate fixed and held-out regional suites, and align comparable evidence over time.

  • Maturity: Development evidence
  • Public API: flagquantum.twin
  • Runtime modes: offline_density_prediction, offline_evidence_assessment, explicit_provider_submission
  • Hardware: cpu, quafu_development_evidence
  • Gradient support: unsupported
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: Frozen mapped models, validation series, calibration and validation histories, prospective candidate comparisons, topology- and depth-qualified evidence, connected structural cell composition, fail-closed regional model composition, repeated exact-circuit validation, fixed regional suite validation, predeclared reference/holdout comparison with simultaneous confidence, and longitudinal fixed-suite regional histories are available. Agreement is total-variation agreement for classical measurement distributions, not quantum-state fidelity. Region-model composition never infers cross-cell correlated noise or combines local bounds into region accuracy. Holdout evidence applies only to the prospectively declared circuits and does not establish arbitrary-circuit accuracy or training generalization. Regional suite histories compare only the same ordered fixed suite, mapping, and full topology; they do not define a trust window. Evidence remains specific to declared circuits, operations, mappings, physical couplers, depth, calibration snapshots, and confidence bounds. Automatic calibration collection, scheduling, arbitrary-circuit generalization, region-wide statistical inference, trust policy, model promotion, global publication, and release-certified provider support remain outside the framework capability.

Interoperability adapter contract

Implement and certify optional external-framework conversion behind one immutable lazy registry and framework-neutral, loss-aware result contract.

  • Maturity: Experimental
  • Public API: flagquantum.ecosystem
  • Runtime modes: control_plane_conversion
  • Hardware: cpu_control_plane
  • Gradient support: adapter_defined
  • Distribution semantics: not_applicable
  • Start: quick example
  • Documentation: guide
  • Known boundary: The framework-neutral protocol is candidate-stable pending API-owner approval; PennyLane and Qiskit implementations remain experimental. Common conformance does not install dependencies, sandbox third-party Python, certify provider hardware or numerical equivalence, or permit external objects to enter runtime and accelerator layers.

PennyLane QuantumScript interoperability

Translate supported immutable PennyLane QuantumScript programs to versioned FlagQuantum IR and back through an isolated, loss-aware control-plane adapter.

  • Maturity: Experimental
  • Public API: flagquantum.ecosystem.pennylane.from_pennylane, flagquantum.ecosystem.pennylane.to_pennylane
  • Runtime modes: control_plane_conversion
  • Hardware: cpu_control_plane
  • Gradient support: bound_parameters_only
  • Distribution semantics: not_applicable
  • Start: quick example
  • Documentation: guide
  • Known boundary: Certified with PennyLane 0.44.1 and 0.45.1 on Python 3.11 or newer for static QuantumScript conversion and complex128 numerical semantics. QNode, device execution, shots, measurements, trainable parameters, arbitrary wire labels without explicit lossy flattening, and idle wire extents are outside v1. PennyLane objects never enter FlagQuantum runtime, Torch-FL, CUDA, vendor accelerator, or QPU layers.

Qiskit IR interoperability

Translate supported Qiskit circuits to versioned FlagQuantum IR and export FlagQuantum IR through an isolated, loss-aware control-plane adapter.

  • Maturity: Experimental
  • Public API: flagquantum.ecosystem.qiskit.from_qiskit, flagquantum.ecosystem.qiskit.to_qiskit
  • Runtime modes: control_plane_conversion
  • Hardware: cpu_control_plane
  • Gradient support: symbolic_parameters_only
  • Distribution semantics: not_applicable
  • Start: quick example
  • Documentation: guide
  • Known boundary: Certified against Qiskit 2.0.x and 2.5.x with Aer 0.17.x through an executable operation, wire-order, statevector, classical-bit, arithmetic-parameter-expression, custom-unitary, and round-trip contract. Six fixed-seed differential programs exercise both conversion directions across three to five wires, mixed one- to three-wire operations, reordered wires, and asymmetric custom unitaries. ParameterExpression import supports the FlagQuantum v1 add/multiply/negate arithmetic subset after Qiskit symbolic simplification; functions, powers, and other operations fail closed. One- to three-qubit custom unitary matrices are converted with explicit local basis-order normalization and validated before export. Qiskit control flow is rejected; named or multiple registers require explicit lossy flattening; custom unitary matrices above three qubits are rejected. Conversion does not make Qiskit a runtime dependency or certify any provider hardware.

Dynamic circuits and backend assessment

Execute dynamic circuits locally, including a bounded bit-flip/readout-noise profile, and assess whether a backend can support their required features.

  • Maturity: Experimental
  • Public API: flagquantum.dynamic.DynamicCircuit, flagquantum.experimental.dynamic.assess_dynamic_backend, flagquantum.experimental.dynamic.run_dynamic
  • Runtime modes: local_statevector_trajectory, local_statevector_noisy_trajectory, backend_assessment
  • Hardware: cpu, provider_profiles_unverified
  • Gradient support: unsupported
  • Distribution semantics: single_process
  • Start: quick example
  • Documentation: guide
  • Known boundary: DynamicCircuit construction is candidate-stable pending API-owner approval; dynamic execution and backend assessment remain experimental. Local dynamic noise is limited to one-wire bit-flip channels after matching executed gates and independent readout confusion on explicit measurements and final sampling. Other Kraus channels, correlated readout, device-profile timing noise, noisy gradients, and provider-noise execution fail closed. Routing, deployment packaging, dialect export and provider integration are internal workflows rather than public experimental APIs. Provider-neutral conformance passes locally and on Qiskit Aer, but no real IQM QPU task was used.

Extension SDK

Build and qualify optional extensions through the Ecosystem extension protocol.

  • Maturity: Experimental
  • Public API: flagquantum.ecosystem.extensions
  • Runtime modes: extension_defined
  • Hardware: extension_defined
  • Gradient support: extension_defined
  • Distribution semantics: extension_defined
  • Start: quick example
  • Documentation: guide
  • Known boundary: The migrated SDK protocol is approved but not frozen; individual extensions remain experimental until separately qualified. Compiler plugins currently exchange CircuitIR only; pulse and native-binary artifacts are not supported.

Validated public performance claims

Every value below is read from a hash-bound raw artifact. Missing or changed evidence makes the source-of-truth check fail closed.

ClaimMaturity and scopeArtifact-derived resultRecorded environmentEvidence identity
Sharded MPS exact-workload capacity
mps-capacity-131072-chi768-20260806
development_evidence
One batch-one, complex64, χ768 MPS training step for the checked-in all-rank and all-boundary workload. This is not arbitrary statevector capacity, fixed-plan strong scaling, or release evidence.
Sites: 131,072
Logical MPS state: 1,236,780,012,864 bytes (1,151.84 GiB)
Maximum elapsed time: 367.37 s
Maximum peak allocated memory per rank: 77,745,407,488 bytes (72.41 GiB)
Cumulative discarded weight: 8.39e-06
Ranks: 16
Reported device memory per rank: 85,093,777,408 bytes (79.25 GiB)
CUDA allocator policy: expandable_segments:True
Topology fingerprint: c74a91e3a224a4dd414cfbfbcb31da7570880d42e6aa4eaccecf4b85427940a5
Metadata boundary: The artifact records rank count, per-rank device memory, topology fingerprint, and CUDA allocator policy. It does not record the exact GPU model or Python, PyTorch, CUDA, NCCL, driver, host, and operating-system versions, so the claim is restricted to the recorded environment fields.
raw JSON
SHA-256 df8c19b74cc173799e3894025fb8e72de72a786d4542a9aa5e67687298f482f5
code 9d56a6ecd78b06f11b9ee6e8aadcbe9644f2c708