dxCompiler Architecture

September 20, 2021 ยท View on GitHub

dxCompiler is a Scala application that performs two functions:

  1. Compiler: translates tasks/workflows written in a supported language (currently WDL and CWL) into native DNAnexus applets/workflows
  2. Executor: Executes tasks and workflow constructs (such as scatter/gather) within the applets generated by the compiler, in a language-specific manner

Compiler

The compiler has three phases:

  1. Parsing: The source code written in a workflow language (WDL or CWL) is parsed into an abstract syntax tree (AST). This depends on a language-specific library:
  2. Translation: The language-specific AST is translated into dxCompiler "Intermediate Representation" (IR), which mirrors the structure of DNAnexus applets/workflows, including DNAnexus-supported data types, metadata, and runtime specifications. The output of this step is an IR "Bundle".
  3. Compilation: Generates native applets and workflows from the IR. The output of this step is the applet ID (for a source file that contains a single tool/task) or workflow ID.

Translation

Compilation

A WDL workflow is compiled into an equivalent DNAnexus workflow, enabling running it on the platform. The basic mapping is:

  • A WDL task compiles to a DNAnexus applet (dx:applet).
  • A WDL workflow compiles to a DNAnexus workflow (dx:workflow)
  • A WDL call compiles to a DNAnexus workflow stage, and sometimes an auxiliary applet
  • Each scatters or conditional block is compiled into workflow stages, plus an auxiliary applet
    • Nested scatters/conditionals may result in nested workflows

The mapping from WDL to native DNAnexus objects requires overcoming two key obstacles:

  • We wish to avoid creating a controlling applet that would run and manage a WDL workflow. Such an applet might get killed due to temporary resource shortage, causing an expensive workflow to fail.
  • We want to minimize the context that needs to be kept around for the WDL workflow, because it limits job manager scalability.

Examining compilation output

After compiling your workflow to a project, you will see the compiled workflows and applets in the project like this:

my_task_1
my_task_2
my_workflow
my_workflow_common
my_workflow_outputs

To look inside the generated job script for an applet, you can run e.g. dx get my_task_1. This will download a folder called my_task_1 with the applet job script inside my_task_1/src.

The workflow source code (afer some processing during compilation, not necessarily matching the input WDL/CWL) can be extracted from a compiled workflow / compiled task applets using scripts/extract_source_code.sh (First have installed: base64, jq, gzip).

After compiling, you will also see dxWDLrt : record-xxxx in the project. This record points to the asset bundle of resources published to the platform during each dxCompiler release. It contains dxExecutorWdl.jar (or dxExecutorCwl.jar for CWL), which is needed during workflow execution.

Type mapping

WDL supports complex and recursive data types, which are not natively supported on DNAnexus. In order to maintain the usability of the UI, when possible, we map WDL types to the DNAnexus equivalent. This works for primitive types (Boolean, Int, String, Float, File), and for single-dimensional arrays of primitives. However, difficulties arise with nested types. For example, a ragged array of strings Array[Array[String]] presents two issues:

  • Type: Which DNAnexus type should we use, so that it will be presented intuitively in the UI?
  • Size: ragged arrays may get very large. A naive approach is to serialize them as strings. However, this has resulted in strings in excess of 100KB for real-world workflows. This is too large to present comfortably on the screen.

The type mapping for primitive types is:

WDL typeDNAx type
Booleanboolean
Intint
Floatfloat
Stringstring
Filefile
Directorynot yet supported

Optional primitives are mapped as follows:

WDL typeDNAx typeoptional
Boolean?booleantrue
Int?inttrue
Float?floattrue
String?stringtrue
File?filetrue
Directory?not yet supportedtrue

Single dimensional arrays of WDL primitives are mapped to DNAx optional arrays, because it allows them to be empty. The default DNAnexus array type is required to have at least one element.

WDL typeDNAX typeoptional
Array[Boolean]array:booleantrue
Array[Int]array:inttrue
Array[Float]array:floattrue
Array[String]array:stringtrue
Array[File]array:filetrue
Array[Directory]not yet supportedtrue

WDL types that fall outside these categories (e.g. ragged array of files Array[Array[File]]) are mapped to two fields: a flat array of files, and a hash, which is a JSON-serialized representation of the WDL value. The flat file array informs the job manager about data objects that need to be closed and cloned into the workspace.

Imports and nested namespaces

A WDL file creates its own namespace. It may import other WDL files, each inhabiting its own namespace. Tasks and workflows from children can be called with their fully-qualified-names. We map the WDL namespace hierarchy to a flat space of dx:applets and dx:workflows in the target project and folder. To do this, we make sure that tasks and workflows are uniquely named.

In a complex namespace, a task/workflow can have several definitions. Such namespaces cannot be compiled by dxCompiler.

Compiling a task

A task is compiled into an applet that has an equivalent signature. For example, a task such as:

version 1.0

task count_bam {
  input {
    File bam
  }
  command <<<
    samtools view -c ${bam}
  >>>
  runtime {
    docker: "quay.io/ucsc_cgl/samtools"
  }
  output {
    Int count = read_int(stdout())
  }
}

is compiled into an applet with the following dxapp.json:

{
  "name": "count_bam",
  "dxapi": "1.0.0",
  "version": "0.0.1",
  "inputSpec": [
    {
      "name": "bam",
      "class": "file"
    }
  ],
  "outputSpec": [
    {
      "name": "count",
      "class": "int"
    }
  ],
  "runSpec": {
    "interpreter": "bash",
    "file": "code.sh",
    "distribution": "Ubuntu",
    "release": "16.04"
  }
}

The code.sh bash script runs the docker image quay.io/ucsc_cgl/samtools, under which it runs the shell command samtools view -c ${bam}.

A Linear Workflow

Workflow linear (below) takes integers x and y, and calculates 2*(x + y) + 1. Integers are used for simplicity; more complex types such as maps or arrays could be substituted, keeping the compilation process exactly the same.

version 1.0

workflow linear {
  input {
    Int x
    Int y
  }

  call add { input: a = x, b = y }
  call mul { input: a = add.result, b = 2 }
  call inc { input: a = mul.result }

  output {
    Int result = inc.result
  }
}

# Add two integers
task add {
  input {
  Int a
  Int b
  }
  command {}
  output {
    Int result = a + b
  }
}

# Multiply two integers
task mul {
  input {
    Int a
    Int b
  }
  command {}
  output {
    Int result = a * b
  }
}

# Add one to an integer
task inc {
  input {
    Int a
  }
  command {}
  output {
    Int result = a + 1
  }
}

linear has no expressions and no if/scatter blocks. This allows direct compilation into a dx:workflow, which schematically looks like this:

phasecallarguments
Inputsx, y
Stage 1applet addx, y
Stage 2applet mulstage-1.result, 2
Stage 3applet incstage-2.result
Outputssub.result

In addition, there are three applets that can be called on their own: add, mul, and inc. The image below shows the workflow as an ellipse, and the standalone applets as light blue hexagons.

Fragments

The compiler can generate applets that are able to fully process simple parts of a larger workflow. These are called fragments. A fragment comprises a series of declarations followed by a call, a conditional block, or a scatter block. Native workflows do not support variable lookup, expressions, or evaluation. This means that we need to launch a job even for a trivial expression. The compiler tries to batch such evaluations together, to minimize the number of jobs. For example, workflow linear2 is split into three stages, the last two of which are fragments.

workflow linear2 {
  input {
    Int x
    Int y
  }

  call add { input: a=x, b=y }

  Int z = add.result + 1
  call mul { input: a=z, b=5 }

  call inc { input: i= z + mul.result + 8}

  output {
    Int result = inc.result
  }
}

Task add can be called directly, no fragment is required. Fragment-1 evaluates expression add.result + 1, and then calls mul.

Int z = add.result + 1
call mul { input: a=z, b=5 }

Fragment-2 evaluates z + mul.result + 8, and then calls inc.

call inc { input: i= z + mul.result + 8 }

Workflow linear2 is compiled into:

phasecallarguments
Inputsx, y
Stage 1applet addx, y
Stage 2applet fragment-1stage-1.result
Stage 3applet fragment-2stage-2.z, stage-2.mul.result
Outputsstage-3.result

Workflow optionals uses conditional blocks. It can be broken down into two fragments.

workflow optionals {
  input {
    Boolean flag
    Int x
    Int y
  }

  if (flag) {
    call inc { input: a=x }
  }
  if (!flag) {
    call add { input: a=x, b=y }
  }

  output {
    Int? r1 = inc.result
    Int? r2 = add.result
  }
}

Fragment 1:

if (flag) {
call inc { input: a=x }
}

Fragment 2:

if (!flag) {
call add { input: a=x, b=y }
}

The fragments are linked together into a dx:workflow like this:

phasecallarguments
Inputsflag, x, y
Stage 1applet fragment-1flag, x
Stage 2applet fragment-2flag, x, y
Outputsstage-1.inc.result, stage-2.add.result

Workflow mul_loop loops through the numbers 0, 1, .. n, and multiplies them by two. The result is an array of integers.

workflow mul_loop {
  input {
    Int n
  }

  scatter (item in range(n)) {
    call mul { input: a = item, b=2 }
  }

  output {
    Array[Int] result = mul.result
  }
}

It is compiled into:

phasecallarguments
Inputsn
Stage 1applet fragment-1n
Outputsstage-1.mul.result

The fragment is executed by an applet that calculates the WDL expressions range(n), iterates on it, and launches a child job for each value of item. In order to massage the results into the proper WDL types, we run a collect sub-job that waits for the child jobs to complete, and returns an array of integers.

Nested blocks

WDL allows blocks of scatters and conditionals to be nested arbitrarily. Such complex workflows are broken down into fragments, and tied together with subworkflows. For example, in workflow two_levels the scatter block requires a subworkflow that will chain together the calls inc1, inc2, and inc3. Note that inc3 requires a fragment because it needs to evaluate and export declaration b.

workflow two_levels {
  input {
  }

  scatter (i in [1,2,3]) {
    call inc as inc1 { input: a = i}
    call inc as inc2 { input: a = inc1.result }

    Int b = inc2.result

    call inc as inc3 { input: a = b }
  }

  if (true) {
    call add { input: a = 3, b = 4 }
  }

  call mul {input: a=1, b=4}

  output {
    Array[Int] a = inc3.result
    Int? b = add.result
    Int c = mul.result
  }
}

It will be broken down into five parts. A sub-workflow will tie the first three pieces together:

Part 1:

call inc as inc1 { input: a = i}

Part 2:

call inc as inc2 { input: a = inc1.result }

Part 3 (fragment A):

Int b = inc2.result
call inc as inc3 { input: a = b }

The top level workflow calls a scatter applet, which calls the sub-workflow. Later, it calls parts four and five.

Part 4 (fragment B):

if (true) {
  call add { input: a = 3, b = 4 }
}

Part 5:

call mul {input: a=1, b=4}

The overall structure is

Executor

The Executor has two branches: TaskExecutor, for executing individual tasks, and WorkflowExecutor, for executing workflow constructs, such as scatter/gather, conditionals, and special applets for input and output expression evaluation, and for output reorganization.

The command line interface for the Executor is fairly simple:

java -jar dxExecutorWdl.jar <task|workflow> <action> <rootdir> [options]

Options:
    -streamFiles [all,none,perfile]
                           Whether to mount all files with dxfuse (do not use the
                           download agent), to mount no files with dxfuse (only use
                           download agent), or to allow streaming to be set on a
                           per-file basis (the default).
    -separateOutputs       Whether to put output files in a separate folder based on
                           the job name. If not specified, then all outputs go to the
                           parent job's output folder.
    -traceLevel [0,1,2]    How much debug information to write to the
                           job log at runtime. Zero means write the minimum,
                           one is the default, and two is for internal debugging.
    -quiet                 Do not print warnings or informational outputs
    -verbose               Print detailed progress reports
    -verboseKey <module>   Detailed information for a specific module
    -logFile <path>        File to use for logging output; defaults to stderr
    -waitOnUpload          Whether to wait for each file upload to complete.

On a DNAnexus worker, rootdir is always /home/dnanexus.

The streamAllFiles option overrides any streaming settings in the compiled apps, so that all files are streamed to the worker using dxfuse.

TaskExecutor

The TaskExecutor (dxExecutorWdl.jar, dxExecutorCwl.jar) implements the commands called in the job script generated for each task applet by the compiler. The progression of events is:

  • TaskExecutor prolog:
    • Evaluate all job inputs
    • Extract the input files that need to be localized to the worker
    • Write dxda and/or dxfuse manifests (depending on the global and file-specific streaming settings)
    • Update the inputs to replace file URIs with the localized paths
    • Serialize the evaluated and updated inputs to a (transient) JSON file
  • The job script executes dxda and/or dxfuse to download the files in the manifest(s) generated by the prolog
  • TaskExecutor instantiateCommand
    • Deserialize the localized inputs
    • Evaluate any private variables
    • Evaluate the task's command block to replace all interpolation placeholders with their values
    • Write the evaluated command block to a script
    • Re-serialize the localized inputs with the private variables included
    • NOTE: the WDL implementation currently evaluates all private variables during this step. This means that dxCompiler cannot currently be used for tasks with private variables (i.e. not in the output {} section) that depend on the execution of the command.
  • The job script executes the task command script
  • TaskExecutor epilog:
    • Deserialize the localized inputs
    • Evaluate the parameters in the output {} block
    • Extract all file paths from the evaluted outputs
    • Upload any files that were generated by the task command
    • Replace all local file paths in the output values with URIs
    • Write the de-localized outputs to the job_output.json meta file