HipHap (previously Diplinator)
July 7, 2026 · View on GitHub
Diploid genome assemblies are now routinely available, but most read aligners were designed for haploid references. When reads are aligned to a diploid assembly, the aligner sees two nearly identical alignments to either haplotype, and thus reduces the mapping quality (MapQ) score to reflect this ambiguity. This can cause downstream tools to discard reads from easily mappable regions. HipHap resolves this issue by aligning reads to each haplotype assembly separately and assigning each read to its best-supported haplotype. We also introduce a haplotype assignment quality score (HapQ) in HipHap to quantify confidence in the haplotype of origin of a read.
HipHap is implemented in Rust, and supports SAM, BAM, CRAM, and PAF formats
Installation
Recommended: Download precompiled binary:
# Todo
Build from source:
git clone https://github.com/jheinz27/hiphap.git
cd hiphap
cargo build --release
./target/release/hiphap
Usage
HipHap: Choose the best alignment to each haploid of a diploid assembly
Usage: hiphap [OPTIONS] <ASM1> <ASM2>
Arguments:
<ASM1> asm1 alignment file (sam/bam/cram/paf)
<ASM2> asm2 alignment file (sam/bam/cram/paf)
Options:
-1, --s1 <NAME> label for asm1 sample (used in output file names and summary) [default: asm1]
-2, --s2 <NAME> label for asm2 sample (used in output file names and summary) [default: asm2]
--paf input files are PAF
--ms use ms:i: tag rather than AS:i: for alignment score
-b, --both write reads with equal alignment scores to both output files
-m, --merge write a single merged output file (hiphap_{s1}_{s2}_merged.*) instead of one file per haplotype
--ref-merged <FILE> combined reference FASTA for merged CRAM output (must contain all contigs of both inputs); required with --merge on CRAM input
-u, --unmapped <DEST> where to write reads unmapped in both assemblies: asm1, asm2, or discard [default: asm1] [possible values: asm1, asm2, discard]
--ref1 <FILE> reference FASTA for cram file (asm1)
--ref2 <FILE> reference FASTA for cram file (asm2)
-A, --match-sc <FLOAT> per-base match score from aligner scoring scheme (auto-estimated from ms:i tags if omitted)
--no-hapq skip HAPQ score calculation and hq tag output (e.g. for comparing GRCh38 vs CHM13)
--no-span-chrom disable writing hiphap_{s1}_{s2}_span_chrom.fastq
-t, --threads <INT> Total thread pool size (min 4). Multiples of 8 recommended for optimal read/write balance. [default: 8]
-h, --help Print help
-V, --version Print version
Each output record is annotated with an hq:i: tag carrying the HapQ score, unless --no-hapq is set.
The per-base match score used by the HAPQ calculation is auto-estimated from the ms:i: tags of the input files. Pass -A/--match-sc <FLOAT> to set it explicitly (e.g. 2.0 for minimap2 long-read defaults).
Example Workflow
HipHap only works on name-sorted files, which is the default minimap2 output. Therefore, coordinate-sorted files need to be name-sorted first, for example:
samtools sort -n -o name_sort.bam index_sort.bam
Diploid assembly alignment
# If needed, split diploid genome assembly FASTA into respective haplotypes
separate_haps_fasta -1 MATERNAL -2 PATERNAL hg002v1.1.fa
# writes hg002v1.1.MATERNAL.fa and hg002v1.1.PATERNAL.fa
# Align reads to each haplotype
minimap2 -ax map-hifi -o asm1_alignments.sam hg002v1.1.MATERNAL.fa reads.fastq
minimap2 -ax map-hifi -o asm2_alignments.sam hg002v1.1.PATERNAL.fa reads.fastq
# Run hiphap
hiphap -1 mat -2 pat asm1_alignments.sam asm2_alignments.sam
# Output: hiphap_mat.sam hiphap_pat.sam hiphap_mat_pat_span_chrom.fastq
# Merge best alignments into one file (if desired)
samtools merge -@ 12 merged.sam hiphap_mat.sam hiphap_pat.sam
# Save as sorted BAM
samtools sort -@ 12 -o hiphap_mat_pat_merged.bam hiphap_mat_pat_merged.sam
By default, HipHap also writes hiphap_{s1}_{s2}_span_chrom.fastq, a FASTQ of reads whose alignments span more than one chromosome. These reads are emitted for easy realignment. Pass --no-span-chrom to skip writing this file. If inputs are PAF files, a tsv of the chromosome spamnign reads rather than a fastq is output, as the read sequences are not stored in the input PAF files.
Merged output (--merge)
Instead of one file per haplotype, --merge writes a single combined output file with a merged header, so the separate samtools merge step is no longer needed:
hiphap --merge -1 mat -2 pat asm1_alignments.sam asm2_alignments.sam
# Output: hiphap_mat_pat_merged.sam hiphap_mat_pat_span_chrom.fastq
# Save as sorted BAM
samtools sort -@ 12 -o merged.bam hiphap_mat_pat_merged.sam
Notes:
-
In merge mode the two assemblies must have unique contig names. The merged header concatenates the two inputs
@SQlists, so any contig name shared between them is ambiguous -
In merge mode HipHap writes only the merged file; the per-haplotype
hiphap_{s1}.*/hiphap_{s2}.*files are not produced. -
--mergecannot be combined with--both. -
For CRAM inputs, merged output is also CRAM and requires a single combined reference (
--ref-merged <FILE>) containing all contigs of both haplotypes, typically the original diploid assembly:hiphap --merge --ref1 hg002v1.1.MATERNAL.fa --ref2 hg002v1.1.PATERNAL.fa \ --ref-merged hg002v1.1.fa mat_alignments.cram pat_alignments.cram
Comparing different reference genomes
HipHap can also be used to select best alignment of a read between different reference genomes (e.g. GRCh38 and CHM13). For this use case, the HAPQ score is generally not meaningful, so pass --no-hapq to skip its calculation:
hiphap --no-hapq -1 grch38 -2 chm13 grch38_alignments.sam chm13_alignments.sam
# Output: hiphap_grch38.sam hiphap_chm13.sam
CRAM input files
If input files are CRAM format, the original reference genomes must be provided, for example:
hiphap --ref1 asm1_hap.fasta --ref2 asm2_hap.fasta asm1_alignments.cram asm2_alignments.cram
Example PAF Usage
Notes:
- It is important to use the
--paf-no-hitflags when aligning with minimap2 to output unmapped reads to the file - If a SAM file is converted to a PAF file with
paftools.js sam2paf, it will NOT have the required AS:i: tag and HipHap will fail to run
minimap2 -cx map-hifi --paf-no-hit -o asm1_alignments.paf hg002v1.1.MATERNAL.fa reads.fastq
minimap2 -cx map-hifi --paf-no-hit -o asm2_alignments.paf hg002v1.1.PATERNAL.fa reads.fastq
hiphap --paf out_asm1.paf out_asm2.paf
# Output: hiphap_asm1.paf hiphap_asm2.paf hiphap_asm1_asm2_span_chrom.txt
--merge also works with PAF
hiphap --paf --merge out_asm1.paf out_asm2.paf
# Output: hiphap_asm1_asm2_merged.paf
Weighted Alignment Scoring Mechanism
For each read, HipHap computes a single weighted alignment score per assembly using all primary and supplementary alignments (secondary alignments are passed through to the output but ignored when scoring).
Let be the number of primary and supplementary alignments for a given read to that reference genome.
Let be the full read length (sum of query-consuming CIGAR operations on a non-secondary record, or the qlen field of the PAF record).
Let denote the interval on the read covered by alignment .
Let
be the number of read bases covered by at least one alignment.
Let be the alignment score for alignment (AS:i: by default, or ms:i: if --ms is set).
Let be the alignment length in read coordinates for alignment .
The first factor is the average alignment score per aligned base. It is multiplied by the number of unique read bases covered, , and then scaled by the read coverage fraction , so that reads which align over a large fraction of their length are weighted more heavily than reads which align only over a small portion.
For each read, the assembly with the higher wins; its full alignment cluster (including secondary alignments) is written to the corresponding output file. If is equal in both assemblies, the "better" assignment is determined by a hash of the read name, or the read is written to both output files when --both is used.
Haplotype Assignment Quality (HapQ)
For each read assigned to a winning haplotype, HipHap reports a HapQ score in the hq:i: tag of the output record. HapQ is a Phred-like confidence [0-60] that the read was assigned to the correct haplotype. The calculation is modeled on BWA-MEM's mem_approx_mapq_se.
Let and be the weighted alignment scores of the winning and losing assemblies, the per-base match score of the alignment software used (-A/--match-sc), and the number of non-secondary alignments (splits) on the winning side.
HapQ is the product of:
d, the approximate difference, in matching bases, between the winning and losing alignments between haplotypes.
The split penalty, , which penalizes reads with more than three split alignments, which often fall in complex/repetitive regions where haplotype assignment is less reliable.
The final HapQ calculation is then:
Clamped to the range [0,60].
Special cases:
- Read mapped in only one assembly: HapQ = 60.
- Read tied between assemblies (winner =
Both): HapQ = 0. - Read unmapped in both assemblies: no
hqtag is written.
If --no-hapq is set, HAPQ is not computed and no hq:i: tag is added (recommended when the two inputs are not haplotypes of the same sample, e.g. GRCh38 vs CHM13).
Citation
If HipHap has helped you in your research, please cite our preprint at: TODO