Banner for ZETOBIT LLC Bioinformatics Blog with a light green background, connected dots and lines, and text listing genomics, transcriptomics, proteomics, and metabolomics.

Expert analysis at the intersection of AI, bioinformatics, and genomics.

Breaking down the science shaping precision medicine — from whole-genome sequencing and liquid biopsy to spatial transcriptomics, proteomics, and metabolomics — grounded in peer-reviewed literature and written for the people building the next generation of life sciences products.

Archive

The archive now runs to roughly 100 articles. Use the search below to find the ones relevant to your work.

Search the blog by entering the keywords below and hitting “Enter“

Disclaimer — Zetobit, LLC
Zetobit, LLC  ·  Bioinformatics

Editorial Notice

Disclaimer

What this blog is, what it is not, and the terms on which it is published.

Scope

The content published on this blog by Zetobit LLC is provided for general educational and informational purposes only. It reflects the professional views of the author, Kanna Nandakumar, PhD, and does not represent the views of any current or former employer, client, or institution.

Not advice

Nothing here constitutes medical, clinical, diagnostic, or legal advice. It is not a substitute for consultation with a qualified physician, genetic counselor, or the laboratory that issued a given test result, and it must not be used to make diagnostic or treatment decisions. Reading this material does not create a client, consulting, or professional relationship with Zetobit LLC.

Illustrative content

All report excerpts, figures, variants, and data shown are synthetic or illustrative. They contain no patient data and describe no real individual.

Currency & accuracy

Scientific literature, professional guidelines, and regulatory requirements change. Articles reflect the state of the field as of their publication date and are not updated. While the author makes reasonable efforts to verify claims and cite primary sources, no warranty is made as to accuracy or completeness, and Zetobit LLC accepts no liability for any action taken in reliance on this content.

Last updated August 2026
Questions about this notice — zetobit.com/contact

Microbiome Data Is Compositional: Why a Taxon Can Appear to Rise When Nothing About It Changed
Kanna Nandakumar Kanna Nandakumar

Microbiome Data Is Compositional: Why a Taxon Can Appear to Rise When Nothing About It Changed

Microbiome sequencing doesn't count bacteria — it reports each taxon's share of a fixed pool of reads, and those shares are forced to sum to one. That single constraint quietly manufactures false findings: correlations biased negative even for independent taxa, and long lists of taxa that "significantly decrease" when the only real event was one organism blooming and squeezing everyone else's slice. Rarefying and subsetting don't fix it. This piece explains the constant-sum constraint, why log-ratio methods like ALDEx2 and ANCOM are the honest alternative, and why answering an absolute question needs an external anchor like qPCR or a spike-in.

Read More
Sample Swaps and Identity QC: Why the Cleanest VCF in Your Batch Can Belong to the Wrong Person
Kanna Nandakumar Kanna Nandakumar

Sample Swaps and Identity QC: Why the Cleanest VCF in Your Batch Can Belong to the Wrong Person

A mislabeled tube, a pipetting error, a transposed row on a plate — and a sample produces a flawless VCF that belongs to someone else, or to two people at once. Every technical metric passes, because technically nothing failed: read quality, coverage, Ts/Tv are all normal. Standard QC is structurally blind to it, because those metrics ask "did the sequencing work?" not "is this the right person?" The only check that catches a swap is a genetic fingerprint — genotyping a panel of common SNPs to ask the genome who it belongs to. This piece explains why identity is a separate QC axis, and why tumor-normal work is especially exposed.

Read More
RNA Velocity: Why a Confident Arrow of Cellular Direction Can Point the Wrong Way
Kanna Nandakumar Kanna Nandakumar

RNA Velocity: Why a Confident Arrow of Cellular Direction Can Point the Wrong Way

RNA velocity promises what a static snapshot can't give: the direction each cell is heading, inferred from the ratio of unspliced to spliced RNA. The arrows are visually compelling and often biologically right — but they rest on kinetic assumptions real data frequently violates (genes with multiple rate kinetics can produce distorted or even reversed velocity), the unspliced signal is faint and noisy, and the smooth streamline you see can be manufactured by the 2D projection rather than read from the biology. One critique produced a coherent-looking flow from three genes with no directional signal at all. This piece explains why a clean velocity stream is a hypothesis, not a measurement.

Read More
Variant Normalization: Why the Same Variant Written Two Ways Silently Fails to Match
Kanna Nandakumar Kanna Nandakumar

Variant Normalization: Why the Same Variant Written Two Ways Silently Fails to Match

A single insertion or deletion can be written in several equally correct VCF representations — different position, different alleles, same biological edit. To a human they're obviously identical. To a computer doing an exact match against ClinVar, they're different variants, and the mismatch is completely silent: every job runs, every file is well-formed, and a pathogenic variant simply fails to annotate. This piece explains why VCF doesn't enforce a unique spelling, how left-alignment and parsimony produce one canonical form (and the uniqueness theorem that makes matching correct), why multiallelic decomposition is a required partner step, and why two different normalizers can still disagree on indels in repeats.

Read More
False Positives in Metagenomic Classification: Why a Handful of Reads Can Invent a Pathogen That Was Never There
Kanna Nandakumar Kanna Nandakumar

False Positives in Metagenomic Classification: Why a Handful of Reads Can Invent a Pathogen That Was Never There

In metagenomic pathogen detection the organism you're hunting may be fewer than 100 reads out of tens of millions — and contamination, a conserved rRNA gene, or a dirty database entry can produce exactly that same handful. Read count alone cannot tell the real signal from the manufactured one. This piece explains why the instinctive metric fails at low abundance, why false-positive reads share a geometry (they pile up in the same few spots while real organisms spread across the genome), and how unique-k-mer counting operationalizes breadth of coverage — the KrakenUniq threshold of ~1000 unique k-mers eliminated many background pathogen identifications that a read-count filter would have passed.

Read More
HLA Typing from Short Reads: Why the Most Polymorphic Region in the Genome Defeats Standard Alignment
Kanna Nandakumar Kanna Nandakumar

HLA Typing from Short Reads: Why the Most Polymorphic Region in the Genome Defeats Standard Alignment

Typing HLA is not variant calling. It's choosing two alleles out of thousands whose sequences differ by less than the ambiguity of a single short read — against a reference database that is itself incomplete. That's why a pipeline that nails the rest of the genome can quietly return the wrong HLA type. This piece walks through why the standard align-to-GRCh38-and-call workflow fails at the MHC, why four-digit resolution collapses on short reads (one early method got only ~32% of four-digit class I calls right), why low-SNP-density loci like DQA1/DPA1/DPB1 can't be phased, and how the working tools model the ambiguity instead of ignoring it.

Read More
Polygenic Score Portability Across Ancestries: Why Accuracy Decays Continuously and a Population Label Cannot Tell You Where
Kanna Nandakumar Kanna Nandakumar

Polygenic Score Portability Across Ancestries: Why Accuracy Decays Continuously and a Population Label Cannot Tell You Where

A polygenic score trained in one population predicts poorly in another — that much is well known. What's less appreciated is that the loss isn't a step change between labeled groups. It's a smooth decline with genetic distance from the training data, and it happens inside a single ancestry label as surely as across the boundary between two. A large biobank analysis found accuracy correlates −0.95 with genetic distance across 84 traits, and the closest Hispanic/Latino individuals can be predicted as well as the farthest Europeans. This piece explains why the population label is the wrong unit — and what to report instead.

Read More
Calling Mitochondrial Heteroplasmy: Why a Low-Level Variant Is as Likely to Be Nuclear DNA as Real
Kanna Nandakumar Kanna Nandakumar

Calling Mitochondrial Heteroplasmy: Why a Low-Level Variant Is as Likely to Be Nuclear DNA as Real

Point a germline pipeline at chrM and it runs without complaint and hands you a VCF. That's the problem. The mitochondrial genome breaks nearly every assumption a standard variant caller is built on — it's circular, it carries a variant at any fraction rather than two clean alleles, and, most treacherously, fragments of it are scattered across the nuclear genome as NUMTs. Reads from those nuclear copies mis-map onto the mitochondrial reference and manufacture low-heteroplasmy variants that were never in a mitochondrion. Worse, whether they clear your threshold depends on the sample's mtDNA copy number. This piece covers why a flat cutoff is the wrong instinct, and what actually controls the confound.

Read More
Liftover Between Genome Assemblies: Why Coordinate Conversion Is Lossy and Silently Wrong for Indels
Kanna Nandakumar Kanna Nandakumar

Liftover Between Genome Assemblies: Why Coordinate Conversion Is Lossy and Silently Wrong for Indels

Moving a variant from GRCh37 to GRCh38 looks like arithmetic on a position — pass the coordinates through a chain file and you're done. It isn't that simple. A genome assembly is a specific reconstructed sequence, not a neutral ruler, and wherever two builds spell the same variant with different alleles, the standard liftover tools quietly drop it or convert it into a record that will never match a natively called genome. SNVs mostly survive; indels and short-tandem-repeat variants are where it breaks, silently. This piece walks through what a chain file actually encodes, why tool choice isn't cosmetic, and when you should realign instead of lift over.

Read More
Panel vs. Exome vs. Genome: Why More Sequencing Isn't Always More Answers
Kanna Nandakumar Kanna Nandakumar

Panel vs. Exome vs. Genome: Why More Sequencing Isn't Always More Answers

It's tempting to treat panel, exome, and genome as a ladder — more territory, more diagnoses, strictly better if you can afford it. But breadth trades against depth, interpretability, turnaround, and the burden of uncertain and incidental findings — and the widest test isn't always the one that answers your question. A deep panel can out-diagnose an exome for a well-defined phenotype; a uniform genome can out-detect a patchy exome for the right coding variant. This piece frames test selection as a design decision driven by the question — the phenotype's specificity, the likely variant class, and the reporting strategy — not a reflex toward "more."

Read More
Genome Assembly Metrics That Mislead: Why a High N50 Can Coexist With Structural Misassemblies
Kanna Nandakumar Kanna Nandakumar

Genome Assembly Metrics That Mislead: Why a High N50 Can Coexist With Structural Misassemblies

N50 is the number everyone quotes to prove a genome assembly is good: the bigger it is, the more contiguous the assembly. But N50 measures only how long the pieces are — not whether they're joined correctly. And the fastest way to raise it is to make exactly the kind of aggressive join that produces a chimera. A misjoin fuses two sequences that don't belong adjacent into one longer contig, which raises N50 — the structural error and the metric improvement are the same event. This piece covers the three Cs of assembly quality, why contiguity isn't correctness, and how to actually verify structure.

Read More
Coverage Uniformity vs. Mean Depth: Why the Average That Reassures You Hides the Dropout That Fails the Sample
Kanna Nandakumar Kanna Nandakumar

Coverage Uniformity vs. Mean Depth: Why the Average That Reassures You Hides the Dropout That Fails the Sample

"Sequenced to 100× mean coverage" sounds like every base was read a hundred times. It wasn't. Mean depth is an average over a distribution, and two samples with the same average can differ enormously in whether a given clinically important base was covered at all — because sequencing dropout is systematic, not random. GC-extreme regions, promoters and first exons, long exons, and repeats drop out in the same places every run, invisible in the mean. This piece covers why the average reassures while the distribution decides which variants get called, and which uniformity metrics — percent-above-threshold, Fold-80, evenness — actually capture what the mean hides.

Read More
Phasing and Haplotype Assembly Limits: Why "Phased" Has a Resolution You Rarely See Reported
Kanna Nandakumar Kanna Nandakumar

Phasing and Haplotype Assembly Limits: Why "Phased" Has a Resolution You Rarely See Reported

A phased genome sounds like a solved genome — two clean parental haplotypes, allele by allele. In reality "phased" means the chromosome was broken into blocks: contiguous stretches the method declares internally phased, separated by gaps it couldn't resolve, each carrying its own error rate, with no reliable phase relationship between blocks. The single number usually reported — block N50 — describes only how long the blocks are, not whether they're correct. And a switch error, one wrong junction, inverts every allele downstream of it, silently flipping a compound-heterozygote call from benign to disease-causing. This piece covers switch vs. flip errors, block boundaries, and the resolution "phased" hides.

Read More
Multiple Testing Across Omics: Why the Genome-Wide Correction You Trust Breaks When You Integrate Data Types
Kanna Nandakumar Kanna Nandakumar

Multiple Testing Across Omics: Why the Genome-Wide Correction You Trust Breaks When You Integrate Data Types

The GWAS threshold of 5×10⁻⁸ feels like a constant of nature. It isn't. It's a Bonferroni correction — 0.05 divided by roughly a million effectively independent tests, a count derived from the genome's linkage-disequilibrium structure. That works because the dependence is understood and confined to one data type. Multi-omic integration removes both comforts: you're testing SNPs, transcripts, proteins, and metabolites with different feature counts, different correlation structures, and correlations between them. The number of tests — the denominator every correction depends on — stops having a clean answer, and the same biological signal gets counted once, thrice, or not at all. This piece covers why.

Read More
Batch Integration in Single-Cell RNA-Seq: Why Removing Technical Variation Can Erase the Biology You Came to Find
Kanna Nandakumar Kanna Nandakumar

Batch Integration in Single-Cell RNA-Seq: Why Removing Technical Variation Can Erase the Biology You Came to Find

Combining single-cell datasets from different runs, labs, or chemistries means removing the technical batch effects between them. But technical and biological variation are entangled in the same numbers, and no algorithm separates them perfectly. Integration is a balancing act between two goals that pull in opposite directions — mixing the batches together vs. conserving real biology — with a dangerous failure mode: push too hard toward blending, and you don't just remove noise, you erase genuine differences. And because the output is a clean, well-mixed embedding, the erasure is invisible. This piece covers the tradeoff, why over-correction is common, and the confounding trap where batch IS the biology.

Read More
Cell-Type Annotation in Single-Cell RNA-Seq: Why a Confident Label Can Be Wrong, and Reference Mapping Can't Tell You It Is
Kanna Nandakumar Kanna Nandakumar

Cell-Type Annotation in Single-Cell RNA-Seq: Why a Confident Label Can Be Wrong, and Reference Mapping Can't Tell You It Is

Clustering gives you groups of cells; annotation gives them names — and that naming step feels like the objective payoff of the whole experiment. But every annotation method assigns a label by relating your cells to prior knowledge: a marker list, a reference atlas, a training set. It inherits that prior's limits without flagging them. The sharpest failure is structural: a standard classifier assigns every cell to one of the labels it was trained on, so a genuinely novel population gets forced to the nearest known type — confidently — instead of flagged as unknown. This piece covers the forced-label problem, why the reference isn't ground truth, and what defensible annotation requires.

Read More
Clustering Resolution in Single-Cell RNA-Seq: Why the Number of Cell Types Is a Parameter You Choose, Not a Fact You Discover
Kanna Nandakumar Kanna Nandakumar

Clustering Resolution in Single-Cell RNA-Seq: Why the Number of Cell Types Is a Parameter You Choose, Not a Fact You Discover

Single-cell analysis answers "how many cell types are in this sample?" with an unsupervised clustering step — and it's the step most easily mistaken for objective. The number of clusters isn't something the data determines; it's set directly by a resolution parameter the analyst chooses. Turn it up and the same cells split into more groups, each with its own confident marker genes. Worse, those markers are tested on the same data used to draw the clusters — a circularity that inflates p-values even when there's only one true population. This piece covers over-clustering, the double-dipping problem, and how to tell a real cluster from a manufactured one.

Read More
Doublets in Single-Cell RNA-Seq: Why Two Cells in One Droplet Fabricate Cell Types That Were Never There
Kanna Nandakumar Kanna Nandakumar

Doublets in Single-Cell RNA-Seq: Why Two Cells in One Droplet Fabricate Cell Types That Were Never There

The entire premise of single-cell RNA-seq is the word single: one droplet, one cell, one barcode. Doublets break it. When two cells share a droplet, their combined mRNA gets one barcode, and the analysis sees a single "cell" with a blended transcriptome. The danger is that the blend doesn't look like noise — it looks like a discovery. Two different cell types merged into one barcode land between their real clusters and read as a novel intermediate or transitional cell type that doesn't exist. This piece covers why the doublet rate is set when you load the chip, why detection is blind to same-type doublets, and what actually works.

Read More
DNA Methylation Calling from Bisulfite Sequencing: Why the Chemistry That Reveals Methylation Also Corrupts the Data
Kanna Nandakumar Kanna Nandakumar

DNA Methylation Calling from Bisulfite Sequencing: Why the Chemistry That Reveals Methylation Also Corrupts the Data

Bisulfite conversion is an elegant trick: it rewrites unmethylated cytosines as thymines and leaves methylated ones untouched, turning an invisible epigenetic mark into a readable C-versus-T base change. But the same reaction fragments the DNA with heat and acid, collapses the four-letter genome to three letters (wrecking alignment), breaks strand complementarity, and — most dangerously — makes an unmethylated cytosine that escaped conversion indistinguishable from a truly methylated one. Incomplete conversion is a false-positive engine that only a spiked-in control can quantify. This piece covers why the chemistry that reveals methylation also corrupts the data, and what a trustworthy methylation pipeline requires.

Read More
Allele-Specific Expression from RNA-Seq: Why the Reference Genome Biases the Very Imbalance You're Trying to Measure
Kanna Nandakumar Kanna Nandakumar

Allele-Specific Expression from RNA-Seq: Why the Reference Genome Biases the Very Imbalance You're Trying to Measure

Allele-specific expression asks a surgical question: within one sample, at a heterozygous site, is the maternal copy of a gene expressed as much as the paternal one? Because both alleles share a cell, any consistent imbalance points to a cis cause — a regulatory variant, imprinting, nonsense-mediated decay, X-inactivation — which makes ASE powerful where standard differential expression is blind. But the measurement has a thumb on the scale: reads are aligned to a reference that contains only one allele, so alternate-allele reads mismap more often and the reference count inflates. This piece covers reference bias, why correcting it is necessary but not sufficient, and what a trustworthy ASE pipeline requires.

Read More