+------------------------------------------------------------------------------+ | | | BIOINFORMATICS PROJECTS | | | +------------------------------------------------------------------------------+

I build tools to study genome evolution, from bacterial gene formation and antibiotic resistance to chromatin contacts and rare disease genetics.

DOCTORAL RESEARCH

Novel genes from genomic deletions

I studied how genomic deletions create new bacterial genes by fusing previously separate sequences into open reading frames.

I conceived and implemented the prefix-suffix k-mer approach during my Chateaubriand Fellowship with Karel Brinda, then applied it across millions of bacterial genomes, varying sampling depth to distinguish biological signal from sequencing bias.

Antibiotic resistance in historical isolates

I analyzed 1,800+ clinical isolates spanning 140 years and found that antibiotic resistance genes became more frequent and mobile after antibiotics entered clinical use. The analysis accounted for uneven sampling, though collection bias could not be fully excluded.

I built CallMeMobile to combine mobile-element predictions and identify potentially mobile genomic regions, then used it to investigate resistance-gene mobility. Published in Microbial Genomics.

Phylogeny-colored de Bruijn graphs

I conceived and implemented phylogeny-colored de Bruijn graphs to query genomic variation alongside evolutionary relationships. I developed the approach with Karel Brinda at Inria GenScale during my Chateaubriand Fellowship and applied it to more than 1,000 MRSA genomes.

TOOLS & COLLABORATIONS

Fit-Hi-C

Finding statistically significant chromatin contacts.

I led continued development of Fit-Hi-C, originally developed by Ferhat Ay, improving speed, memory use, packaging, and accessibility. The tool identifies statistically significant chromatin contacts in Hi-C data; our FitHiC2 protocol explains its use.

HiCKRy

Normalizing Hi-C contact maps with matrix balancing.

I implemented this Python tool in the Ay Lab to normalize Hi-C contact maps using Knight-Ruiz matrix balancing. It corrects coverage differences under an equal-visibility assumption without modeling each source of bias separately.

Simdigree

Simulated pedigrees for questions about rare disease.

I built Simdigree in the Sunyaev Lab to generate pedigrees from SLiM simulations and explore what family history reveals about monogenic and polygenic rare diseases.

NovaSplice

Predicting new splice sites caused by genetic variants.

I built NovaSplice in the Sunyaev Lab to predict splice sites created by non-coding variants. It compares candidate sites with nearby canonical sites and ranks potential novel splicing events.

Megatron

Comparing clones in lineage-tracing experiments.

I contributed to the Pinello/Morris collaboration on Megatron, a tool for measuring distances between clones in lineage-tracing datasets. Presented at CZI Seed Networks in 2020.

EARLIER COURSE PROJECTS

HotWASp

Exploring association signals through gene networks.

A project for UC San Diego’s BNFO 286 course with Trey Ideker. HotWASp propagates association signals through a gene network to explore whether network context can help identify signals in genome-wide association studies.

TErex

Finding transposable elements in DNA sequences.

A final project for UC San Diego’s CSE 180, Biology Meets Computing. TErex searches FASTA sequences for transposable elements using the Dfam database.



________________________________________________________________________________

Arya Kaul (C) now - forever