Lens species genome assemblies and the role of centromere in Lens chromosome stability

Overview
Objectives
  • Assemble, annotate and compare the genomes of the seven Lens species.

  • Perform gene synteny analysis to characterize the large genome structural variation in Lens species.

  • Identify any associations between genomic features known to be linked to genome stability/instability—such as centromeres, repetitive sequences, and cytosine methylation patterns—and the observed structural changes.

Experiment Summary

Lens species are proposed to exhibit increased levels of genome structural variation both within and between species, which has implications for plant breeding and evolutionary genomics. Plant centromeres are key regions involved in chromosomal stability, where structural changes can be linked to meiotic abnormalities, resulting in inversions, translocations, and deletions of chromosomal segments. We sequenced, assembled, and performed gene synteny analysis in the genomes of all seven Lens species using long-read sequencing technologies. We also located, characterized and analyzed centromere synteny in these genomes through ChIP-Seq experiments to better understand their role in Lens chromosomal stability. We performed repeat analysis, including cytosine methylation calls at CpG context, genome size estimation, species divergence analysis, and an immunofluorescence assay using anti-CenH3 antibodies for all genomes and centromeres to test for any associations with genome structural changes.

Germplasm
Germplasm Genus
Lens
Germplasm Scientific Name
  • Lens culinaris
  • Lens orientalis
  • Lens tomentosus
  • Lens odemensis
  • Lens lamottei
  • Lens ervoides
  • Lens nigricans
Experimental Design
Biochemical Technique
  • Chromatin Immunoprecipitation Sequencing (ChIP-Seq)
  • Immunofluorescence assay using anti-CenH3 antibodies
  • Flow Cytometry to Estimate Genome Size
Phenomics
Methodology
  • Pseudomolecule assembly of genomes

    PacBio HiFi reads greater than 10kb at 36x coverage were used to assemble CDC Greenstar (Lcu.1GRN). The contigs were scaffolded using an optical map and aligned to the Lcu.2RBY reference assembly using mummer v3.23 in order to determine pseudomolecule placement. Assembly of Ler.1DRT is documented by Ramsay et al. (2021). All other wild assemblies used the GenomeFocus workflow.

    The following table lists the genome assemblies that were assembled and the technologies that were used for sequencing. You can click on the name of a specific genome assembly to go to its genome assembly page and learn more.

    Table 1. Sequencing Data
    Genome AssemblyPacBioNanoporeHi-CIlluminaChIP-Seq
    Lcu.1GRNx   x
    Ler.1DRT xxxx
    Lla.1ESP xxxx
    Lod.1TUR xxxx
    Lor.1WPS xxxx
    Lni.1VIC xxxx
    Lto.1BIG xxxx
  • Analysis of Repeats

    For each genome, the EDTA pipeline v1.9.4 was used to identify various classes of transposable elements including long terminal repeats. TEsorter was used to determine superfamily-level classification of the annotated sequences, by searching against the REXdb database. The resulting repeat library was then used to mask the genome using RepeatMasker v4.1.6.

  • Gene Annotation

    Annotation of genes combined both ab initio and evidence-based methods. Using the BRAKER2 pipeline, protein sequences from L. culinaris and L. ervoides and RNA-Seq data were used to annotate protein-coding genes on the soft-masked genomes. The gen models were further refined using PASA v2.5.3 by incorporating transcript alignment evidence. BLAST+ and InterProScan searches against UniProt and Pfam databases completed the functional annotation of the predicted genes.

  • Synteny Analysis

    Synteny analysis was done with a customized workflow in Python (https://github.com/mcbbaker/genome-synteny). For each genome, an All-by-All BLAST is performed using annotated protein sequences. Genes with repeat-related functional annotation are removed and then provided to MCScanX using default parameters. The output is a collinearity file that was then filtered by score for each genome pair and visualized using Synvisio.

  • Methylation Calls

    For all genomes except for Lcu.1GRN, raw signal data from Oxford Nanopore Sequences were aligned to its corresponding genome sequence using minimap2. To identify methylated (5mC) and unmethylated cytosines in the CpG context, these alignments were given to Nanopolish's call-methylation submodule.

    For Lcu.1GRN, pbmm2 was used to align raw PacBio reads against the genome (using parameters --sort -j 100). The aligned reads were then given to Pb-CpG-tools to perform methylation (5mC) calling in the CpG context.

  • Species Divergence Analysis

    To estimate the level of synonymous substitution (Ks) between paralogous and orthologous gene pairs between Lens species, the python script Dated was utilized to perform the maximum likelihood method, implemented in the PAML package, on these sequences. Mutation rates for each Lens species were estimated based on average Ks values associated with a shared whole-genome duplicated event (dated 56.5 million years ago). Divergence times between species was then calculated using the mean mutation rate. Finally, a Gaussian mixture model analysis was performed to identify the major peaks in Ks distributions.

  • Genome Size Estimation

    Illumina short-read sequences were filtered for 40x coverage and processed with Jellyfish v2.2.7 to generate a 21-mer frequency distribution. GenomeScope was used to analyze the resulting K-Mer histograms in order to estimate genome size.

  • Centromere Characterization

    Illumina short-reads from ChIP-Seq were pre-processed using Trimmomatic to remove adapters and low-quality reads, then mapped to each species’ genome using Bowtie2 and SAMtools. A custom Python script generated by Petr Novak (https://github.com/kavonrtep) processes BAM files generated from Bowtie2 using input (whole DNA) and chipped (centromeric reads) data. The script employs three software tools— Epic2, MACS3, and bamCompare —to identify read enrichment and ChIP-Seq peaks. Visualization was done with IGV and JBrowse2.

    ChIP-Seq validation and centromere composition analysis were performed using RepeatExplorer2 with the ChIP-Seq mapper tool on a Galaxy server, with enrichment ratios used to identify centromeric repeat clusters (≥5× enrichment). Satellite sequence similarity was assessed using MAFFT and Muscle, with identity calculated via P-distance. TideCluster was used for satellite mapping, and Bedtools helped analyze centromere synteny, transposable elements, and methylation.

    Previously developed genetic maps (LR-26 and LR-100) supported centromere repositioning for specific chromosome combinations using JBrowse2 and Synvisio.

Assay Protocol
  • Flow Cytometry to Estimate Genome Size

    Approximately 1 cm² of young leaves were harvested from all species excluding L. nigricans. These were finely chopped using a razor blade in 100 µL of Extraction Buffer (CyStain UV precise P automate nuclei extraction and DNA staining kit, Sysmex) to prepare a nuclei suspension. For the endogenous control (Raphanus sativus), young leaves measuring 10 cm² were similarly chopped in 6 mL of Extraction Buffer to produce a nuclei suspension.

    Each sample was filtered through a CellTrics™ filter equipped with a 30 mm nylon mesh. For DNA size analysis, the control samples were combined with 150 µL of DAPI DNA staining solution, 20 µL of the control nuclei suspension, and 80 µL of the sample nuclei suspension.

    Measurements were performed using a CytoFlex cytometer (Beckman Coulter, Brea, CA), with each sample analyzed in three repetitions on two separate days.

  • Immunofluorescence

    Seeds from all Lens species were germinated in darkness and collected for nuclei and chromosome slide preparation, following the protocol described by Oliveira et al. (2024). Root tissues were fixed in a 3% formaldehyde solution prepared in Tris buffer (10 mM Tris, 10 mM Na₂EDTA, 100 mM NaCl) for 25 minutes, with the initial 5 minutes under vacuum. After fixation, the roots were rinsed in Tris buffer for 30 minutes on ice. The fixed roots were then prepared for nuclei and chromosome suspensions. The tissue was homogenized using an Ultra-Turrax T8 homogenizer in a solution containing: 15 mM Tris, 2 mM Na₂EDTA, 80 mM KCl, 20 mM NaCl, 0.5 mM spermine, 15 mM mercaptoethanol, and 0.1% Triton X-100.

    For primary antibody detection, a custom rabbit antibody targeting CenH3 (P45) and a mouse-derived anti-α-tubulin antibody (Neumann et al., 2015) were used at dilutions of 1:1000 and 1:100, respectively, in 1× PBS containing 0.1% Tween. Slides were incubated with both antibodies overnight at 4 °C. Post-incubation, slides were washed twice in 1× PBS for 5 minutes each, followed by a single 5-minute wash in 1× PBS with 0.1% Tween. For secondary antibody detection, Rhodamine-red-conjugated anti-rabbit and Alexa488-conjugated anti-mouse antibodies were applied and incubated in the dark at room temperature for at least one hour.

    Slides were then washed again with 1× PBS and 1× PBS containing 0.1% Tween, fixed in 4% formaldehyde in 1× PBS for 10 minutes at room temperature, and subjected to a final round of washes. Finally, slides were mounted with DAPI (4′,6-diamidino-2-phenylindole DNA stain) for visualization. Images were captured using a ZEISS Axio Imager Z1 epifluorescence microscope and processed with Adobe Photoshop version 24.1.1.

  • Chromatin immunoprecipitation sequencing (ChIP-Seq)

    Intact nuclei were extracted from fresh leaves of all Lens species to optimize chromatin digestion using Micrococcal nuclease, targeting mononucleosome-sized fragments (~147 bp). Digestion conditions such as enzyme concentration and duration were adjusted per species.

    Post-digestion, chromatin was split into two fractions: one retained as whole-genome DNA, and the other used for immunoprecipitation of centromeric DNA with the CenH3-2 antibody (P45; Neumann et al., 2015). A standard antibody concentration of 15 μg/mL was used, but for L. tomentosus, L. lamottei, L. ervoides, and L. nigricans, a second round of precipitation was performed using 3 μg/mL due to unexpectedly high DNA yield.

    Both input and immunoprecipitated (chipped) DNA samples from all species were sent for Illumina paired-end sequencing (150 bp reads) at Admera Health, New Jersey, USA.

Attribution
Researchers
Data Collector
Data Custodian
Kirstin E Bett
Data Curator
Associated Datasets
Dataset
Funding
Research Organization
Funding Grant
Title
EVOLVES: Enhancing the Value of Lentil Variation for Ecosystem Survival
Data Custodian
  • Kirstin E Bett
  • Albert Vandenberg
Research Organization
Funding Range

2019-2023