Biology, Pre-medical studies
PhD student,
Department of Biology
Professor,
Department of Biology
Chicken (Gallus gallus) is both a key model organism and a species of major agricultural significance, yet their genome annotation remains incomplete due to limited RNA sampling across diverse tissues, developmental stages, and conditions. Transcript isoforms, which are distinct mRNA versions that arise from the same gene via alternative splicing, promoter usage, or polyadenylation. They greatly expand proteomic diversity while playing critical roles in gene regulation. Abnormal isoform expression is implicated in cancer, neurological disorders, and cardiovascular disease, underscoring the importance of accurate isoform characterization. Using bulk RNA-seq data, this study aims to identify novel transcript isoforms in the chicken genome through comprehensive gene and transcript-level analysis. We expect this approach to uncover previously unannotated isoforms, advancing our understanding of gene regulation, phenotypic diversity, and functional complexity in Gallus gallus.
The primary focus of this study is to address several key gaps in the current chicken genome annotation. Specifically, we ask whether there are novel genes in the chicken genome that do not overlap with any annotated features in the reference genome.
This project will use bulk RNA-seq data of 24 chicken samples obtained from Dr. Katia Del Rio-Tsonis' Lab at Miami University. The data will be analyzed to identify and characterize novel transcript isoforms that are not currently annotated in the chicken reference genome. The project will utilize open-source bioinformatics tools and computational techniques specifically designed for RNA-seq analysis. The workflow will begin by using the transcript annotation file that is already created by the PhD student, Hiruni. She generated that file by preprocessing raw RNA-seq data using fastp (v.0.23.4), which trimmed adapters, removed low-quality reads, and generated quality control reports, ensuring data reliability. The cleaned reads were then aligned to the chicken reference genome (GRCg7b, ENSEMBL v.115) using STAR (v.2.7.11b), a splice-aware aligner designed for mapping RNA-Seq reads to the genome. After alignment, the resulting BAM files were processed with StringTie2 (v.2.2.1), which assembled transcripts, estimated expression levels, and identified potential novel isoforms. The 24 output files from Stringtie were merged into a single transcriptome file (merged.gtf), which will serve as a base for this project. Currently, we are focusing on classifying and identifying one specific type of novel genes with novel isoforms from that GTF file, which is called C1 (category 1). C1 includes novel genes with novel transcripts that do not overlap with any known reference gene, even on the antisense strand. These novel genes with transcripts will be extracted using Pybedtools, a python library used for genomic manipulation. The resulting genes will then be validated against the existing annotations in the reference genome (GRCg7b, ENSEMBL v.115) to confirm novelty. We will be using Integrative Genomic Viewer (IGV), a genomic browser, to examine isoform structures and splicing patterns. This computational approach is efficient and appropriate for the project’s goals because it leverages existing RNA-Seq data and open-source tools to uncover new biological insights without requiring additional wet-lab experiments. The methods used will generate reproducible, data-driven results that contribute to improving the chicken genome annotation and understanding transcript diversity. Ultimately, this will allow us to contribute to the scientific community.
Using pybedtools in Python, we identified 462 novel genes, 1,037 novel transcripts, and 4,181 novel exons through systematic comparison of our assembled transcriptome against the reference genome annotation. These findings suggest a substantial degree of previously uncharacterized transcriptional complexity in Gallus gallus, highlighting the limitations of the current genome annotation.
To validate our findings, we visualized a subset of the identified novel genes using the Integrative Genomics Viewer (IGV). Inspection confirmed that these candidate genes did not overlap with any annotated features in the reference genome, supporting their classification as genuinely novel loci rather than artifacts of misassembly or annotation redundancy.
As an additional layer of validation, we performed Blastx analysis on a selection of novel gene sequences, which translates nucleotide sequences and compares them against known protein databases. Several of these novel genes returned significant matches to proteins from other species, providing evidence of evolutionary conservation and functional relevance. The cross-species protein similarity further supports the biological authenticity of these candidates, suggesting that these genes may encode functionally important proteins that have been conserved across vertebrate lineages but were previously unannotated in the chicken genome.
Overall, these results demonstrate that gene-level RNA-seq analysis can substantially expand the known gene repertoire of Gallus gallus. The combination of computational identification, IGV-based visualization, and Blastx homology searching provides a multi-tiered validation framework.
This study demonstrates that comprehensive gene-level RNA-seq analysis can significantly expand the annotated gene repertoire of Gallus gallus. By leveraging pybedtools for systematic comparison against the reference genome, we identified 462 novel genes, 1,037 novel transcripts, and 4,181 novel exons, revealing a substantial degree of transcriptional complexity that remains uncharacterized in the current annotation. Multi-tiered validation through IGV visualization and Blastx homology searches confirmed the biological authenticity of several novel candidates, with cross-species protein matches suggesting functional conservation across vertebrate lineages. These findings underscore the inadequacy of existing chicken genome annotations and highlight the power of isoform-level analysis in uncovering hidden layers of gene regulation and proteomic diversity. Ultimately, this work contributes a valuable foundation for improving genome annotation in Gallus gallus and advancing our understanding of gene expression regulation in both agricultural and biomedical contexts.
Several avenues of research emerge naturally from this work. First, investigating overlapping reference genes that overlap with novel genes represents an important next step, because such genes may reveal complex regulatory relationships and expand our understanding of how transcriptional output is coordinated across genomic loci. Second, de novo transcript assembly, conducted without reliance on the reference genome, could uncover entirely new genes that fall outside the scope of reference-guided approaches, potentially identifying transcribed regions that are absent from current annotations altogether. Finally, experimental validation can be used to confirm the biological relevance of some of the important novel isoforms identified during the project. This includes RT-PCR and qPCR confirmation of top novel isoform candidates to verify their expression at the RNA level, as well as mass spectrometry-based proteomics to determine whether these isoforms can be translated into stable protein products. Together, these efforts will strengthen the functional interpretation of our computational findings and move the field toward a more complete and accurate annotation of the chicken genome.
The following is an image of poster presented at the 2026 Undergraduate Research Forum.
This project was funded by the Undergraduate Research Office of Miami University
[1] S. Wu et al., “Annotations of four high-quality indigenous chicken genomes identify more than one thousand missing genes in subtelomeric regions and micro-chromosomes with high G/C contents,” BMC Genomics, vol. 25, no. 1, May 2024, doi: https://doi.org/10.1186/s12864-024-10316-z.
[2] W. Jiang and L. Chen, “Alternative splicing: Human disease and quantitative analysis from high-throughput sequencing,” Computational and Structural Biotechnology Journal, vol. 19, pp. 183–195, Dec. 2020, doi: https://doi.org/10.1016/j.csbj.2020.12.009.
[3] P. Ren, L. Lu, S. Cai, J. Chen, W. Lin, and F. Han, “Alternative Splicing: A New Cause and Potential Therapeutic Target in Autoimmune Disease,” Frontiers in Immunology, vol. 12, Aug. 2021, doi: https://doi.org/10.3389/fimmu.2021.713540.
[4] Y. Zhang, J. Qian, C. Gu, and Y. Yang, “Alternative splicing and cancer: a systematic review,” Signal Transduction and Targeted Therapy, vol. 6, no. 1, pp. 1–14, Feb. 2021, doi: https://doi.org/10.1038/s41392-021-00486-7.
[5] S. Chen, Y. Zhou, Y. Chen, and J. Gu, “fastp: an ultra-fast all-in-one FASTQ preprocessor,” Bioinformatics, vol. 34, no. 17, pp. i884–i890, Sep. 2018, doi: https://doi.org/10.1093/bioinformatics/bty560.
[6] A. Dobin et al., “STAR: ultrafast universal RNA-seq aligner,” Bioinformatics, vol. 29, no. 1, pp. 15–21, Oct. 2012, doi: https://doi.org/10.1093/bioinformatics/bts635.
[7] S. C. Dyer et al., “Ensembl 2025,” Nucleic Acids Research, vol. 53, no. D1, pp. D948–D957, Dec. 2024, doi: https://doi.org/10.1093/nar/gkae1071.
[8] S. Kovaka, A. V. Zimin, G. M. Pertea, R. Razaghi, S. L. Salzberg, and M. Pertea, “Transcriptome assembly from long-read RNA-seq alignments with StringTie2,” Genome Biology, vol. 20, no. 1, Dec. 2019, doi: https://doi.org/10.1186/s13059-019-1910-1.
[9] R. K. Dale, B. S. Pedersen, and A. R. Quinlan, “Pybedtools: a flexible Python library for manipulating genomic datasets and annotations,” Bioinformatics, vol. 27, no. 24, pp. 3423–3424, Sep. 2011, doi: https://doi.org/10.1093/bioinformatics/btr539.
[10] J. T. Robinson et al., “Integrative genomics viewer,” Nature Biotechnology, vol. 29, no. 1, pp. 24–26, Jan. 2011, doi: https://doi.org/10.1038/nbt.1754.
[11] S. F. Altschul, W. Gish, W. Miller, E. W. Myers, and D. J. Lipman, “Basic local alignment search tool,” Journal of Molecular Biology, vol. 215, no. 3, pp. 403–410, Oct. 1990, doi: https://doi.org/10.1016/S0022-2836(05)80360-2.
Critical Thinking & Problem Solving: Analyzed RNA-seq datasets to identify novel gene isoforms and troubleshoot annotation inconsistencies.
Communication: Presented research findings through posters and reports, translating complex genomic data into accessible insights.
Teamwork: Collaborated with lab members to refine analysis pipelines and interpret results.
Technology: Utilized bioinformatics tools and programming (e.g., RNA-seq analysis software) to process large datasets.
Reference genome data were obtained from ENSEMBL database, and a collaborating laboratory provided RNA-seq data under a data use agreement. Analyses followed institutional research guidelines.