Back to Programs

Science and medicine

Genome Assembly and Annotation

Unleashing the potential of big data in biology

databasedeep learningdevopsdockermachine learningmysqlnextflowpythonpytorch

Participation history

4 GSoC years

2026

7 projects

Official year page

Expose a Subset of ENA REST Services as MCP

The goal of this project is to provide a Model Context Protocol (MCP) server that is suitable for production and exposes expected REST endpoints from...

Building a Perturbation-Aware LLM for Multimodal In Silico Perturbation Modelling

Perturbation biology datasets across CRISPR screens, MAVE, and scPerturb-seq remain siloed in incompatible formats, making cross-modal reasoning...

Browser-Native Genomic Feature Search using SQLite WASM and JBrowse

This project implements a browser-native genomic feature search system for efficiently querying large GFF3 annotation datasets without requiring...

Annotation Metrics Reporting and Analysis Modules for the Ensembl Assembly/Annotation Tracking App

The Ensembl Assembly/Annotation tracking application stores rich quality metrics for thousands of genome annotations but currently lacks tooling to...

Expand genome metadata in Ensembl with AI tools

The Ensembl Plants and Metazoa platforms face a significant metadata gap where critical biological context, such as ploidy, strain, and sex, is...

Sequence similarity networks for the visualisation and exploration of MGnify Proteins

My proposal is to develop a scalable computational pipeline to construct, annotate, and visualise Sequence Similarity Networks (SSNs) for a...

Ask VEPai. Trained chatbot interface for Ensembl VEP web

Ensembl VEP's web interface offers dozens of configuration options for variant annotation, overwhelming new users and generating recurring helpdesk...

2023

5 projects

Official year page

Interactive Visualization for Comparative Metagenomics in MGnify

The project aims to improve the visualisation tools for metagenomics data in the MGnify platform by identifying and using new technologies that can...

Using Deep Learning to Identify Features of Protein-Coding Genes

Accurate gene annotation in eukaryotes solely based on genomic data has been a significant obstacle in biology since the introduction of...

Differentiating Real and Misaligned Introns with Machine Learning

The advancement in the accuracy of long-read sequencing technology has allowed us to explore novel transcript variants of known genes. Preventing...

A Nextflow Pipeline for Repeat Annotation

My proposal is to develop a NextFlow pipeline that will efficiently and accurately perform repeat annotation and masking on large genome sequences...

Expand the species search functionality for the ensembl beta website (Metazoa).

The objective of this project is to create a standalone Elasticsearch tool that can handle taxonomic-related requests. This tool helps to expand the...

2022

6 projects

Official year page

Accessing Ensembl data with Presto and AWS Athena

The goal of this project is to build a nextgen replacement for the BioMart tool that provides a way to download custom reports of genes, transcripts,...

New FAANG backend with Elasticsearch and GraphQL

Current limitations: The current Back End for the Functional Annotation of Animal Genomes project (FAANG) provides users with a public rest API to...

Using Machine Learning to Identify and Classify Repeat Features

A number of tools exist for identifying repeat features, but it remains a problem that the DNA sequence of some genes can be identified as being a...

Investigating and Implementing Compact Data Representation of Homology Relationship

A key challenge surrounding modern bioinformatics is to manage and store the growing amount of biological data with both space efficiency and...

Extract important information from scientific papers

During GSoC 2021, BioBERT and RegEx/string matching technique based “Named Entity Recognition” (NER) system was developed to recognize and extract...

GSoC 2022 Proposal - Extract text from tables in Scientific Papers by Kshitij Soni

PyTesseract is really helpful, the first time I knew PyTesseract, I directly used it to detect some a short text and the result is satisfying. Then,...

2021

3 projects

Official year page

Deep learning homology inference

Many genes both within and across species share a common origin. Homologoy inference is concerned with disentangling the precise nature of this...

Extract important information from scientific papers

Current limitations with the variant detection using wbtools (and entity extraction in Wormbases’s AFP pipeline) is that it relies on regular...

An orchestration system for MGnify running on distributed heterogeneous compute clusters

MGnify is a freely available online service hosted by the European Bioinformatics Institute (EMBL-EBI). It helps researchers to do exploration and...