Back to Programs

Science and medicine

Apertium

A free/open-source machine translation platform

c++javascriptpythonshellxml

Participation history

8 GSoC years

2024

3 projects

Official year page

Dictionary Induction from Parallel Corpora

The aim is to construct bidirectional dictionaries for a language pair, given a pair of parallel corpora - i.e., the same content in two different...

Capitalization Handling Module for es-pt

This project aims to add the Capitalization Handling Module for the es-pt language pair, as well as to create more rules expanding the tool. After...

Spell-Checking Interface for Apertium's Web Tools

Developing a spell-checking interface for Apertium's web tools to enhance user experience and accessibility. The project aims to update the existing...

2023

5 projects

Official year page

Internationalization of Apertium Tools

This project aims to internationalize Apertium tools so that they can be localized easily to other languages, which makes usage of Apertium tools...

Leveraging Morphological Data from Linguistic Software Tools for Computational Resource Generation

This proposal aims to leverage the language documentation data compiled by linguists in popular fieldwork software tools for extraction of...

Tokenization for spaceless orthographies in Japanese

Investigating the suitable tokenizer for east/south Asian languages which usually do not use spaces and implementing it. Besides, improving...

Develop a morphological analyser

I will be working on creating a morphological dictionary for Kumaoni language and then use it for implementing a morphological analyzer. There are...

Develop a language pair for Highland Puebla Nahuatl (azz) and Western Sierra Puebla Nahuatl (nhi)

My idea is to develop a language pair for Highland Puebla Nahuatl (`azz`) and Western Sierra Puebla Nahuatl (`nhi`). Both are endangered variants of...

2021

10 projects

Official year page

Adopting the Hindi-Bengali language pair (unreleased language pair).

In this project, I aim to create a hin-ben repository in Apertium that also includes the task of creating/expanding the transfer rules, creating the...

Apertium Browser Plugin

My project has been to develop the Apertium Browser Plugin. The previous Geriaoueg plugin is out of date, with the official link given in the wiki...

Adopt an unreleased language pair, Hindi-Bhojpuri

I plan on developing the Bhojpur-Hindi language pair in both directions i.e. bho-hin and hin-bho. This will involve building a monolingual...

Implementing new language pair: Kazakh - Uzbek

Having seen the benefits of the open-source Rule-Based Machine Translation platform - Apertium as an alternative to other free/commercial online...

Ideas for Google Summer of Code/Morphological analyser

• Creating a high-accuracy morphological analyser for Ibo by contributing to the currently existing one; • Increasing WER on the eng-ibo pair...

User friendly lexical training

The procedure for lexical selection training is a bit messy, with various scripts involved that require lots of manual tweaking, and many third party...

Unipertium

3 mostly unrelated smaller projects that all happen to start with "uni": UNIcode, UNIt testing, and UNIversal dependencies transfer (the latter being...

A morphological analyzer for Bagvalal

Bagvalal is an endangered typologically rare Caucasian language from the Nakh-Daghestanian family. Its conservation and study are constrained by the...

Finnish, Olonets-Karelian and Karelian lexicon development

The three languages that this application targets are closely related Balto-Finnic languages spoken in geographical proximity to one another. Finnish...

Develop a prototype MT system for a strategic language pair uzb->kaa

In this project I'm going to continue developing translation pair Uzb-Kaa languages. In the list of different pairs of Turkic languages, I analyzed...

2020

7 projects

Official year page

Bilingual Dictionary Discovery via Graph Exploration

A crucial step in developing a language pair is writing its bilingual dictionary, which maps a lemma X in language A to a lemma Y in language B if X...

Adopt an unreleased language pair : Hindi-Punjabi

I plan on developing the Hindi-Punjabi language pair in both directions i.e. hin-pan and pan-hin. This'll involve improving the monolingual...

State-of-the-art Morphological Analayser for Uzbek language and improved language pairs uz-kk, uz-ky, uz-tr.

Creating the State-of-the-art HFST-based Morphological Analayser for Uzbek language, contributing on the Karakapak and Uyghur Morphological...

Adopting the French-Arpitan language pair

I propose to create a bidirectional French-Arpitan translator. Arpitan (often called Franco-Provençal) is an endangered and heavily under-resourced...

Adopting an unreleased language pair of Uzb-> Kaa

In this project I am going to create a new language pair uzb-kaa. Last year I have helped with translations to GSoC 2019 student,as I am native...

Extending Ve’rdd for Apertium Needs

This proposal is targeted to the task named "A Web Interface to expanding dictionary lemmas integrate with GitLab/GitHub". As I have already...

Modifying the apertium stream format and solving the markup reordering problem using wordbound blanks

Markup handling has been a problem in Apertium for a long time. It was done using superblanks that encapsulate markup information inside them during...

2019

10 projects

Official year page

Recursive Transfer

Build a GLR parser-generator as an alternative to the current chunking system to better support long-distance phrasal reordering.

Anaphora Resolution

Anaphora resolution is the problem of resolving references to earlier items in the discourse. This most commonly appears as pronoun resolution where...

Develop a releasable Uzbek-Qaraqalpaq translation pair

In this project I am going to create a new translation pair between Uzbek and Qaraqalpaq. There is no other single translator between these two...

Improve/Extend weighted transfer rules module

Ambiguous patterns are ones that more than one transfer rule could be applied to. Apertium resolves this ambiguity by applying the left-to-right...

Turkic MT improvements

Refining four Turkic MTs: uig-tur, kyr-tur, uzb-tur and tat-tur

Python API/library for Apertium

Apertium is a free/open-source rule-based machine translation platform implemented in C++. Right now, the project is calling Apertium binaries as...

English-Lingala language pair

An ‘English-Lingala’ language pair using Apertium rule-based machine translation system.

Unsupervised weighting of automata

Finite state automata/ transducers are currently used in lots of application including machine translation. One of the most challenging parts of...

Improvement of Annotatrix project

Bug fixes and feature implementations for Annotatrix tool

Improving the Catalan-Italian and Catalan-Portuguese language pairs

In this project there are two major goals: 1) improving the existing translators from Italian to Catalan, from Portuguese to Catalan and from Catalan...

2018

11 projects

Official year page

Fra-oci/oci-fra translator

I intend to work on a French-Occitan translation pair in order to provide a new translator, which will be useful first to the Occitan community but...

Adoption of Guarani - Spanish pair

Guarani is one of the most widely spread indigenous languages of southern South America. It is spoken by 6 million people in Paraguay (where it is...

Kannada-Marathi language translation

I am adding a new language pair (Kannda-Marathi) to Apertium.

Apertium translation pair for Kazakh and Sakha

I would like to develop Apertium translation pair for Kazakh and Sakha languages. It would benefit society in whole by keeping diversity supporting...

Bilingual dictionary enrichment via graph completion

Graph representation is very promising because it represents a philosophical model of a metalanguage knowledge. Knowing several languages, I know...

Adopting the unreleased Romanian-Catalan pair and upgrading other pairs to the monolingual module system

Currently there are no machine translation systems offering direct translation between Romanian and Catalan available to the general public. English...

Uyghur-Turkish MT

An MT for the closely related Turkic languages, Uyghur of the Karluk branch and Turkish of the Oghuz branch.

Improving language pairs by mining MediaWiki Content Translation postedits

The purpose of this proposal is to create a toolbox for automatic improvement of lexical component of a language pair. This toolbox might become a...

Tatar and Bashkir: developing a language pair

The tat-bak language pair already exists in Apertium, but is now in the nursery state. The aim of my project is to develop this language pair, fill...

Extend lttoolbox to have the power of HFST

The aim of this project is to implement the support for morphographemics and weights in the lttoolbox transducer. The proposal focuses on extending...

UD-Annotatrix

This project aims to extend the functionality of the UD Annotatrix tool. This tool allows researchers to annotate universal dependency trees right...

2017

10 projects

Official year page

UD-annotatrix

The aim of my project is to create an easy-to-use, quick and interactive interface tool for UD annotation based on the existing Apertium project. The...

Implementing a shallow syntactic function labeller

In many pairs it is useful to know in addition to the morphological tags of a word, syntactic function tags in order to make an adequate translation....

Automatic blank handeling

Our current handling of formatting/markup (HTML, odt, docx, latex) is brittle, requiring transfer rules to explicitly deal with blanks (e.g. markup),...

Development of the Czech to Russian Language Pair

I plan to assist in the development of the Czech to Russian language pair in order to bring the Czech to Russian translation capabilities to release...

Improvements to the Apertium Website Interface

Apertium is a free/open-source platform for rule-based machine translation and language technology which is aimed providing support for...

Crimean Tatar-Turkish MT

Creating a new rule based translation pair between Crimean Tatar and Turkish. This involves disambiguation, transfer and lexical selection.

Discontiguous Multiwords

Discontiguous multiwords are multi-word expressions that are separated by something in the middle (e.g. "take the garbage out") . Apertium currently...

Chukchi morphological analyser using HFST

Chukchi is a language with rich and complicated morphology and incorporation. By now morphological parsers using regular expressions were not able to...

“Proposal apertium cat-srd and ita-srd”

Catalan to Sardinian (apertium-cat-srd): The project has already started and currently the bidix is in the Staging section. The Catalan language is...

Adopting English-Catalan language pair to bring it close to state-of-the-art quality

Apertium currently has an English-Catalan language pair in trunk, but there is a lot of room for improvement. One of the existing problems is the use...

2016

11 projects

Official year page

Apertium website improvements

New features provide benefits both to Apertium users and Apertium team. Apertium website users will get the improved tool which provides a new...

Investigation of new ways to combining Constraint-grammar and apertium-tagger & a new averaged perception based tagger

Many less resourced languages don’t have a tagged corpus or only have a small amount of poorly tagged material. In such cases unsupervised learning...

Sardu, abbarra vivu!

The project I intend to carry out is the creation of a MT engine for the language pair Italian-Sardinian based on the Apertium platform. As pointed...

Kurmanji (Kurdish)-English MT

I propose to work on the Kurmanji-English language pair, with the aim of improving it to the state of the art level, in terms of coverage and...

Project: Adopting a language pair

Currently the pol-rus language pair is in the beginning state (in the incubator). There are very few words in the bilingual dictionary and no rules....

Adopt an unreleased Kazakh-English language pair

These days the translating text automatically by using machine translation is very important, because it helps people from whole world to understand...

Lint For Apertium

My draft proposal for the Lint for Apertium project.

Automatic Blank Handling

Our current format handling is brittle, requiring transfer rules to explicitly deal with blanks, and some times inevitably outputting them in the...

Machine Translation for Sicilian-Spanish Language Pair

The project goal is to create a machine translation package for Sicilian-Spanish language pair on the base of Apertium’s machine translation system....

New Belarusian-Russian language pair

I propose to create Belarusian-Russian language pair.

Apertium Weighted Transfer Rules

The idea of the task is to implement a mechanism of resolving transfer rule conflicts using previously obtained rule weights. The tool for obtaining...