Back to Programs

Data

Red Hen Lab

Research on Multimodal Communication

computer visiondata sciencemachine learningnlppythonai

Participation history

9 GSoC years

2024

12 projects

Official year page

Red Hen TV News Multilingual Chat - LLM

Red Hen Lab has access to a large archive of news transcript, which is perfect for training a foundational large language model about the world. To...

Modeling Wayfinding

Stentor roeselii is a free-living ciliate species of the genus Stentor. It is a unicellular organism but shows remarkable hierarchy of decisions when...

Detection of Intonational Units - an Overhaul of AuToBI

Detection of Intonational Units- An overhaul of AuToBI aims to refurbish the AuToBI system first proposed by Andrew Rosenberg in 2009, with today’s...

Speech and Language Processing for a multimodal corpus of Farsi

A multimodal corpus is a treasure trove of language data. In this project, we’re diving into Farsi-language content from public broadcasts. Through...

Chatty AI

Red Hen Lab focuses on various aspects of linguistics and cognitive science, including Construction Grammar and FrameNet. With a voluminous...

Frame Blending by LLMs

Despite LLMs' proficiency in various tasks, they struggle with incorporating characteristics like frame blending into sentence generation, a concept...

Super Rapid Annotator - Multimodal vision tool to annotate videos with LLaVA framework

This proposal outlines the development of an innovative system designed to enhance video annotation capabilities through the integration of a...

A Model Predictive Control & Deep Q Learning Approach to Wayfinding

For the last 200 years research into human cognition and decision-making has revolved around rational choice theory, which assumes that actors are...

Multimodal Sentiment and Stance Detection

This project aims to develop a unified system for sentiment and stance detection in televised news content. Leveraging text, audio, and video...

Red Hen TV News Multilingual Chat - LLM

Large Language Models serve as the foundation for recent advancements in AI. Leveraging the extensive news transcripts data archive from Red Hen...

Visual-Aware End-to-End Speech Recognition in Noisy Setting

Human communication is multi-modal and includes both the visual and audio cues. Modern technology makes it possible to capture both aspects of...

Quantum Wave Function for Information-Processing

Neural networks are descendants of McCulloch & Pitts' threshold-based mathematical model of a binary neuron, but there is ample evidence that...

2023

5 projects

Official year page

Semantic search in video datasets

I propose to create a semantic multimodal search engine for collections of transcribed and aligned videos using state-of-the-art artificial...

Extraction of Gesture Features

The way humans interact with each other occurs in multimodality. We not only articulate words but also, show them. Expressing different concepts such...

Multi-modal Stance Detection on Television News

With the increased American viewership of cable news stations, it has become important to understand the stance of the news stories, relating to...

Red Hen Anonymizer

The Red Hen Anonymizer is a software that uses deep learning and signal processing techniques to anonymize audio-visual data. The proposed project...

Classification of body-keypose trajectories of gesture co-occurring with time expressions using GNNs

This project aims to improve the previous iteration of the project where the hypothesis of it was the existence of a relation between the body...

2022

12 projects

Official year page

The Émile Mâle pipeline at RedHenLab

A knowledge extraction pipeline for artworks from christian iconography that will be used to create a knowledge graph for Christian Iconography. The...

Machine Detection of Film Edits

Shot Boundary Detection of films is an important task that is mostly performed manually. There are some available tools and existing research which...

Multimodal Image Captioning and a dataset for Christian Art

Christian Iconography refers to the study of identifying the saints in a painting by using the attributes such as a crucifix, pedestal or key. Art...

Improving the Visual Recognition of Aztec Hieroglyphs (Decipherment Tool)

Our aim is to enlarge the data set with added iconographic and hieroglyphic examples, varying the angles on the ones we have and adding examples from...

Pipeline for Multimodal Television Show Segmentation

This proposal proposes a multi-modal multi-phase pipeline to tackle television show segmentation on the Rosenthal videotape collection. The two-stage...

Tagging Sound Effects

The objective is to develop a machine learning model to tag sound effects in streams (like police sirens in a news-stream) of Red Hen’s data. A...

Classification of body keypoint trajectories of gesture co-occurring with time expressions

The way humans interact with each other occurs in multimodality. We not only articulate words, but also we show them. Expressing different concepts...

TV Show Segmentation (Final Stage)

We are already getting segmented shows from the past year's algorithm. We need to focus on improving accuracy and time complexity now. The results...

Simulating Representational Communication in Vervet Monkeys using Agent-Based Simulation

Communication within any species is very crucial to its survival. One of the primitive forms of communication featuring mere alarm calls is claimed...

11. Tools for improving subtitle/caption quality

I am going to use some already proposed NLP methods to implement sub-word based spelling checker, T5 based grammar checker, and merge open source...

Red Hen Rapid Annotator

Extending the Red Hen Rapid Annotator through adding are some new features based on the suggestions I am proposing and the app-users requests, in...

Machine Detection of Film Edits

Automatic film comprehension has recently gained increased attention due to the rapid development of the streaming services and the need to reduce...

2021

12 projects

Official year page

Machine detection of film edits

A film can be fundamentally broken into innumerous shots, placed after one another. These shots are divided by cuts. Film cuts can be broadly divided...

Multimodal TV Show Segmentation

I will continue from last year’s work and improve the clustering algorithm to the in-production code and enhance the previous work. The main problem...

Utilizing Speech-to-Speech Translation to Facilitate a Multilingual Text, Audio, and Video Message Board and Database

We design a simple pipeline for using state-of-the-art speech-to-text, text-to-text, and text-to-speech to create a speech-to-speech translation...

Machine detection of film edits

By using digital video, every day the number of people needs to edit and manipulate video content is to increase. This requires from us to have...

Gesture temporal detection pipeline for news videos

Gesture recognition becomes popular in recent years since it can play an essential role in non-verbal communication, emotion analysis as well as...

Simulating Representational Communication in Vervet Monkeys

Vervet monkeys (Cercopithecus aethiops) are said to give acoustically different alarm calls to different predators, evoking contrasting, seemingly...

Anonymizing Audiovisual Data

The Audiovisual data present with Red Hen is not shareable. The visual data clearly shows the speakers involved, and a person can be recognized with...

Detecting Joint Meaning Construal by Language and Gesture

The project aims to develop a prototype that is capable of meaning construction using multi-modal channels of communication. Specifically, for a...

Depicting of Graphical Communication Systems (GCS) in Aztec/Central Mexican manuscripts with Deep Learning: glyphic visual recognition and deciphering using Keras

Among all the human writing communication systems and inspired by Google Arts & Culture project Fabricius, we propose the creation of a framework to...

Red Hen Rapid Annotator

Continuing the work done on Rapid Annotator 2.0 by Vaibhav Gupta, There were some improvements and new functionalities can be added to the current...

Create a Red Hen OpenDataset for gestures with performance baselines

Red Hen OpenDataset contains annotated gesture videos from talk shows. The project requires to systematize the data for computer science researchers...

Development of a Visual Recognition model for Aztec Hieroglyphs

This project aims to Develop a Visual Recognition model for Aztec Hieroglyphs. Aztec language is pictographic and ideographic photo-writing. It has...

2020

8 projects

Official year page

Understanding Messages to Underrepresented Racial, Ethnic, Gender, and Sexual Groups on Social Media by Democratic Politicians and their Electoral Implications

Social media has continued to proliferate, not only as a space for self-expression, but also for greater communication between politicians and...

Pipeline for posture, and posture and gesture embeddings including addition of query by gesture functionality into vitrivr

This proposal concerns the addition of pose data extraction using OpenPose and the generation of posture and gesture embeddings to Red Hen’s...

Hand gesture detection and recognition in news videos

The goal is to implement a reliable pipeline which can take a raw RGB video as input and output information such as whether there is a hand gesture...

Image and audio clustering

The projects aim to design a system that clusters the images and the audio from the media broadcasts and then re-orders them accordingly in the red...

Red Hen Rapid Annotator

With Red Hen Lab’s Rapid Annotator we try to enable researchers worldwide to annotate large chunks of data in a very short period of time with least...

2020 GSoC Red Hen Lab Proposal for AI Recognizers of Frame Blends, Especially in Conversations About the Future

For the project idea "AI Recognizers of Frame Blends, Especially in Conversations About the Future," Wenyue developed multiple approaches that can...

Multimodal TV Show Segmentation

The main research question or problem is to split the videos into named show and dated show. Identify anchor/show names or recognizes what show it...

Age Group Prediction in TV news

It has been raising industrial and research interest to analyze the age biasing in various video formats. Here I'm proposing a project to make use of...

2019

15 projects

Official year page

Annotating NewsScape with FrameNet 1.7 and Expanding FrameNet with BabelNet and Deep Structured Semantic Models

This project sets out to achieve two goals. The first objective is to update the annotation system for Red Hen’s NewsScape dataset to FrameNet 1.7...

OCR for Chinese, Arabic, Hindi,Urdu,Bengali

For this project I wish to develop OCR for television news in a tri-phased implementation model. The first phase will be consist of successfully...

Multimodal Show Segmentation

The goal of the proposal is to create an algorithm that can automatically find boundaries between TV shows in unannotated recordings and also find...

Feature Recognition in works of Art and Iconographic Artwork Captioning

In this project I propose to include a tool to the Red Hen Lab's Art pipeline based on Gradient Activated Class Maps which can be used for...

Gesture Recognition in works of art

I have worked on building an annotation tool which helps to correct the extracted pose from images and build a pipeline to retrieve the extracted...

Speaker Adapted ASR Pipeline

The aim of the project would be to develop an​ ASR pipeline utilizing the existing news conversation dataset and audio pipeline codebase....

Red Hen Rapid Annotator

With Red Hen Lab’s Rapid Annotator we try to enable researchers worldwide to annotate large chunks of data in a very short period of time with least...

Automatic Speech Recognition for European Language (German)

This project aims to build an ASR pipeline for European Language (German) and it must be built as a Singularity container on the Case HPC and put...

GSoC 2019 | Red Hen Lab OCR

OCR is a very wide application which translates characters in the image to an editable format. OCR on television news shows would recognize any text...

Design and develop an online deep learning course for humanists

This Project goal is to design and develops an online course, to teach deep learning for students in the humanities and social sciences. The course...

Speech Recognition for Indian English & Hindi

This project aims to build an Automated Speech Recognition engine for Indian English and Hindi using Deep Learning( RNN-CTC, TDNN, LDA-MLLT, CNN )....

Cockpit : The Red Hen Monitoring System

This project automates the task of sensing the health of the many Red Hen Lab remote capture stations, which are Raspberry Pi devices, and provides a...

Chinese Pipeline

Red Hen gathers Chinese broadcasts to make data sets for NLP, OCR, audio, and video pipelines. Currently, Red Hen have a preliminary ASR pipeline but...

Semantic Art from Big Data

In the proposal, I described and demonstrated my ideas about the visualization tasks: 1) Clusters of Event Category Over Time 2) Distribution of...

Accountability Classifier from Annotated Data

The objective of this project is to automatically detect types of accountability in news articles when describing crimes such as shootings, using a...

2018

10 projects

Official year page

Multimodal Television Show Segmentation

University and libraries of social science and literature department have a large collection of digitized legacy video recordings but are...

Multilingual Neural Machine Translation System

The aim of this project is to build a single Machine Translation system using Neural Networks (RNNs-LSTMs, GRUs,Bi-LSTMs) to translate between...

Russian Ticker Tape OCR

We are proposing an OCR framework for recognizing ticker text in Russian Videos. We do this by solving two main problems, improving the OCR by...

Multimodal Egocentric Perception (with video, audio, eyetracking data)

Hey, I have been in constant touch with Mehul regarding my project on Multi-modal Egocentric Perception. I have already had a skype meet with him...

Automatic Speech Recognition for Speech-to-Text on Chinese

In this project, a Speech-to-Text conversion engine on Chinese is established, resulting in a working application. There are two leading candidates...

Multi modal Egocentric Perception (with video and eye tracking data)

This project aims to tackle the problem of egocentric activity recognition based on the information available from two modalities which are video and...

Emotion detection and characterization in video using CNN-RNN

This project aims to develop a pipeline for emotion detection using video frames. Specifically, we detect and analyze faces present in the video...

Rapid Annotator

With Red Hen Lab’s Rapid Annotator we try to enable researchers worldwide to annotate large chunks of data in a very short period of time with least...

Chinese Pipeline

This project is roughly divided into three parts: OCR Recognition, which uses existing tools to extract captions from videos to text; Speech...

Arabic Speech Recognition and Dialect Identification

The project proposed aims to implement an Arabic speech recognition model using training data from the MGB-3 Arabic datasets to perform speech...

2017

9 projects

Official year page

Multilingual Corpus Pipeline

This project aims to build a pipeline for a searchable corpus on multiple languages. We will be using NewsScape data for the project and tools like...

Sentiment Analysis of Social Media Data

Whilst crime in general has been falling for decades, hate crime has gone in the other direction. Especially after the US election 2016, it has risen...

Multimodal television show segmentation

I aim to build a general system that detects natural boundaries of TV shows. This task has long been under the realm of manual approach by skilled...

Multimodal Emotion Detection on Videos using CNN-RNN, 3D Convolutions and Audio features

This is a deep learning approach which uses both image and audio modality from the videos to detect emotion and characterize it. It uses a...

Neural Network Models to Study Framing and Echo Chambers in News

An interesting study is to construct a model of the media representations of the world, considering features from social discourse such as crime,...

Audio Visual Speech Recognition System based on Deep Speech

Current Red Hen Lab’s Audio Pipeline can be extended to support speech recognition. This project proposes the development of a deep neural-net speech...

Learning Embeddings for Laughter Categorization

I propose to train a deep neural network to discriminate between various kinds of laughter (giggle, snicker, etc.) A convolutional neural network can...

Large-scale Speaker Recognition System for CNN News

This project aims to build a large-scale speaker recognition system for tagging speakers in CNN news recordings upon the existing Red Hen audio...

Audio embedding space in a MultiTask architecture

Auditory stimuli like music, radio recordings, movie soundtracks or the regular speech are widely used in research. While it is easy for a human to...

2016

5 projects

Official year page

Gesture Recognition Using Machine Learning

Gesture Recognition using template matching, motion history image and machine learning. The project is basically divided into 3 phases involving...

Gesture recognition using multimodal deep learning

Use video and text data of the speaker to recognise gesture of the speaker on TV using LSTMs.

Computer Vision and Machine Leaning Applications on Artwork Images

The proposal is inspired by the idea G on RedHen's GSoC 2016 idea page. The main purpose is to develop models and code helping domain experts to...

To construct Bootstrapping Human Motion Data for Gesture Analysis

The Project aims at detecting the Human gesture with the help of classifiers.The project consists of 1) Database, has the Segmented Frames of...

Gestures, Machine learning and other things

The proposal aims to identify elements of co-speech gestures in a massive data of television news. The steps will include building a flawed data-set,...