The inaugural edition of the "From Theory to Practice" workshop (T2P) took place at the University of Rwanda in Kigali from March 30 to April 1, 2022. This event was held in conjunction with the QLA Doctoral Training School in Data Science and brought together a diverse group of attendees including academics from institutions such as CMU-Africa and the University of Pretoria, industry representatives from companies such as InstaDeep and HENCE, NGO representatives from organizations such as UN Global Pulse and DAIR Institute, as well as young researchers and practitioners. The following is the list of speakers from the first edition of T2P.
Carine Pierrette Mukamakuza | Scholar-in-Residence Faculty-Carnegie Mellon University-Africa
Abstract:
Social recommenders seek to improve prediction accuracy by exploiting social connections between users. But when and how should social connections be used? Is there room for improvement in existing techniques? This study investigates various types of association between two views: the social structure and the rating behavior. It finds that there are associations at three levels: individual, pairwise, and community. The strength of association depends on the specific attributes examined, and is often directional, rather than bi-directional. When looking at communities, analysis shows that there are individuals who are more heavily influenced by the community than others. Based on these results, a social recommender that is as effective as existing techniques but treats users more fairly is proposed.
Chabi Elegbede | Senior Assistant professor | Faculté des Sciences et Technologies (FAST) - Université Nationales des Sciences, Technologies, Ingénierie et Mathématique (UNSTIM) | Bénin.
Abstract:
Background: Traditional methods of analysing social issues focus on statistical and econometric approaches. With the repetition of multiple surveys produced by international institutions around the world, and particularly in Africa, difficulties have arisen in analysing the large volumes of data generated. This is why, before applying traditional methods, it is often interesting to identify the most important variables that explain the problem studied. Machine learning (ML) provides tools for selecting important variables in the analysis of a given problem. In this case, ML is used as a complementary approach to traditional methods. Objective: This work aims to analyse cigarette consumption with a focus on gender differences, using an ML approach followed by econometric analysis.
Data: Cross sectional data were used (MICS, 2018) in which 2,412 men and 10,300 Women. Only cigarette consumption is treated while other tobacco products are excluded.
Methods: To explain the choice to smoke and the number of cigarettes consumed per day when smoking, two variables of interest related to cigarette consumption were considered. Then, two complementary approaches were used successively. The first one is based on a machine learning method that allowed an efficient selection of the main determinants useful for the prediction of our two variables of interest. The second approach is an econometric one that explains the link between each of the two variables of interest and the exogenous variables.
Results: Disparities between genders in cigarette consumption choice in Tunisia have been noted. The most important variables that emerge quite often are whether or not they have health insurance, the number of children, the individuals wealth and their place of residence, which have a positive effect on the smoking phenomenon in Tunisia. Nevertheless, the number of cigarettes smoked per day is not affected by the individuals wealth. Finally, it turns out that the observable individual characteristics explain only 3% of the smoking phenomenon in Tunisia against 97% for the unobservable characteristics.
Conclusion: Individuals characteristics explain only a small part of the smoking phenomenon in Tunisia. For a better understanding of the problem addressed in this article, unobservable characteristics, especially sociological aspects, should be taken into account.
Aurelle Tchagna Kouanou | Assistant Lecturer | College of Technology- University of Buea- Cameroon/ Department of Training, Research, Development and Innovation, InchTech’s Solutions | Yaoundé - Cameroon
Abstract:
To mitigate the spread of the virus responsible for COVID-19, known as SARS-Cov-2, there is an urgent need for massive population testing. Due to the constant shortage of PCR (Polymerase Chain Reaction) test reagents, which are the tests for COVID-19 by excellence, several medical centers have opted for immunological tests to look for the presence of antibodies produced against this virus. However, these tests have a high rate of false positives (positive but actually negative test results) and false negatives (negative but actually positive test results) and are therefore not always reliable. In this paper, we proposed a solution based on data analysis and Machine Learning to detect Covid-19 infections. We based on two clinical data set from literature to perform our algorithm. We preprocessed the datasets to select the best features that mostly influence the target, and we noticed almost all of them are blood parameters. We carried out a comparative study of supervised Machine Learning models, after which we selected the Support Vector Machine (SVM) as the one with the best performance. Optimizing is applied to SVM and we obtained a 99.29% of accuracy, 92.79% of sensitivity and 100% of specificity with data set from Kaggle (https://www.kaggle.com/einsteindata4u/covid19). We did the same work with another dataset taken from the San Raffaele Hospital (https://zenodo.org/record/3886927#.YIluB5AzbMV). Once more, the SVM presented the best performance among other machine learning algorithms: 92.86%, 93.55% and 90.91% for accuracy, sensitivity and specificity respectively. These results are compared and outperformed all literature works that based on these same datasets. We can conclude that our proposed solution is reliable for the COVID-19 tests.
Dr. Abebe Geletu | German Research Chair | AIMS-Rwanda
Abstract:
A machine learning model is as good as the quality of its training data. Uncertainties naturally arise in training dataset, for instance, due to measurement noise, data inconsistency, missing data, data insufficiency, etc. Consequently, in the face of data uncertainty, a machine learning model could have poor performance on previously unknown data, rendering the model unreliable and uncertain. Incorrect predictions and decisions made based-on machine learning models whose behaviors are unpredictable and unreliable may incur serious consequences in practical applications, e.g., in health and mission critical applications. This talk is intended to give an overview on recent probabilistic approaches to designing machine learning models. The main focus will be on strategies for reliability of machine learning models, by suggesting a novel chance constrained optimization approach for designing reliable neural network models.
Arun Shanmuganthan | Hence | Co-Founder
Abstract:
At Hence, we help companies harness high-signal data to enhance the way businesses engage with their knowledge providers. To do this, we have to build a data foundation that enables users and algorithms alike to solve business problems. We spend a lot of time thinking about the right ontology for the problem at-hand. This involves a lot of data design, data engineering and data strategy. From there, we have the foundation to do AI and Machine Learning to solve the problems at hands.
Slyvia Makario | Hepta Analytics | Co-Founder and Head of Business
Abstract:
Skills development in STEM are a very key aspect to the development of the african continent. We have barely scratched the surface to the possibilities with the few professionals that are already in the job market at the moment. Most African governments, still rely on outside contractors or contractors whose systems have existed for far too long and have been optimized to work in certain markets compared to Africa. When it comes to customizing to the African Market, we conterminously have to have trade offs. Why should we have tradeoffs when we can build for the problems that exist today.
Hepta Analytics was born out of the need to build tools and or technologies for Africa, with an African context. Like any startup company, we have gone through all the processes and still feeling like we are not out of the woods yet. When we started off, the idea was solely based on working with civil society to improve public voter mechanisms in Kenya. Several things after that led to us pivoting and focusing on verticals other than political.
We have since worked with Non-governmental organizations focused on improving education outcomes in various countries, supporting governments improve Job outcomes and opportunities for young people in Africa, scaling ways in which FGM is publicized and supporting authorities in rescuing young girls from marriages etc. We have also worked with various private institutions to help them rank various policy tools in their market for their work and several other example.
This has not been without challenges and especially being aware each waking day of the challenges of navigating the regulations and processes withing the startup ecosystem in Africa.
There are times we considered closing down to go work full time jobs to make enough money to have more hands on deck but each moment we felt like things were not working out, one more thing came through. It also comes with the burden of not being sure when you are going to be paid for your work and still, due to the still maturing technology ecosystem in Africa, we still don't have as good policies to protect young companies like ours.
Nevertheless we have navigated that though various ways. These includes participating in government funded incubators to gain advantage and benefit from the tax rebates as well a market linkages and grants. It has also helped that we have had a backing from our former institution, Carnegie Mellon University Africa, that offers stipend as compensation for us to continue building when cash is low or when we feel like it is getting to the end of the road. This, and many other acts of Kindness have mostly kept us afloat during the pandemic which has been our most challenging period as a company.
In addition, we have a very strong team of co-founders who have found ways not to quit by tapping into their networks and seeking help or opportunities to improve themselves.
Vukosi Marivate | Chair of Data Science | University of Pretoria | Masakhane NLP | Deep Learning Indaba
Abstract:
In this talk I will go through some of our recent work in both Local Language NLP as well applications in better fighting misinformation.
Martin Mubangizi | Officer - Data Science and Acting Head of Office | UN Global Pulse Kampala
Abstract:
In this digital age, individuals contribute data in the form of ''digital footprints'' knowingly or unknowingly, as people go about their daily activities, making phone calls, sending mobile money, searching the internet, registering for digital services, interacting with ATMs, and calling into radio stations to contribute to discussions. In addition, the deployment of sensors also generates data, contributing to vast amounts of data available to utilize.
The private sector utilizes this data to create new services and products and improve their customer's experience. In the same vein, the government can benefit from this data to inform its policies and programmes, leading to improved service delivery to its citizens. In my talk, I will discuss case studies in which this data is safely harnessed for development and humanitarian work. I will also highlight such data's potential, challenges, and lessons learnt.
Karim Beguir | Co-Founder & CEO InstaDeep
Abstract:
Presenting concrete use cases of AI innovation from InstaDeep, a deep tech startup from Africa.
Arnu Pretorius | Research Scientist | InstaDeep
Abstract:
In this talk, we will introduce high-level concepts of multi-agent reinforcement learning (MARL) and motivate why MARL might be a useful approach to building large-scale decision-making AI systems. We will look at one particular real-world use case. Finally, we will discuss current developments and research around the scaling and interpretability of MARL systems.
Franck Kalala Mutombo | Associate Professor | University of Lubumbashi
Abstract:
The identification of community structure in network data/graphs continues to attract great interest in several fields. A community is a group of nodes that have a higher likelihood of connecting to each other than to nodes from other communities. Network neuroscience is particularly concerned with this problem considering the key roles communities play in brain processes and functionality. In social networks, communities could represent circles of friends, a group of individuals who pursues the same hobby together, or individuals living in the same neighborhood. Communities also play a particularly important role in our understanding of how specific biological functions are encoded in cellular networks. We review algorithms for detecting community in-network data. These methods suffer from the high computational cost and excessive memory requirements when the very large and heterogeneous network. Embedding techniques are introduced as an alternative.
Collen Farrelly | Senior Data Analyst | Datasembly, Inc
Abstract:
Written language sample are ubiquitous in industry--from product/website feedback to tweets to chatbot interactions to creative writing samples to document collections related to a specific task. The analysis of these writing samples allows us to find user groups for marketing purposes, identify main themes content, track sentiment on product feedback to identify problems, understand a writer's state of mind, understand language evolution, or predict outcomes of interest based on the content and linguistic traits of the writing samples. This talk will include a couple of case studies, along with some general code and mathematical application aspects of the projects discussed.
Ernest Mwebaze | Director | Sunbird AI
Abstract:
The process or urbanization in many developing country cities has led to an explosion of noise pollution. Research indicates a strong link between long-term noise exposure and the risk of a negative health outcome particularly cardiovascular and metabolic defects. While many governments have formed some type of regulation to curb noise pollution, they are faced with the insurmountable task of being able to collect data and detect the type of noise pollution. At Sunbird AI we are attempting to solve this challenge by deploying autonomous noise sensors throughout the city. This talk will focus on the process of noise training data collection and models for predicting different types of noises in a developing world urban city like Kampala in Uganda. I will focus on the training data collection and preprocessing and show some early results from various modeling options we have considered.
Timnit Geberu | Founder | DAIR Institute
Abstract:
I will talk about why I started The Distributed AI Research Institute as an independent community rooted research institute focused on the needs of people in Africa and the African diaspora.
Olivier Kraft | Independent Advisor on Financial Crime Analytics | FNA (Financial Network Analytics)
Abstract:
Network analytics is a rapidly growing area of data science and is used to extract insights from any dataset that can be represented through nodes and edges (including but not limited to social networks). This presentation will focus on a real-world application of network analytics, namely the detection and investigation of anomalies in efforts against financial crime. The presentation will begin with a high-level overview of anti-money laundering and then discuss case studies to illustrate the relevance of network analytics in this context.
Aisha Walcott | IBM Research Africa
Abstract:
Artificial Intelligence (AI) has for some time stoked the creative fires of computer scientists and researchers world-wide -- even before the so-called AI winter. After emerging from the winter, with much improved compute, vast amounts of data, and new techniques, AI has ignited vast possibilities to transform Global Health. Against this backdrop, I will discuss how specific AI methods can be leveraged to support decision-makers and the scientific community with disease-control. Specifically, I will focus on how the Malaria control problem inspired a similar formulation to understand the Covid-19 response.
Franck Kalala Mutombo | Facilitator