John Bagiliko | QLA
Abstract:
Rainfall estimation from satellite data is of great importance, especially in the African rainfed agriculture setting where gauge stations are sparse. However, most of the precipitation estimation products rely on the relationship between cloud-top brightness temperature and actual rainfall, assuming that precipitation originates from convective clouds with cold tops. This leaves most of them underestimating rainfall in areas with warm cloud tops while overestimating rainfall in areas of cold cloud tops. This presentation we will be on fit-for-purpose validation of four satellite/reanalysis rainfall products for nine stations in Zambia.
Eka Arua | QLA | PhD
Abstract:
Non-communicable diseases (NCDs) disproportionately affect low-and-middle income countries (LMICs). Much of the health and socio-economic impact caused yearly by NCD-related mortality and morbidity is preventable through well-understood, cost-effective, and feasible interventions. In resource-constrained settings, knowing where and when to allocate health resources in an efficient manner to maximize their use is necessary. Additionally, understanding why some areas may be more burdened than others is crucial for shaping health policy in combating NCDs. Model-based geostatistics (MBG) methods are well established for mapping and monitoring tropical diseases in LMICs, however, there is a paucity of research on their applications to NCDs within the same resource-constrained settings. Type 2 diabetes mellitus (T2DM) is a major NCD globally and in LMICs and is associated with several adverse health outcomes. The primary aim of my thesis is to understand the spatial variation of T2DM burden in an urban city, Cape Town, South Africa using MBG methods.
The purpose of this talk is to highlight the progress made so far, the challenges faced, and the possible avenues going forward in mapping T2DM in the city of Cape Town.
Abigail Baido | QLA | PhD
Abstract:
In classical survival analysis one focuses on a single event for each individual, describing the occurrence of the event by means of survival curves and hazard rates and analyzing the dependence on covariates by means of regression models. In other situations, more than one type of event is of interest. Such situations, with more than one possible type of event for each individual, may conveniently be described by multistate models. In credit risk analysis, multistate models recognize that "default" is only one of the possible credit rating grades to which the loan contract could migrate over the planning horizon. As such, these models estimate the probability that the borrower’s credit quality deteriorates, including a change to default status.
In this talk, we will discuss the application of multi-state models in predicting transition probabilities for a microfinance institution in Ghana. The models consider the transitions of borrowers among different credit status categories, such as full monthly payment, partial monthly payment, no monthly payment, and default on loan, and can incorporate covariates that may influence the transition probabilities. We compare the performance of various multi-state models, including the Markov chain model and the Cox regression model, in predicting the transition probabilities.
Theonille Mukamana | QLA | PhD
Abstract:
Recently, neural networks techniques have evolved into a powerful tool for dealing with a number of problems for which classical solution approaches reach their limits. This study develops a robust neural network approximation framework and tools for efficient training to solve one of the most challenging problems in applied mathematics: The approximation of solutions to high-dimensional nonlinear Partial Differential Equations arising in quantitative finance via classical PDE approaches suffer from the so-called curse of dimensionality, that is, the computational cost goes up exponentially with the dimensionality. We will investigate deep learning of Backward Stochastic Differential Equations and related variational inequalities in order to overcome this difficulty. The goal of this research project is to provide a framework to assess the strengths and weaknesses
of different neural network architectures for the problem at hand and to provide both the theoretical and computational foundation for efficient training of these networks. We believe this study will have a strong impact as it will help clarify how these new tools can and should be applied in practice in a controlled way.
Peguy Kem Meka | QLA | PhD
Abstract:
Climate change is a big challenge in all continents including Africa. Africa is particularly vulnerable to climate change impacts. Some of these impacts include floods, droughts, and storms which are the main sources of agricultural risk. This study examines the climate pattern in selected African countries (Kenya,
Cameroon, Somalia, Tunisia, Burundi, Rwanda, Madagascar, and Zambia) for a period of 1905 to 2021. To detect climate change patterns, precipitation and
temperature are analyzed using persistent homology, which is the main tool in topological data analysis (TDA) that focuses on qualitative information known as
topological features of data. Results of the analyses show that there are changes in topological features during climate change. The study also found that the first persistence landscape (a functional topological summary of persistent homology) consistently detects climate change. This paper offers the possibility to develop an effective early detection system for climate change based on our approach.
Chanelle Matadah | QLA | PhD
Abstract:
This project examines the application of quantum machine learning techniques to binary classification tasks, specifically focusing on the classification of pulsar data. Pulsars are highly magnetized neutron stars that emit electromagnetic radiation at regular intervals (milli-seconds to seconds), and their classification is a crucial task in astrophysics. Traditional machine-learning approaches for pulsar classification have been limited by the large amounts of data and complex feature engineering required. Quantum machine learning, which utilizes quantum computing to enhance computational power and efficiency, offers a promising solution to these challenges. In recent works, some researchers proposed a novel technique for pulsar classification with quantum neural networks using a single qubit architecture. In this talk, we first present their method, then we describe our current work which would be an extension of their results.
Jeremiah Fadugba | QLA | PhD
Abstract:
Deep Learning models have shown high performance in medical image analysis, specifically for retinal vessel segmentation. However, the current neural network models have not been able to represent the uncertainties that exist in images. The use of Bayesian neural networks has been shown across several studies to estimate the uncertainty of deep learning models. In this study, we briefly introduce uncertainty estimation and evaluate the predictive performance of some uncertainty estimation methods on retinal segmentation. We also show that using the estimates of the network in the loss function does not necessarily improve the segmentation performance. We end the talk by briefly discussing our current approach to get a reliable estimate of the uncertainty in this task. .
Everlyn Chimoto | QLA | PhD
Abstract:
In recent years, neural networks have achieved remarkable performance in language translation for numerous languages. However, African languages have not seen adequate benefits from this progress, despite being spoken by large populations. Insufficient digitized language data is a key barrier to developing effective translation models, leading to most African languages being considered low-resource. This talk aims to explore the techniques used to create neural machine translation models for low-resource African languages. The techniques include transfer learning, active learning, and data pruning, all of which can assist in developing effective machine translation models for low-resource languages.
Sekou Remy | IBM-Research Africa
Abstract:
Global health is a domain which presents significant and pressing large scale challenges, any incremental improvement may have significant benefit at a global scale. In this work we focus on an important challenge in the space, calibration for disease models, and show that Reinforcement Learning (RL) is an effective algorithm for calibration problems at a scale which traditionally applied Bayesian approaches struggle. This work uses synthetic data, so has access to ground truth parameters and it can be seen that RL learns different, arguably better information for different parts of the learning process. These exciting results set the foundation for deeper consideration of RL in this space.
Louis Fendji | University of Ngaoundere
Abstract:
Sustainable agriculture amounts to provide novel solutions to farmers and extension workers that use resources wisely and promoting biodiversity while increasing the yield. Performing sustainable agriculture requires paying attention to the whole lifecycle of crops. Several processes from crop selection to the post-harvest transformation of crops can be improved by leveraging advances in AI, mobile phones, sensors, and drones. But agriculture in Sub-Saharan Africa (SSA) is generally performed in regions experiencing a lack of infrastructure such road, electricity, Internet, in addition to weak purchase power and low literacy level of farmers that constitute a barrier to digital service adoption. Moreover, the lack of good quality data prevents the leverage of AI-based techniques, despite the efforts to produce relevant information on crops such as the PlantVillage project. Beyond developing models and evaluating their performance, AI-based solutions should overcome aforementioned limitations to reach farmers. This talk will attempt to identify gaps and provide ideas to connect the dots to reach farmers.
Ciira Maina | Centre for Data Science and Artificial Intelligence, Dedan Kimathi University of Technology
Abstract:
In numerous applications such as weather monitoring, it is desired to estimate a spatial field by observing it at a discrete set of locations. Often these observations are obtained using sensors and they can be contaminated by noise and anomalous values. In this talk, we describe the use of Gaussian process models for anomaly detection when observations are from a spatial field. Gaussian processes provide a natural way to deal with the spatial correlation of sensor observations. We present a variational Bayesian approach to inference in this setting and demonstrate the performance of the algorithm in detecting blocked rain gauges using data from a large weather station network deployed in Africa.
Jonathan Shock | University of Cape Town
Abstract:
The world is filled with massive amounts of data of agents (be they humans or algorithms), interacting in complex environments. To be able to use these as training datasets for reinforcement learning will open up huge opportunities to advance the abilities of artificial agents. The OG-MARL (Off the Grid Multi-Agent Reinforcement Learning) platform, aims to give the community tools, dataset and benchmarks with which to work together to drive the development of MARL systems which can use these offline datasets instead of having to create bespoke environments of huge cost and complexity.
Sebastian Lee | Imperial College and UCL
Abstract:
One of the major challenges of machine learning and obstacle to more general intelligence is so-called continual learning. This refers to the regime of learning tasks in sequence. Humans are very good continual learners. We can learn X and then learn Y without forgetting X, and in some cases also transfer learned aspects of X to Y. On the other hand artificial neural networks suffer from catastrophic forgetting, where information relevant to previously learned tasks is overwritten by the information required to effectively learn later tasks. In this talk I will discuss recent work trying to understand how the similarity of tasks in a continual learning setting affects the amount of forgetting observed in artificial networks.
Sahar Vahdati | Institutes Fur Angewandte Informatik
Abstract: Deep learning approaches have been used very successfully to automatically find appropriate representations of input data in order to solve machine learning tasks. One particularly relevant, but also challenging, type of input data are knowledge graphs (KGs). Knowledge graphs encode human knowledge, which in turn is often structured according to potentially complex underlying patterns. Currently, most deep learning approaches for representation learning in knowledge graphs are empirically driven. There is a lack of a clear mathematical understanding of how deep learning approaches can capture the complexity of human knowledge. Furthermore, the connection between practical performance and mathematical properties of representation learning approaches is not well researched yet. Building and understanding mathematical foundations for representation learning in knowledge graphs can help to advance Artificial Intelligence in general.
Abstract:
Abstract: The capabilities of generative models have heavily improved in different domains (images, text, graphs, molecules, etc.), partly due to larger training data, sophisticated generative models, and sampling techniques. However, evaluation metrics for generative models largely remain based on simplified quantities or manual inspection, which often limit their adoption in practical settings. In this talk, I will present our recently proposed framework for Multi-level Performance Evaluation of Generative models (MPEGO), which could be employed across different domains. MPEGO aims to quantify generation performance hierarchically, starting from a sub-feature-based low-level evaluation to a global features-based high-level evaluation. MPEGO offers great customizability as the employed features are entirely user-driven and can thus be highly domain/problem-specific while being arbitrarily complex (e.g., outcomes of experimental procedures). MPEGO is validated using multiple generative models across several datasets from the material discovery domain. An ablation study is conducted to study the plausibility of intermediate steps in MPEGO. Results demonstrate that MPEGO provides a flexible, user-driven, and multi-level evaluation framework, with practical insights on the generation quality.
Abstract: It was recently shown that the quantum annealing paradigm could be emulated in a variational Monte Carlo (VMC) framework using autoregressive neural networks [1]. There, only stoquastic driver Hamiltonians were considered. Here, we leverage the fact that the VMC algorithm is inherently sign-problem-free to simulate quantum annealing with non-stoquastic drivers. We show that the variational quantum annealing method is able to capture the dynamics of of non-stoquastic Hamiltonians and can provide an advantage for annealing paths that are hampered by exponentially closing gaps [2].
[1] M. Hibat-Allah, E. M. Inack, R. Wiersema, R. G. Melko, J. Carrasquilla, Nature Machine Intelligence volume 3, pages 952–961 (2021) [2] Nishimori and Takada,Frontiers in ICT 4,2 (2017)
Bubacarr Bah | Medical Research Council Unit The Gambia
Abstract:
Testing populations using for instance Quantitative Polymerase Chain Reaction (qPCR) testing is a key component in controlling an epidemic. The testing can be significantly fastened by pooling samples together. This pool testing has its theoretical foundations in Group testing (GT). There are limitations in sampling rates and performance of algorithms in GT that the series of works stated below, attempt to improve through compressed sensing (CS) and machine learning (ML) approaches. Firstly, we propose a CS-based testing approach with a practical measurement design and a tuning-free and noise-robust algorithm for detecting infected persons. Due to nonnegativity of virus loads and an appropriate noise model, the compressed sensing problem can be solved with the non-negative least absolute deviation regression (NNLAD) algorithm. NNLAD requires the same number of tests as current state of the art GT methods and provably detects a small number of infected persons among a possibly large number of people. Moreover, CS approaches also allow to recover the viral load, but the underlying noise models are not common in CS and not well understood. Therefore, in our second improvement work, we study hybrid approaches that combine methods from CS and GT on various noise models. We compare the performance of suchapproaches with classical decoders from CS and GT. Our results show that combined strategies can improve the error rates and provide viral load estimation. These combined strategies exposed the strengths and limitations of each of these methods. In our final improvement attempt, we sought to exploit the strengths and overcome the limitations of these approaches by developing a decoding pipeline which terminates with a ML decoder using various noise models and their combinations. This framework allows for the ML decoder to learn a data-driven non-linear decision boundaries that improves upon the GT and CS predictions. COVID-19 pandemic was used as a case study, but these results will apply generally to all epidemics with similar characteristics to COVID.
Marcel Atemkeng | Rhodes University
Abstract:
This talk will address the data format and archiving/compession issues for the Square Kilometre Array (SKA). One of the major contributors to the large data volumes is the maximum baseline of an array since it sets how well the data needs to be sampled to avoid smearing and decorrelation of the signal. Now, the sampling rate is dependent on the baseline length as shorter baselines can be sampled a lot more coarsely which leads to smaller data volumes. In fact, baseline-dependent averaging (BDA) is an established technique for compressing radio interferometric data, however, this technique results in irregularly sampled data which while supported by the database format requires the data to be restructured in ways that reduce performance when processing. But more importantly, radio-astronomy software tools implicitly assume uniformly sampled datasets. Recent research in M. Atemkeng, in prep. in collaboration with the Wits University and SARAO; shows an approach that uses local low-rank matrix approximation to achieve similar compression rates while significantly minimizing smearing.
Samuel Maina | Microsoft African Research Institute
Abstract:
Joint modeling of densities of random variables given a set of data is one of the core tasks in machine learning and statistics. Despite its crucial role in understanding dependency structures, classical correlation is constrained by its Gaussian assumption. Copula functions are being applied in generating synthetically augmented datasets where generative models have limitations in capturing tail dependencies. Copula functions have been applied to machine learning forecasting algorithms and their performance evaluated using several well-known datasets and compared against widely used methods giving competitive results. In this paper, we examine the potential of bivariate and vine copulas to improve weather forecasting. We demonstrate that the dependency structure between weather variables exhibits tail dependency as well as symmetry, and that different copula functions are best able to represent these relationships.
Younes Boutaib | University of Liverpool
Abstract:
We investigate the role of stochasticity in linear recurrent neural networks (RNN). Time-allowing, I will discuss one or both of the following aspects:
A) The role of the noise in the computations in guaranteeing the robustness of learning (Based on a joint work with W. Bartolomeus, S. Nestler & H. Rauhut).
B) The role of the random choice of the connectivity matrix of the RNN in guaranteeing its separation capacity (based on an ongoing work).
Oladimeji Samuel Sowole | AIMS- Senegal
Abstract:
In demonstrating the potential of AI in addressing key societal issues in Africa by providing reliable and accurate information and resources to the public, the Run Am Mobile App was developed by my team, a platform that utilizes artificial intelligence to promote fact-based and authentic news and encourage investigative and responsible journalism, as well as activate citizen agencies for change and encourage the active engagement of citizens in issues of accountability and anti-corruption. With features such as natural language processing for news source verification and image verification using the Google Reverse Image Search API, the app is well-equipped to address these issues.
In particular, the News and News Source Verification feature allows users to easily identify credible sources of information as regards the 2023 general election in Nigeria, which is crucial in addressing issues related to misinformation and disinformation. The Photo/Image Verification feature helps users determine the authenticity of images, which is important for combating the spread of fake news and propaganda.
Additionally, by providing voter education resources based on laws and regulations from INEC, the app empowers users to make informed decisions and engage in meaningful political discourse, which can contribute to the development of a more accountable and transparent political system.
Overall, the Run Am Mobile App is a valuable tool that demonstrates the potential of AI in addressing key societal issues in Africa by providing reliable and accurate information and resources to users.
Abstract:
Digital and data-driven agriculture using Artificial Intelligence (AI), Blockchain, big data analytics, remote sensing and Internet of Things (IoTs) in farming industry necessitates the integration of various tools across the entire agricultural production system, as well as the use of enhanced systems-based techniques (i.e. a System of Systems based approach (SoS)) to ensure that it is scalable, adaptive, and sustainable. Water scarcity, pests and diseases, climatic uncertainty, natural disasters, and labor shortages are all putting strain on crop and livestock agricultural production systems in arid and semi-arid African countries. AI can accelerate the process of creating region specific crops and livestock by harnessing the tools of machine learning to mine existing and new geospatial and temporal, genotypic and phenotypic datasets to analyze relationships and predict outcomes. AI and remote sensing can speed up the process of developing region-specific crops by using machine learning methods to mine existing and new geographic and temporal, genotypic and phenotypic datasets to assess relationships and forecast outcomes. Real-time adaptive decision support systems that incorporate several data streams are required to optimize operations ranging from analyzing animal health to planning future pasture usage. Furthermore, developing AI for large-scale applications that can provide better short- and long-term management and planning for regional water and crop management is crucial.
This paper provides a data-driven techniques for improving agriculture productivity of small-holder farmers in Africa. We provide our experience of developing digital agriculture platforms, climate-smart agriculture and farming as a business approach. We present various techniques for improving food security and productivity of agriculture using key technologies including: (a) Internet of Things (IoT) based agricultural systems, (b) blockchain for food supply chain monitoring and management, (c) AI/ML for increasing agricultural productivity, and (d) remote sensing techniques for digital agriculture (e.g., imagery treatment and prediction in mapping).
Abstract:
We present COVID-19 vaccination models represented by a system of first order non-autonomous differential equations. In this talk, we present a qualitative analysis of the model including sensitivity parameter estimates. However, we will see that the epidemiological parameters of the model are better estimated via neural network. We present an approach that combines residual neural network with variants of recurrent neural network and analyze them for reliable and accurate prediction of daily cases. The data-driven simulations reveal several details and strategies that can be effected in vaccinating Ghana.
Abstract:
A recent report issued by the WWF states that there has been a catastrophic decline in wildlife population in recent years. A large number of species are threatened with extinction due to a number of factors such as over-exploitation of resources, deforestation and climate change. Certain species have been placed on the IUCN Red List for several years, but further conservation efforts are still urgently required to ensure the survival of the remaining individuals. While it is true that the number of individuals in threatened populations is decreasing, there have been considerable conservation efforts to put a halt to this. Ecologists, researchers, and rangers closely monitor the populations, in certain cases by placing microphones and camera-traps into the environment and searching through the recorded audio/images for the species of interest. This non-invasive approach comes with a cost in that huge datasets of audio or images are produced and is difficult to manually process. This talk will discuss efforts in using machine learning to monitor certain species through the use of machine learning. The theme of this colloquium is: "Helping Africa Listen to Herself" where I will showcase some of our latest efforts as well as efforts by my students in producing machine learning models for conservation ecology.
Abstract:
This article draws on the thinking about trust in African scholarship to describe the exact problems Black Box clinical AI generates in health professional-patient relationships. Notably, under the assumption of a Black Box problem, the view of trust as inherently relational implies that health professionals would be unable to explain whether and how a clinical AI incorporates a patient’s values or leverages the same (in its outputs) to honour fiduciary relations. Additionally, the African view of trust as experienced-based and accepting responsibility implies that health professionals can neither be held accountable for Black Box clinical AI outputs that they can hardly understand nor provide material information (concerning what the clinical AI does and why). Finally, given the understanding of trust as a normative concept, health professionals cannot accept patients’ vulnerabilities, and patients cannot give the same. Given that trust will play a vital role in the global acceptance of clinical AI, more studies are required to research how Black Box problem will challenge the relationship of trust.
Abstract:
Foundation Models are a family of models that use unsupervised objectives to learn representations of data units. These models demand re training on mammoth volumes of data at scale and are adaptable to a wide range of downstream tasks. Research and Industry has already experienced tremendous and burgeoning development in the use of these models for natural language and image processing e.g., with the models such as BERT, LaMDA, DALL-E , and GPT already showing viable products in the public domain. Nonetheless, the range of application possibilities for these models is enormous and can only be viewed as unexplored. Research begins to extend these models with concepts from physics and thermodynamics for application in healthcare, and discovery science - drug discovery, molecule discovery etc. We take a look into what foundation models are and their background but provide a brief peek into the applications.
Abstract:
Interpreting deep learning models typically relies on post-hoc saliency map techniques. However, these techniques often fail to serve as actionable feedback to clinicians, and they do not directly explain the decision mechanism. We propose an inherently interpretable model that combines the feature extraction capabilities of deep neural networks with advantages of sparse linear models in interpretability. Our approach relies on straightforward but effective changes to a deep bag-of-local-features model (BagNet). These modifications lead to fine-grained and sparse class evidence maps which, by design, correctly reflect the model's decision mechanism. Our model is particularly suited for tasks which rely on characterising regions of interests that are very small and distributed over the image. In this paper, we focus on the detection of Diabetic Retinopathy, which is characterised by the progressive presence of small retinal lesions on fundus images. We observed good classification accuracy despite our added sparseness constraint. In addition, our model precisely highlighted retinal lesions relevant for the disease grading task and excluded irrelevant regions from the decision mechanism. The results suggest our sparse BagNet model can be a useful tool for clinicians as it allows efficient inspection of the model predictions and facilitates clinicians' and patients' trust.
Abstract:
Visualising high-dimensional datasets is a crucial step in exploratory data analysis. However, there has been some controversy about the relation of the two most prominent non-linear visualisation methods, t-SNE and UMAP. They have seemingly unrelated objective functions and distinct key ingredients, producing qualitatively different embeddings.
We find the exact relation between UMAP and t-SNE and show that UMAP is essentially negative sampling applied to the t-SNE loss function. Our main insights are the derivation of UMAP’s true loss function and its connection to noise-contrastive estimation, which is used by the method NCVis to approximate t-SNE.
Our analysis shows that UMAP applies stronger attraction on neighbouring points than t-SNE and thus explains the more compact embeddings produced by UMAP. Varying the attraction strength further leads to a whole spectrum of neighbour embedding methods on which t-SNE, UMAP and NCVis are only three points.
Gauging attraction and repulsion has important practical consequences. On the one hand, stronger attraction can highlight continuous structures, like trajectories. On the other hand, more repulsion resolves discrete clusters better and thus shows the finer local structure of a dataset. We recommend looking at different points on the attraction-repulsion spectrum to observe how (possibly artifactual) structure arises and decays in order to faithfully explore high-dimensional datasets
Kalala Mutombo Franck | University of Lubumbashi
Abstract:
We review briefly concepts of Networks and then we focus on the diffusion process of the network using the Laplacian of the network. We then look at the heart heat kernel of the diffusion process which is obtained by exponentiating the Laplacian of the network. This concept is extended by considering a novel network-theoretic approach developed in recent years and consisting of the definition of the k-path Laplacian operator for networks. We use the idea of incorporating long-range interactions (LRI) in the transmission of information through the nodes and the edges of the network. This generalized heat kernel is then used to characterized networks and applied to object clustering.
Abstract:
Most cancer patients die due to metastasis, and the early onset of this multistep process is usually missed by current staging tumour modalities [1]. Advanced technics exist to enrich disseminated tumour cells from patient blood and bone mallow as cancer progression marker. However, these cells present high heterogeneity [2], only some of them can exhibit stem cell phenotype and tumour development potential, other can have plasticity potential to reprogram into cancer stem cells [3]. So, detection and characterization is challenging due to lack of clear phenotypic markers [4]. Therefore, there is a critical need to find new ways to anticipate and predict metastasis development at an early stage of patient care. Cancer progression involves many cellular morphological effects, which have been revealed by biophysical studies [5]. The relevance of the characterization of cancer cells by their electromechanical phenotype is attested by reports pointing out their physical alteration as reduced deformability of circulating lymphocytes in the case of chronic lymphocytic leukaemia, and increased size and stiffness of labelled circulating tumor cells (CTC). The physical properties allow identifying different malignant breast epithelial cells lines by their viscoelastic behavior [6] or their electrical impedance [7]. Even though, the analysis capability of cancer cell by their physical characteristics has been demonstrated [5], the current state of the art is far from being relevant for clinical practices. This project aims to develop and push the concept of cell physical phenotyping to categorize disseminating cells, evaluate their metastatic potential towards cancer diagnosis and prognosis. We aim to combine MEMS (Micro-Electro-Mechanical Systems), biological assays, statistical learning ([8]) with following objectives:
• Physical phenotyping: A high-throughput method for physical characterization of cells
• Biological phenotyping: Characterization of cells to analyze their metastatic potential
• Modeling: Statistical functional methods for physical/biological cell classification
• Predicting the metastatic potential of cells
This work offers a complete mapping and correlation of different characteristics (biological, physical and genetic) of selected cells lines to model their metastatic potential for improved diagnostic and prognostic capabilities. It aims providing the required statistical/ML tools to predict metastatic potential of a cell by
means of physical properties; a method that is faster, cheaper and better suited for point-of-care applications. For data processing, the needed statistical and computational classification and prediction methods require training with large-scale, heterogeneous, continuous (functional) and spatial datasets including
high numbers of cells’ physical and biological properties. A very limited number of models are available for description, visualization, classification and prediction of quantitative data involving continuous (functional) heterogeneous data. Challenges are, in one hand, to use statistical (including functional
classification, regression methods) tools able to compute in a non-costly way correlation among huge amounts of data, for training and accuracy prediction. A main issue when analyzing our massive data is to use statistical tools including functional classification, regression methods) able to compute in a non-costly way correlation among huge amounts of data and the ultimate goal of predicting the metastatic potential of cells by physical characterization of cells.
The work is interdisciplinary (biophysicists, biologists with analysts and statisticians) and is in collaboration with University of Tokyo (Bio-MEMS technology in SMMiL-E ; Seeding Microsystems in Medecine in Lille – European-Japanese Technologies against Cancer) project: http://www.ircl.org/programme-smmil-e/).
References:
[1] Fidler IJ. Nat Rev Cancer, 2003, 3, 453–458.
[2] Punnoose EA, Atwal SK, Spoerke JM, Savage H, Pandita A, et al.,. PLoS One. 2010, 5:e12517.
[3] Lagadec C; Vlashi E; Della Donna L; Dekmezian C; Pajonk F. Stem Cell, 2012, 30(5), 833-44.
[4] Millner L, Linder M, Valdes R. Ann Clin Lab Sci. 2013, 43(3), 295–304.
[5] S. E. Cross, Y.-S. Jin, J. Rao, and J. K. Gimzewski, Nature Nanotech, 2007, 2, 12, 780–783.
[6] J. Guck, S. Schinkinger, B. Lincoln, F. Wottawah, S. Ebert, et al.,, Biophys. J, 2005, 88, 5, 3689–3698.
[7] A. Han, L. Yang, and A. B. Frazier, Clinical Cancer Research, 2007, 13, 1, 139–143.
[8] A. Bahram et al. MEMS, 2022, 2022, 317-320
Abstract:
Locust outbreaks are notoriously difficult to deal with and their consequences are felt for years. In 2020 alone, over 23 million people were made severely food insecure and over 12 million were forcibly displaced due to an upsurge of desert locusts. In this presentation, we will give an overview of some of our latest efforts towards building an AI-driven early warning system for detecting desert locusts. This project is a collaboration between InstaDeep and the Google research team based in Accra, Ghana. Such an early warning system would allow control operations to take the proper precautions when facing potential threats from locusts and be able to assist in mitigating risks to food security posed by locust outbreaks.
Abstract:
Progress toward the United Nations Sustainable Development Goals (SDGs) has been hindered by a lack of data on key environmental and socioeconomic indicators, which historically have come from ground surveys with sparse temporal and spatial coverage. Recent advances in machine learning have made it possible to utilize abundant, frequently-updated, and globally available data, such as from satellites or social media, to provide insights into progress toward SDGs. Despite promising early results, approaches to using such data for SDG measurement thus far have largely evaluated on different datasets or used inconsistent evaluation metrics, making it hard to understand whether performance is improving and where additional research would be most fruitful. Our contributions our two-fold. First, we have compiled one of the largest ML-ready datasets for tracking progress towards UN SDGs, which we call SustainBench. This is a collection of 15 benchmark tasks across 7 SDGs, including tasks related to economic development, agriculture, health, education, water and sanitation, climate action, and life on land. It significantly lowers the barriers to entry for the ML community to contribute to measuring and achieving the SDGs and provides standard benchmarks for evaluating ML models on tasks across a variety of SDGs. Second, we have developed novel machine learning methods that have directly facilitated progress towards the SDGs, such as enabling targeted cash-transfer programs at the start of the COVID-19 pandemic. Nonetheless, there remains significant room for developing new AI methods for further accelerating progress towards the SDGs.
Abstract:
This talk will address the data format and archiving/compession issues for the Square Kilometre Array (SKA). One of the major contributors to the large data volumes is the maximum baseline of an array since it sets how well the data needs to be sampled to avoid smearing and decorrelation of the signal. Now, the sampling rate is dependent on the baseline length as shorter baselines can be sampled a lot more coarsely which leads to smaller data volumes. In fact, baseline-dependent averaging (BDA) is an established technique for compressing radio interferometric data, however, this technique results in irregularly sampled data which while supported by the database format requires the data to be restructured in ways that reduce performance when processing. But more importantly, radio-astronomy software tools implicitly assume uniformly sampled datasets. Recent research in M. Atemkeng, in prep. in collaboration with the Wits University and SARAO; shows an approach that uses local low-rank matrix approximation to achieve similar compression rates while significantly minimizing smearing.
Issa Karambal | Facilitator