Please click the link below to join the session
Please click the link below to watch the recording of the opening session
11:00 - 11:05
11:05 - 11:15
11:15 - 11:20
11:35- 12:30
Please click the link below to watch the recording of the keynote session
Abstract: Q-learning has been one of the most commonly used methods for optimizing dynamic treatment regimes (DTRs) in multi-stage decision-making. Right-censored survival outcome poses a significant challenge to Q-Learning due to its reliance on parametric models that are subject to misspecification and sensitive to missing covariates. In this paper, we propose an imputation-based Q-learning (IQ-learning) where the semiparametric Cox proportional hazard (CPH) model is employed to estimate optimal treatment rules for each stage, and then weighted hot-deck multiple imputation (MI) and direct-draw MI are used to predict optimal potential survival times. Missing data are handled using inverse probability weighting and MI, and the non-random treatment assignment among the observed is accounted for using a propensity-score approach. We investigate the performance of IQ-learning via extensive simulations and show that it is robust to model misspecification, imputes only plausible potential survival times contrary to parametric models, and provides more flexibility in terms of baseline hazard shape. We demonstrate IQ-Learning by developing an optimal DTR for leukemia treatment based on a randomized trial with observational follow-up.
12:30 - 2:00
2:00 - 2:25
Please click the link below to watch the recording of this talk
Abstract: Data plots have been being used widely for exploratory data analysis, model checking and diagnosis. The question is, if it also helps in the statistical decision making process. In this talk we will present how statistical graphics or data plots can be used in inferential procedures. We call this method visual statistical inference. Surprisingly, visual tests have higher power than the conventional UMP tests when the effect size is large. We will also present some scenarios to demonstrate how visual inference can be applied to make decisions with practical significance instead of just statistical significance.
2:30 - 2:55
Please click the link below to watch the recording of this talk
Abstract: Multiple testing has been an area of active statistical research in the past decade mainly because of its wide scope of applicability in modern scientific investigations. Currently research in multiple testing is mainly focused on developing powerful methods even when the number of tests is very large. This talk briefly reviews modern multiple testing methodologies before focusing on its primary goal of making further contributions to the field of controlling false discovery proportion (FDP). More specifically, we propose four newer step-up procedures controlling the γ-FDP, the probability of FDP exceeding γ, given some γ∈ [0,1). The first of these procedures is developed by modifying the Benjamini and Hochberg (1995, J. Roy. Statist. Soc., Ser. B) critical constants, which controls the γ-FDP under both independent and positively dependent test statistics. The second one is a two-stage adaptive procedure developed from these modified Benjamini and Hochberg critical constants and controls the γ-FDP under independence. The third and fourth procedures are also two-stage adaptive procedures controlling the γ-FDP under independence but developed using critical constants in Lehmann and Romano (2005, Ann. of Statist.) and Delattre and Roquain (2015, Ann. of Statist.), respectively. Results of simulation studies examining performances of the proposed procedures relative to their relevant competitors will be presented. We also show the performance of our proposed procedures on high throughput genomic data.
3:00 - 3:25
Please click the link below to watch the recording of this talk
Abstract: This talk is based on joint research I am currently conducting with a member of the BSU statistics faculty, Professor Rebecca Pierce. The motivation for our work is the need to prepare the modern student in our increasingly data driven world with the ability to understand how data is used to make decisions. While our students now have access to amazing technology, the question is, “will they know how to use it properly”? Studies indicate they don’t. Although the ASA guidelines (GAISE) for teaching introductory statistics were clearly articulated over a decade ago, our study indicates there is an issue in achieving the first guideline, “to teach statistics as an investigative process of problem-solving and decision-making”. Our survey results indicates students seem to be learning the steps, but not the logic of the inferential process. We propose a solution to this problem based on George Cobb’s advice to “put the logic of inference at the center of our curriculum”. The goal of my talk is to introduce a new approach, called QED (Question, Explain, Do) to teaching inference in an introductory statistics course, which provides a solid foundational understanding of how data science works, thereby helping our students see “The Big Picture of Statistics”. I will explain the basic theory of this approach and a demonstration of its use.
3:30 - 3:55
Please click the link below to watch the recording of this talk
Abstract:
Aims: The aim of this study is to provide statistical inference of the sojourn time and the transition probability from the disease free to the preclinical state of lung cancer for male and female smokers using lung cancer data from the National Lung Screening Trial (NLST).
Materials and Methods: We applied a likelihood function to the lung cancer data, to obtain Bayesian inference of the transition probability and the sojourn time distribution. A log-normal distribution was used for the transition probability density function multiplied by 30%, and a Weibull distribution was used to model the sojourn time in the preclinical state.
Results: The estimate of screening sensitivity is 0.61 for males and 0.62 for females. Early transition happened before age 50, and lasted until after age 90. The transition probability from the disease free to the preclinical state has a single maximum at around age 73 for males and 72 for females. For male, the Bayesian posterior mean and median sojourn time are 1.28 and 1.23 years, respectively. For female, the corresponding posterior mean and median sojourn time are 1.23 and 1.21 years respectively.
Conclusion: Our estimation showed that male smokers are more vulnerable to lung cancer, because they have a higher transition probability density than the same aged female smokers. The female smokers have a slightly shorter mean sojourn time than the male, meaning that they are quicker to develop clinical symptom of lung cancer.
4:00 - 5:10
5:15 - 6:15
Please click the link below to watch the recording of the panel discussion session
Abstract: In accordance with the conference theme: statistics and data science in today’s driven world, the panel discussion session will feature our notable alumni working at pharmaceutical industries, government, and other institutions. In this session panelists will discuss strategies to be successful as a statistician/data scientist at these or similar organizations.