Showing posts with label Prediction Model. Show all posts
Showing posts with label Prediction Model. Show all posts

Saturday, October 27, 2012

Student Individualized Growth Model and Assessment (SIGMA)

As I mentioned in my previous posts student dropout from educational programs is becoming the most pressing issues in current time. According to the Carnegie Corporation of New York they said that 

“Today, young people who leave high school without excellent and flexible reading and writing skills stand at a great disadvantage. In the past, those students who dropped out of high school could count on an array of options for establishing a productive and successful life. But in a society driven by knowledge and ever-accelerating demands for reading and writing skills, very few options exist for young people lacking a high school diploma.”

Same like the Sakai Learning Management System Microsoft have introduced a new learning analytic platform which can be used identify different students according to their performance activities which is basically using the predictive analysis approach. 

According to the study conducted by the U.S. Department of Education the most common reasons for student to be dropping out of school are 

  • Lack of educational support. Many students decided to drop out of high school due to lack of sufficient parental support and educational encouragement.
  • Outside influences.  
  • Special needs. Students often drop out of high school because they require specific attention to a certain need, such as dyslexia or other learning disabilities
  • Financial problems.  
Out of those four mentioned above the Lack of educational support and the Special needs reasons can be easily managed using predictive analysis approach since student who are at risk of failing can be identified at early stage by analyzing their historical data of learning behavior. 

According to the Microsoft the Student Information Systems and Learning Management Systems (LMS) have a  shortcoming of inability to perform integrated analysis of large amounts of data. As they mentioned in their report to manage this type of reporting infrastructure requires a different type of data analysis system that is highly optimized for rapid and comprehensive analysis of large amounts of data. Online Analytical Processing can be used as an approach for fast analysis of large amounts of data offering greater insight into student performances.

Microsoft Education Analytics Platform (EAP) or SIGMA offers both business intelligence and predictive analytics data management services which consider about different aspect of the students where it categorize the influence factors in to individual and family. Based on these factors they use predictive analysis approaches to identify the students who are at risk of dropping out from the educational programs.

Reference : Student Individualized Growth Model and Assessment (SIGMA), A Microsoft Education Analytics Platform Approach to Students at Risk, May 2010

Can we use Learning Analytics Approaches to Prevent Dropping out of Students ???

In any educational institution the learning capacities and the learning capabilities of the students are varying in different levels due to the way their arranging their learning behaviors. Due to this learning nature applying the same teaching principle on each and every student in the same manner does not provide the required level of knowledge in students. This brings a indication to the teachers why many students are dropping out from schools or educational institutions. 

In year 2009, Massachusetts Department of Elementary and Secondary Education started a project known as Dropout Prevention Planning Project to implement an approach to understand the different factors on student's dropout incidents. They implemented a Student Information Management System which contains information about the student within the State and they initiate an index known as Early Warning Indicator Index for measure the risk of dropout in each student. 

According to the Cash, Dawicki, Sevick [2011] the early warning system to be useful and effective and it should allow districts to achieve the following goals and objectives in their dropout prevention efforts:

  • Goal 1: Accurately define and uncover students’ problems and needs
  • Goal 2: Successfully identify interventions and improvement strategies
  • Goal 3: Effectively target and initiate programs and reforms
  • Goal 4: Truthfully monitor ongoing efforts and progress with at-risk students
In order to achieve these goals either Business Intelligence or Predictive Analytics can be used. Business intelligence model focus on analyzing historical and current data in order to provide a look at operations or conditions at a given period of time. Predictive analytics approach attempts to incorporate historical data into statistical models in order to make predictions about future events or outcomes.

According to them currently there are several early warning systems available which are using predictive analysis approaches to analyze the students. Microsoft SIGMA is a early warning system  which capitalizes on its Education Analytics Platform (EAP) to provide a new data-based approach to managing students who are at-risk and it is known as the Student Individualized Growth Model and Assessment. 

Mizuni’s Data Warehouse and Dashboard Suite is a transactional and aggregation data store for managing and analyzing data and offers education stakeholders insight into student performance by monitoring key indicators to increase student achievement.

VERSI-FIT is also based on Microsoft has also developed its own early warning system based upon the Education Analytics Platform which is known as the Edvantage At-Risk Early Warning System and Credit Recovery System.


Reference: T. Cash, C. Dawicki, and B. Sevick, “Springfield Public Schools Dropout Prevention Program Assessment & Review (PAR),” 2011.

Wednesday, October 17, 2012

Using K-Nearest Neighbor Algorithm on Student Data

The k-nearest neighbor data mining method is a classical prediction method among the machine learning techniques available in data mining. It has been widely used due to its simplicity and adaptability in predicting many different types of data. The main advantage of using KNN in prediction processes is that the KNN is a lazy method which does not require a model to represent the statistics and distribution of the original training data rather it can be apply on the actual instances of the training data. Even though the KNN is a simple predictive algorithm which we can rely on and it does not make any assumption about the prior probabilities of the training dat. Also the KNN is satisfactorily used on the situations when the data set is included with noisy and incomplete data.

Due to the advantages and the simplicity of the algorithm T. Tanner and H. Toivonen [2010] has tried to implement a model using K nearest neighbor data mining algorithm to identify the student who are at high risk of failing in a specific course. By this research they suggested that good results in predicting final scores indicate that students with learning problems can be found reliably. What they have been using on the student data is to make prediction on the performance of a given student based on the similarity to all instances in the training set and find the k most similar objects in the data set. This similarity is calculated by using a Euclidean distance between the features of the test subject and the corresponding features of each instant in the training set.

In their research they showed that KNN can produce predictions accurately for the final scores even after the first lesson. Another interesting result they found is that in any skill based courses early tests on skills can be used as the predictors for the final scores and they suggested that predicting final scores for the courses can be used to identify the students with learning problems and can be used directly to implement as a early warning features for the teachers so that the students can be altered if they are likely to fail the final tests. Based on the information or features they used for the experiment with the KNN algorithm they suggest that KNN method could be just as effective in other learning management systems (LMSs) such as Moodle where only a single lesson score is available for student assessment and especially other skill-based courses could be a good fit for the KNN method.

Reference : T. Tanner and H. Toivonen, “Predicting and preventing student failure–using the k-nearest neighbour method to predict student performance in an online course environment,” International Journal of Learning Technology, vol. 5, no. 4, pp. 356–377, 2010.

How we can implement Learning Analytics???

Academic analytics or Learning analytics is a wide term used recently to describe the use of data mining in educational data sources. Researches have used various data mining methodologies in different ways to understand or identify the learning models and learning patterns of the students. Various learning management systems can be used for providing the study programs for the students those who can be connected in open distance mode or either in a much more hybrid manner. But applying the same teaching principle on all the students in the same way will not cope with the ultimate goal of any learning management systems available which is to help the students to learn rather helping them to pass.

By understanding the way these students are learning it will help the teachers and the academics to change the teaching methodologies they used to cope with the requirements of the students. In order to identify how these students are behaving performance models can be used. It can be used as a monitoring tool to take necessary actions for the issues related to student learning. As the academic analytics suggest the knowledge can be created about the students by applying the statistical analysis and predictive modeling on educational data sources.

For building the predictive model we can use the continue streams of data which is created within the learning management systems with the data mining techniques and it can be used as a decision making tool for teachers. Before building the predictive model it is a must to acquire proper information out from the student data since not all information available in the data sources is important for the analyzing. The outcomes of these prediction models will be to provide guidelines or forecasting about the event that can be occurring based on the observations or the scenarios.

In any data mining research some steps can be given as the basic steps that we need to follow in the knowledge discovery process. The same principles can also be applied in the learning analytics processes where the data collecting and data pre-processing steps will be continued in the same manner. According to the requirements in the data mining methodology data reductions or data cleaning steps can be done. After the prediction model is been implemented using data mining approaches it needs to verify that the predictive accuracy is acceptable on the given student data. The main advantage which I see in the given context is that most of the data which are available in the data sources are labeled and they contains information about the student characteristics as well as the course management events where the modeling process can be applied in many ways.