Showing posts with label Decision Trees. Show all posts
Showing posts with label Decision Trees. Show all posts

Sunday, October 28, 2012

Can we Classify students with Data Mining ???

In web-based educational environments predict student’s performances is very important where the students who are at the risk of failing examinations can be identified at the early stage of the course modules and the educators can take necessary actions to improve their knowledge to a more higher level and to increase their learning capacities as well. 

In data mining context use of classification on the educational data is a upcoming research area where to discover potential student groups with similar characteristics and to identify learners with low motivation and find corrective actions to lower drop-out rates. 

C. Romero, S. Ventura, P. G. Espejo, and C. Hervás [2008] have tried to used different classification approaches on  the student data to compare the applicability on data mining techniques for classifying the students in to groups and to predict the final marks obtained in the course modules. 

In their research they used a framework which is known as KEEL which is an open source framework for building data mining models including classification, regression, clustering, pattern mining and based on this framework they developed an data mining tool which can be integrated in to the moodle environment. 


Friday, October 26, 2012

What is the level of accuracy can be achieve in predicting student performances ???

With the recent developments in the Internet allowed many leading educational institutions to offer online teaching and learning through Learning Management Systems. Systems with different capabilities and approaches have been developed to deliver online education which makes the communication channel between student and teacher into a much more virtual one.
 
The most important consideration would be to identify the student’s performance accurately and provide the necessary support to the students to improve their knowledge levels. But the main problem of measuring the students in accurately and to classify them correctly in to group to predict their performance levels is a huge challenge.
Behrouz Minaei-Bidgoli, Deborah A. Kashy, Gerd Kortemeyer, William F. Punch [2003] have researched on applying data mining methodologies to classify the students in to different groups and try to predict their performance achievement for the future.
For this research they used Quadratic Bayesian classifier, 1-nearest neighbor (1-NN), k-nearest neighbor (k-NN), Parzen-window, multilayer perceptron (MLP), and Decision Tree as the data mining mechanisms and by combining multiple classifiers they hoped to improve classifier performance. For the decision trees they have used C5.0, CART, QUEST, CRUISE algorithms.
In this paper they focused on using a Genetic Algorithm to optimize a combination of classifiers. They used GAToolBox for MATLAB to implement a Genetic Algorithm to optimize classification performance and to find a population of best weights for every feature vector which minimize the classification error rate.
Finally they concluded that using Genetic Algorithm more than a 10% performance improvement can be achieved and having the information generated the instructor would be able to identify students at risk early.
Reference : B. Minaei-Bidgoli, D. A. Kashy, G. Kortmeyer, and W. F. Punch, “Predicting student performance: an application of data mining methods with an educational web-based system,” in Frontiers in Education, 2003. FIE 2003 33rd Annual, 2003, vol. 1, p. T2A–13.

The use of Genetic Algorithms and Decision Trees in Distance Education

In the research of Dimitris Kalles,Christos Pierrakeas [2006] they tried to used genetic algorithm and decision tree based classification on student data to understand the different learning capacities of the students. In their research they based the applicability of these algorithms on different sets of students under different course modules. 

In this research they mainly used the genetic algorithm based decision tree implementation of GATREE which is built using the GALIB library. The genetic operators on the tree representations are relatively straightforward where a mutation may modify the test attribute at a node or the class label at a leaf and a cross-over may substitute whole parts of a decision tree by parts of another decision tree.

For creating the dataset the students’ key demographic characteristics of students such as age, sex, residence and their marks in written assignments and their presence or absence in plenary meetings were considered to create the training dataset. 

In this research they used the GATREE system and experimented with to 150 generations and up to 150 members per generation. To ensure the validity of the experimentation they used the same data sets of the original experimentation which includes demographic data and quantized data.

They observed that GATREE induced trees provide good accuracy estimation, even without the cross-validation testing phase. Their initial findings suggested that when compared to conventional decision-tree classifiers this approach produces significantly more accurate trees.

However it was noted that GATREE has been generating closer estimations even with the quantized formats which gives an indication that GATREE can produce quality results even in the presence of noise.

Reference: D. Kalles and C. Pierrakeas, “Analyzing student performance in distance learning with genetic algorithms and decision trees,” Applied Artificial Intelligence, vol. 20, no. 8, pp. 655–674, 2006.

Thursday, October 25, 2012

Performance Prediction Models on Educational Data

In educational systems like Learning Management Systems Students’ academic performance depends on diverse factors like personal, socio-economic, psychological and other environmental variables. Each of these factors can affect the student overall performance in different weights. Based on the level how each of these factor is appearing in the student education several learning patterns can be identified on each of these students. Based on these learning patterns prediction models can be implemented such that they include all these variables for the effective prediction of the performance of the students. The prediction of student performance with high accuracy is beneficial to identify the students with low academic achievements which enable the educators to assist those students individually.

In M. Ramaswami and R. Bhaskaran [2010] research they argued that the student performance could depend on diversified factors such as demographic, academic, psychological, socio-economic and other environmental factors. Based on these factors they constructed a CHAID prediction model with highly influencing predictive variables obtained through feature selection technique to evaluate the academic achievement of students.

Can we use Decision Tree for the predication models ??? Better or Worse ??

In most of my previous posts were focused on use of various data mining or machine learning approaches on educational data to understand the learning patterns of different students. Educational data mining is a new emerging practice of data mining that can be applied on the data related to the field of education. Process of transforming raw educational data which are collected by education learning systems could be used to take informed decisions on students learning problems. The various techniques of data mining like classification, clustering and rule mining can be applied to bring out various hidden patterns from the educational data.
 
In overall student learning environments upstanding the overall student performance will be beneficial in examinations which is playing a vital role in any student’s life. The marks obtained by the student in the examinations for different course modules will decide the overall grade he obtain in finally. Therefore it is becoming essential for any tutor to understand whether different students will pass or fail in the examinations in beforehand they face them. Based on these predictions tutors can help the students in prior to the examinations and extra efforts can be taken to improve their studies and help them to pass the examinations.
 
S. Anupama Kumar and  M Vijayalakshmi [2011] have tried  to predict the student overall performance based on their internal assessments in the learning environment. In their research they considered five course modules which were offered in a semester and overall student count for the selected analysis were about 117. The algorithms they used for this research is J48 and ID3 decision tree algorithms. According to their discussion the accuracy level of the J48 was higher than the ID3 algorithm since it has predict more correct prediction results than the ID3 algorithm. Based on the different accuracy levels on each of the decision tree they created they concluded that classification techniques can be applied on educational data for predicting the student’s outcome and improve their results and the efficiency of various decision tree algorithms can be analyzed based on their accuracy and time taken to derive the tree. Finally they argue that the application of data mining brings a lot of advantages in higher learning institutions so that these techniques can be applied in the areas of education to optimize the resources allocations as needed with the student learning capacity .
Reference: S. A. Kumar and M. N. Vijayalakshmi, “Efficiency of decision trees in predicting student’s academic performance,” in First International Conference on Computer Science, Engineering and Applications, CS and IT, 2011, vol. 2, pp. 335–343.

Can we prevent school dropout in Distance Learning???

Distance and open education is a emerging educational principle where many universities and institutions are using e-learning in distance education for their study programs. Due to the nature of virtual communication the teachers and students are not meeting each other face to face. Because of this learning nature many students are dropping out from the educational programs since they cannot cope with the requirement of the study programs. Therefore understanding the performance level of each student will help the teachers to identify different capability levels of the student which will help them to climb up the ladder with their peers. 

S. B. Kotsiantis, C. J. Pierrakeas, and P. E. Pintelas [2003] have tried to apply data mining methodologies on educational data to limit student dropout in university-level distance learning. According to them the dropout can be caused by professional, academic, health, family and personal reasons and varies depending on the education system adopted by the institution providing distance learning, as well as the selected subject of studies.

They based their research on a course module which was offered in Hellenic Open University which is based their educational programs mainly on distance mode. The built a data set of 365 student instances and based on the data the attributes were divided in to two groups which were the ‘Curriculum-based' group and the ‘Students' performance' group. The ‘Curriculum-based' group represented attributes of students' sex, age, marital status, number of children and occupation and the group represented attributes concerning students' marks on the first two written assignments and their presence or absence in the first two face-to-face meetings.  

In this research they used six machine learning techniques which are Decision Trees, Neural Networks, Naive Bayes algorithm, Instance-Based Learning Algorithms, Logistic Regression and Support Vector Machines. For each of these algorithms they used a representative algorithm as C4.5 algorithm for the decision trees algorithm and to estimate the values of the weights of a neural network the Back Propagation (BP) algorithm was used. The Naive Bayes (NB) algorithm was used for the Bayers algorithm and 3-Nearest Neighbour algorithm was also used. Maximum Likelihood Estimation (MLE) was the used statistical method for estimating the coefficients of the logistic model and finally, the Sequential Minimal Optimization (or SMO) algorithm was the representative of the Support Vector Machine.

Based on these six algorithms they found that Naive Bayes algorithm and Back Propagation (BP) algorithm had the best accuracy with the data sets. However they mentioned that the differences were generally small and because they were only based on one course module and it may possible that the ranking in another data set of the same domain is different. Also they concluded that Naive Bayes has the  short training time and effective communicated way of predicting and the small programming cost than the other algorithms.

Reference: S. Kotsiantis, C. Pierrakeas, and P. Pintelas, “Preventing student dropout in distance learning using machine learning techniques,” in Knowledge-Based Intelligent Information and Engineering Systems, 2003, pp. 267–274.

Saturday, October 20, 2012

Predicting Student Performance in Distance Learning Systems

In any distance learning environment ability of predicting a student’s performance is very important which is advantageous for the teachers and tutors to identify the students with different capabilities and their capacities. When it comes to University education where many students are accessing or following their studies through open and distance environment it requires a identification process upon the students to measure whether they achieve the required level of performance. Otherwise due to the nature of the distance education some students can be lagging behind while peer students have passed them by miles. If teachers and tutors can recognize them at the early stage of the course module necessary steps or decisions can be made in order to prevent them from dropping out from the course modules.

S. Kotsiantis, C. Pierrakeas, and P. Pintelas [2003] have suggested an approach which has used machine learning algorithms with the LMS data to prevent, student dropouts in university distance education. They tried to investigate the efficiency of machine learning techniques in such an environment with trained data sets provided by the “informatics” course of the Hellenic Open University.

In their research they used five different algorithms to study student data and they found that these algorithms can be used more appropriately to predict the student dropouts in study programs. In this research they used most common machine learning techniques which are Decision Trees, Bayesian Nets, Perceptron-based Learning, Instance-Based Learning and Rule-learning.

In their data collection process they collected student data under two categories of attributes which are Demographic attributes and Performance attributes. The Demographic attributes were collected by concerning students’ sex, age, marital status, number of children and occupation and Performance attributes represents attributes which were collected from tutors’ records concerning students’ marks on the written assignments and their presence or absence in face-to-face meetings.

In the above mentioned algorithms categories they used C4.5 algorithm for representing the decision tree, Naive Bayes algorithm was the representative of the Bayesian networks, the RIPPER algorithm was the representative of the rule-learning techniques, WINNOW as the representative of perceptron-based algorithms and finally 3-NN or 3- Nearest Neighbor as the Instance-Based Learning algorithm.

In order to rank the representative algorithms they used the prediction accuracy criterion was used. In the evaluation of the algorithms they found that there was no statistically significant difference between algorithms, but it showed that the Naive Bayes algorithm and the RIPPER had the best accuracy than the others. Among the Naive Bayes algorithm and the RIPPER, Naive Bayes has the advantage short computational time requirement and importantly Naive Bayes classifier can use data with missing values as inputs, whereas RIPPER cannot work with which gives a indication that the Naive Bayes is the most appropriate learning algorithm to be used for the construction of a software support tool in Learning Management Systems.

Other than the above it was found that there exist some obvious and some less obvious attributes that demonstrate a strong correlation with student performance where some gives the higher importance in consideration. Also it can be argued that the learning algorithms could enable tutors to predict student performance with satisfying accuracy long before final examination. 

Reference: S. Kotsiantis, C. Pierrakeas, and P. Pintelas, “Efficiency of Machine Learning Techniques in Predicting Students’ Performance in Distance Learning Systems,” Citeseer, 2002.

Thursday, October 18, 2012

Classification Approaches in Learning Analytics, Does it always give the better results ????

In educational data mining many researchers have tried many different approaches available in data mining context to predict the learning patterns of the students to achieve better and quality results. All these approaches are mainly focusing on getting the results according to a particular student domain which highlights various specific features indicated through the Learning Management Systems.

Virtual learning is growing enormously and the student population those connect with these Learning management systems are increasing by numbers every day. With the necessity of understanding of each student learning pattern teachers should have a better way of predicting the performance of their students. In response to this necessity different classification techniques can be used to compare and interpret the educational data and improve the modeling of students in to different categories.

In order to make the student modeling process much easier Diego Garcia Saiz and Marta Zorrilla [2011] have researched on applying different classification techniques on the student data to predict their performances. In their research they tried to implement a tool known as Elearning Web Miner (EIWM) to discovering how the students are behaving and progress in the courses which is very helpful for the tutors to identify the students who need more attention among from a larger set of students.

One of the main reason that applying learning analytics in educational data sources is challengeable because of the dataset becomes very small comparing to the other application we see around us. Even though the number of student information which contains in a database is huge, most of these are dynamic and contain many variations among them. Since for this research they found it difficult to collect required data which made them to use the data for past three academic years for average student enrollment of 70  per year for a specific course module. For all of these student instances they considered attributes with mean values such as total time spent, number of sessions carried out, number of sessions per week, average time spent per week and average time per session.   

With the intention of analyzing and choosing  best classification algorithms for educational datasets they  analyzed four of the most common machine learning techniques, namely Rule-based algorithms, Decision Trees, Bayesian classifiers and Instance-based learner classifiers which are mainly were OneR, J48, Naive Bayes, BayesNet TAN and NNge

They tested these five algorithms using different parameter settings and different numbers of folds for cross validation, in order to discover whether they have a great effect on the result. In the evaluation process they found that Bayes algorithms perform better in accuracy and is comparable to J48 algorithm although it is worse at predicting than Naive Bayes which is the best in this aspect. Due to the results they achieved they highlighted that OneR suffered from over-fitting in this dataset, so that it should be discarded as a suitable classifier for very small datasets.

They also observed that NNge improves its performance in this dataset although the great number of rules which it offers as output makes it less interpretative for instructors than the rest of the models. Finally they conclude that Bayes Networks are suitable for small datasets in performing better than the Naive Bayes when the sample is smaller. As consequence of the fact that BayesNet TAN model is more difficult to interpret for a non-expert users and J48 is similar in accuracy to it.

One significant result which I see in their research is that the pre-processing step which they followed. In the dataset they found that there are instances which can be considered as outliers in the statistical sense and they suggested a mechanism to remove or eliminate those outliers in the data set which can improve the results by 20%. This makes a huge advantage when the data set is larger in size and provide with better quality results for the users.

What I believe about this research is that even though they suggested these approaches in classify the students, it cannot be proved that the same algorithm is suit for every situation we have in the educational domains. Some algorithms can perform well with small datasets and some can perform well with larger data samples and some are providing more interpretable results and some are not. Therefore depending on the problem situation and the context we have to choose the best algorithm that can be used for the specific process so that we get more acceptable quality output as final results.

Reference: D. García-Saiz and M. Zorrilla, “Comparing classication methods for predicting distance students’ performance,” 2011.
 

Learning Analytics Approach with Sakai

As I have discussed in my previous posts the Sakai is an Open Source Learning Management System which is been widely used in academic context for their study programs. Lauria and Joshua [2011] have tried to implement a predictive model within the Sakai for predicting the performance of the students and to take the decisions for making corrective actions. In their research they came up with a methodology which contains six phase on the knowledge discovery process.
 
When collecting the required data they extracted information from diverse sources and followed several pre-processing steps to handle the missing value, outliers and incomplete records. All the data which were logged through Sakai were aggregated to produce consolidated records per course and student. In order to remove the variations in different course contents all the data were collected as ratio values rather than an absolute value.
 
After the data collection process they followed some steps to reduce the dimensionalities on the available data. In order to maintain a proper level of query accuracy and efficiency the number of variables and parameters requiring for the estimation were selected properly and unnecessary features were removed.
 
After the necessary data was selected the transformation and rescaling phase was carried out to make sure that all the attribute data were formatted according to the requirement of the data mining algorithms they used. After the data was converted or transformed the partitioning step was used to divide the data in to several groups. They carried out this partitioning process on the data set to make sure that required amount of data is available for the training of the data model and for the validation with testing step. 
 
For building the data models four different types of data mining approaches were selected. Logistic regression, C5.0 decision tree, support vector machine and Bayesian networks were used for creating the train models with the data set. After the models were created they were validated using the validation data set. For validating the data models they measure the prediction accuracy on the data to verify that the required level of accuracy or the quality can be achieved by the models.
 
Reference :
Eitel J.M. Lauría, Joshua Baron, Mining Sakai to Measure Student Performance: Opportunities and Challenges in Academic Analytics, 2011