Statistical and Biological Data Analysis Using Programming Languages
摘要
Rapid progress in high throughput sequencing technologies has generated an unprecedented huge amount of complex and heterogeneous biological data and has opened the door for researchers to carry out different types of analysis in various areas of biological sciences like genomics, transcriptomics, proteomics, metabolomics, etc. Hence, the need for fast and efficient computational tools have been felt for performing various types of in silico analysis of these biological data. Over the years, many computer programming languages like Java, Perl, R, Python, etc. have emerged for software development which helps in the visualization of generated data and subsequent statistical data analysis through which proper inference can be drawn. The main endeavour in bringing out this book is to provide wide-range programming solutions for genomics data handling and analysis to address complex problems associated with crop improvement programs. Special emphasis has been provided on R and Perl programming languages as these are highly popular amongst biologists, especially biometricians and bioinformaticians. Key features of these two programming languages, viz., installation procedures, basic syntax, etc. and their applications in areas of biological data analysis supported by suitable examples are described in successive sections. The targeted audiences of this book are primarily students, data analysts, scientists and researchers from various fields of biological sciences. This book will guide the readers in solving various emerging bioinformatics problems, such as assembly of NGS data, annotation, interpretation, visualization, mapping of QTLs, association mapping, genomic selection, etc.