Skip to Content
Bioinformatics with Python Cookbook - Second Edition
book

Bioinformatics with Python Cookbook - Second Edition

by Tiago Antao
November 2018
Intermediate to advanced
360 pages
9h 36m
English
Packt Publishing
Content preview from Bioinformatics with Python Cookbook - Second Edition

Preparing the dataset for analysis

Our starting point will be a VCF file (or equivalent), with calls made by a genotyper (Genome Analysis Toolkit (GATK) in our case), including the annotations. As we will be filtering NGS data, we need reliable decision criteria to call a site. So, how do we get that information? Generally, we can't, but if we need to do it, there are three basic approaches:

  • Using a more robust sequencing technology for comparison; for example, using Sanger sequencing to verify NGS datasets. This is cost-prohibitive and can only be done for a few loci.
  • Sequencing closely related individuals, for example, two parents and their offspring. In this case, we use Mendelian inheritance rules to devise if a certain call is acceptable ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Start your free trial

You might also like

Bioinformatics with Python Cookbook

Bioinformatics with Python Cookbook

Tiago Antao

Publisher Resources

ISBN: 9781789344691Supplemental Content