Skip to main navigation Skip to search Skip to main content

SNPQC: an R pipeline for quality control of Illumina SNP genotyping array data

Cedric Gondro, Laércio R Porto-Neto, Seung Hwan Lee

Research output: Contribution to journalArticlepeer-review

26 Citations (Scopus)

Abstract

In genome-wide association studies, quality control (QC) of genotypes is important to avoid spurious results. It is also important to maintain long-term data integrity, particularly in settings with ongoing genotyping (e.g. estimation of genomic breeding values). Here we discuss SNPQC, a fully automated pipeline to perform QC analyses of Illumina SNP array data. It applies a wide range of common quality metrics with user-defined filtering thresholds to generate a comprehensive QC report and a filtered dataset, including a genomic relationship matrix, ready for further downstream analyses which make it amenable for integration in high-throughput environments. SNPQC also builds a database to store genotypic, phenotypic and quality metrics to ensure data integrity and the option of integrating more samples from subsequent runs. The program is generic across species and array designs, providing a convenient interface between the genotyping laboratory and downstream genome-wide association study or genomic prediction.
Original languageEnglish
Pages (from-to)758-761
JournalAnimal Genetics
Volume45
Issue number5
DOIs
Publication statusPublished - 31 Dec 2014

Keywords

  • Genomics
  • Animal Breeding

Fingerprint

Dive into the research topics of 'SNPQC: an R pipeline for quality control of Illumina SNP genotyping array data'. Together they form a unique fingerprint.

Cite this