Popülasyon Genomiği: Temel Kavramlar ve Yöntemler
Özet
Bu bölüm, popülasyon genomisinin temel kavramlarını, veri türlerini ve analiz yöntemlerini kapsamlı ancak uygulamaya dönük bir yaklaşımla tanıtmaktadır. Öncelikle SNP, STR, CNV, indel, yapısal varyant ve haplotip gibi genomik veri türleri ile PLINK ve VCF başta olmak üzere yaygın veri formatları açıklanmaktadır. Güvenilir analiz sonuçları elde etmek amacıyla kalite kontrolü, filtreleme, eksik genotiplerin tamamlanması ve format dönüşümü süreçleri ele alınmaktadır. Bölümde ayrıca Temel Bileşenler Analizi, ortak soy ve karışma analizi, genetik çeşitlilik ölçütleri, FST, AMOVA, genetik uzaklık, akrabalık ve genomik ilişki matrisleri açıklanmaktadır. Doğal ve yapay seçilimin genomda bıraktığı izlerin belirlenmesine yönelik Tajima’s D, Fay-Wu’s H, EHH, iHS ve XP-EHH gibi yöntemlere de yer verilmektedir. Son olarak PLINK, GATK, bcftools, ADMIXTURE, STRUCTURE ve çeşitli R paketleri kullanılarak genomik analizlerin biyoenformatik iş akışı sunulmaktadır. Bölüm; tarım, hayvan ve bitki ıslahı, tıp, biyoçeşitliliğin korunması ve adaptasyon araştırmaları için temel bir başvuru niteliğindedir.
This chapter introduces the fundamental concepts, data types, and analytical methods of population genomics through a comprehensive yet application-oriented framework. It first describes genomic data types—including SNPs, STRs, CNVs, indels, structural variants, and haplotypes—and commonly used formats such as PLINK and VCF. To ensure reliable results, the chapter explains quality control, filtering, genotype imputation, and data-format conversion procedures. It then presents major approaches for investigating population structure and evolutionary processes, including Principal Component Analysis, ancestry and admixture analysis, genetic diversity metrics, FST, AMOVA, genetic distance, kinship, and genomic relationship matrices. Methods for detecting genomic signatures of natural and artificial selection, such as Tajima’s D, Fay and Wu’s H, EHH, iHS, and XP-EHH, are also discussed. Finally, the chapter outlines a practical bioinformatics workflow using PLINK, GATK, bcftools, ADMIXTURE, STRUCTURE, and R packages. It provides a foundational resource for studies in agriculture, animal and plant breeding, medicine, biodiversity conservation, and adaptation.
Referanslar
Luikart G, England PR, Tallmon D, et al. The power and promise of population genomics: From genotyping to genome typing. Nature Reviews. 2003; (4): 981-994. doi: 10.1038/nrg1226
Begun, DJ, Holloway, AK, Stevens K, et al. Population genomics: Whole-genome analysis of polymorphism and divergence in Drosophila simulans. PLOS Biology. 2007; 5 (11), e310. doi:10.1371/journal.pbio.0050310
Pecoraro C, Babbucci M, Franch R, et al. The population genomics of yellowfin tuna (Thunnus albacares) at global geographic scale challenges current stock delineation. Scientific Reports. 2018; 8 (1): 13890. doi:10.1038/s41598-018-32331-3
Purcell S, Neale B, Todd-Brown K, et al. PLINK: A tool set for whole-genome association and population-based linkage analyses. The American Journal of Human Genetics. 2007; 81(3): 559–575. doi: 10.1086/519795
Chang CC, Chow CC, Tellier LC, et al. Second-generation PLINK: Rising to the challenge of larger and richer datasets. GigaScience, 2015; 4(1). doi: 10.1186/s13742-015-0047-8
Danecek P, Auton A, Abecasis G, et al. The variant call format and VCFtools. Bioinformatics. 2011; 27(15): 2156–2158. doi: 10.1093/bioinformatics/btr330
DePristo MA, Banks E, Poplin R, et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nature Genetics. 2011; 43(5): 491–498. doi: 10.1038/ng.806
McKenna A, Hanna M, Banks E, et al. The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Research. 2010; 20(9): 1297–1303. doi: 10.1101/gr.107524.110
Cebeci Z, Şakiroğlu M, Bayraktar M. GWAS: Genom boyu ilişkilendirme çalışmaları. Ankara: Nobel Akademik Yayıncılık; 2024.
Meyer H. plinkQC: Genotype quality control with PLINK. R package version 1.1.0. 2026. (04/04/2026 tarihinde https://CRAN.R-project.org/package=plinkQC adresinden erişilmiştir).
Zheng X, Levine D, Shen J, et al. A high-performance computing toolset for relatedness and principal component analysis of SNP data. Bioinformatics; 2014, 28(24): 3326-3328. doi: 10.1093/bioinformatics/bts606
Jombart T. adegenet: a R package for the multivariate analysis of genetic markers. Bioinformatics. 2008; 24: 1403-1405. doi: 10.1093/bioinformatics/btn129
Gogarten SM, Bhangale T, Conomos MP, et al. GWASTools: an R/Bioconductor package for quality control and analysis of genome-wide association studies. Bioinformatics. 2012; 28(24): 3329-3331. doi: 10.1093/bioinformatics/bts610
Wickham H. ggplot2: Elegant graphics for data analysis. New York: Springer-Verlag; 2016.
Xie Y (2026). knitr: A general-purpose package for dynamic report generation in R. R package ver 1.51 (06/04/2026 tarihinde https://yihui.org/knitr adresinden erişilmiştir).
Pritchard JK, Stephens M, Donnelly P. Inference of population structure using multilocus genotype data. Genetics, 2000; 155(2): 945–959. doi: 10.1093/genetics/155.2.945
Alexander DH, Novembre J, Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Research, 2009; 19(9): 1655–1664. doi: 10.1101/gr.094052.109
Frichot E, François O. LEA: An R package for landscape and ecological association studies. Methods in Ecology and Evolution. 2015; 6(8): 925–929. doi: 10.1111/2041-210X.12382
Evanno G, Regnaut S, Goudet J. Detecting the number of clusters of individuals using the software STRUCTURE: a simulation study. Molecular Ecology, 2005; 14(8): 2611-2620.
Holsinger KE, Weir BS. Genetics in geographically structured populations: defining, estimating and interpreting Fst. Nature Reviews Genetics. 2009; 10(9): 639-650.
Nei M. Genetic distance between populations. The American Naturalist. 1972; 106(949), 283-292. doi: 10.1086/282771
Excoffier L, Smouse PE, Quattro JM. Analysis of molecular variance inferred from metric distances among DNA haplotypes: Application to human mitochondrial DNA restriction data. Genetics. 1992; 131(2): 479–491. doi:10.1093/genetics/131.2.479
Cebeci Z, Gökçe G. Genetik parametreler ve genomik tahmin. 1. Basım. Ankara: Nobel Akademik Yayıncılık; 2023.
Van Raden PM. Efficient methods to compute genomic predictions. Journal of Dairy Science. 2008, 91(11): 4414–4423. doi:10.3168/jds.2007-0980
Astle W, Balding DJ. Population structure and cryptic relatedness in genetic association studies. Statist. Sci. 2009; 24(4): 451-471. doi: 10.1214/09-STS307
Yang J, Lee SH, Goddard ME, et al. GCTA: a tool for genome-wide complex trait analysis. The American Journal of Human Genetics. 2011; 88(1): 76-82.
Endelman JB, Jannink JL. Shrinkage estimation of the genomic relationship matrix. G3: Genes, Genomes. Genetics. 2012; 2(10): 1269–1313. doi: 10.1534/g3.112.004259
Vitti JJ, Grossman SR, Sabeti PC. Detecting natural selection in genomic data. Annual Review of Genetics. 2013, 47: 97-120. doi: 10.1146/annurev-genet-111212-133526
Voight BF, Kudaravalli S, Wen X, et al. A map of recent positive selection in the human genome. PLoS Biology. 2006, 4(3): e72. doi: 10.1371/journal.pbio.0040072
Sabeti PC, Varilly BF, Lohmueller J, et al. Genome-wide detection and characterization of positive selection in human populations. Nature. 2007; 449(7164): 913-918. doi: 10.1038/nature06250
Kim Y, Stephan W. Detecting a local selective sweep in continuous efficiency landscapes. Genetics 2002; 160(2): 765-777. doi: 10.1093/genetics/160.2.765
Tajima F (1989). Statistical method for testing the neutral mutation hypothesis by DNA polymorphism. Genetics. 1989; 123(3): 585-595. doi: 10.1093/genetics/123.3.585
Watterson GA. On the number of segregating sites in genetical models without recombination. Theoretical Pop. Biol.. 1975: 7(2): 256-276.
Fay JC, Wu CI. Hitchhiking under positive Darwinian selection. Genetics. 2000; 155(3): 1405-1413. doi: 10.1093/genetics/155.3.1405
Nielsen R, Slatkin M. An introduction to population genetics: theory and applications. Sunderland, MA: Sinauer Associates; 2013.
Durinck S, Moreau Y, Kasprzyk A, et al. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis. Bioinformatics. 2005; 21 (16): 3439-3440. doi: 10.1093/bioinformatics/bti525
Wright S. The genetical structure of populations. Annals of Eugenics. 1951; 15(1): 323-354. doi: 10.1111/j.1469-1809.1949.tb02451.x
Jost L. GST and its relatives do not measure differentiation. Molecular Ecology. 2008; 17(18): 4015-4026. doi: 10.1111/j.1365-294X.2008.03887.x