Popülasyon Genomiği: Temel Kavramlar ve Yöntemler

Yazarlar

Zeynel Cebeci
https://orcid.org/0000-0002-7641-7094

Özet

Bu bölüm, popülasyon genomisinin temel kavramlarını, veri türlerini ve analiz yöntemlerini kapsamlı ancak uygulamaya dönük bir yaklaşımla tanıtmaktadır. Öncelikle SNP, STR, CNV, indel, yapısal varyant ve haplotip gibi genomik veri türleri ile PLINK ve VCF başta olmak üzere yaygın veri formatları açıklanmaktadır. Güvenilir analiz sonuçları elde etmek amacıyla kalite kontrolü, filtreleme, eksik genotiplerin tamamlanması ve format dönüşümü süreçleri ele alınmaktadır. Bölümde ayrıca Temel Bileşenler Analizi, ortak soy ve karışma analizi, genetik çeşitlilik ölçütleri, FST, AMOVA, genetik uzaklık, akrabalık ve genomik ilişki matrisleri açıklanmaktadır. Doğal ve yapay seçilimin genomda bıraktığı izlerin belirlenmesine yönelik Tajima’s D, Fay-Wu’s H, EHH, iHS ve XP-EHH gibi yöntemlere de yer verilmektedir. Son olarak PLINK, GATK, bcftools, ADMIXTURE, STRUCTURE ve çeşitli R paketleri kullanılarak genomik analizlerin biyoenformatik iş akışı sunulmaktadır. Bölüm; tarım, hayvan ve bitki ıslahı, tıp, biyoçeşitliliğin korunması ve adaptasyon araştırmaları için temel bir başvuru niteliğindedir.


This chapter introduces the fundamental concepts, data types, and analytical methods of population genomics through a comprehensive yet application-oriented framework. It first describes genomic data types—including SNPs, STRs, CNVs, indels, structural variants, and haplotypes—and commonly used formats such as PLINK and VCF. To ensure reliable results, the chapter explains quality control, filtering, genotype imputation, and data-format conversion procedures. It then presents major approaches for investigating population structure and evolutionary processes, including Principal Component Analysis, ancestry and admixture analysis, genetic diversity metrics, FST, AMOVA, genetic distance, kinship, and genomic relationship matrices. Methods for detecting genomic signatures of natural and artificial selection, such as Tajima’s D, Fay and Wu’s H, EHH, iHS, and XP-EHH, are also discussed. Finally, the chapter outlines a practical bioinformatics workflow using PLINK, GATK, bcftools, ADMIXTURE, STRUCTURE, and R packages. It provides a foundational resource for studies in agriculture, animal and plant breeding, medicine, biodiversity conservation, and adaptation.

Referanslar

Luikart G, England PR, Tallmon D, et al. The power and promise of population genomics: From genotyping to genome typing. Nature Reviews. 2003; (4): 981-994. doi: 10.1038/nrg1226

Begun, DJ, Holloway, AK, Stevens K, et al. Population genomics: Whole-genome analysis of polymorphism and divergence in Drosophila simulans. PLOS Biology. 2007; 5 (11), e310. doi:10.1371/journal.pbio.0050310

Pecoraro C, Babbucci M, Franch R, et al. The population genomics of yellowfin tuna (Thunnus albacares) at global geographic scale challenges current stock delineation. Scientific Reports. 2018; 8 (1): 13890. doi:10.1038/s41598-018-32331-3

Purcell S, Neale B, Todd-Brown K, et al. PLINK: A tool set for whole-genome association and population-based linkage analyses. The American Journal of Human Genetics. 2007; 81(3): 559–575. doi: 10.1086/519795

Chang CC, Chow CC, Tellier LC, et al. Second-generation PLINK: Rising to the challenge of larger and richer datasets. GigaScience, 2015; 4(1). doi: 10.1186/s13742-015-0047-8

Danecek P, Auton A, Abecasis G, et al. The variant call format and VCFtools. Bioinformatics. 2011; 27(15): 2156–2158. doi: 10.1093/bioinformatics/btr330

DePristo MA, Banks E, Poplin R, et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nature Genetics. 2011; 43(5): 491–498. doi: 10.1038/ng.806

McKenna A, Hanna M, Banks E, et al. The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Research. 2010; 20(9): 1297–1303. doi: 10.1101/gr.107524.110

Cebeci Z, Şakiroğlu M, Bayraktar M. GWAS: Genom boyu ilişkilendirme çalışmaları. Ankara: Nobel Akademik Yayıncılık; 2024.

Meyer H. plinkQC: Genotype quality control with PLINK. R package version 1.1.0. 2026. (04/04/2026 tarihinde https://CRAN.R-project.org/package=plinkQC adresinden erişilmiştir).

Zheng X, Levine D, Shen J, et al. A high-performance computing toolset for relatedness and principal component analysis of SNP data. Bioinformatics; 2014, 28(24): 3326-3328. doi: 10.1093/bioinformatics/bts606

Jombart T. adegenet: a R package for the multivariate analysis of genetic markers. Bioinformatics. 2008; 24: 1403-1405. doi: 10.1093/bioinformatics/btn129

Gogarten SM, Bhangale T, Conomos MP, et al. GWASTools: an R/Bioconductor package for quality control and analysis of genome-wide association studies. Bioinformatics. 2012; 28(24): 3329-3331. doi: 10.1093/bioinformatics/bts610

Wickham H. ggplot2: Elegant graphics for data analysis. New York: Springer-Verlag; 2016.

Xie Y (2026). knitr: A general-purpose package for dynamic report generation in R. R package ver 1.51 (06/04/2026 tarihinde https://yihui.org/knitr adresinden erişilmiştir).

Pritchard JK, Stephens M, Donnelly P. Inference of population structure using multilocus genotype data. Genetics, 2000; 155(2): 945–959. doi: 10.1093/genetics/155.2.945

Alexander DH, Novembre J, Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Research, 2009; 19(9): 1655–1664. doi: 10.1101/gr.094052.109

Frichot E, François O. LEA: An R package for landscape and ecological association studies. Methods in Ecology and Evolution. 2015; 6(8): 925–929. doi: 10.1111/2041-210X.12382

Evanno G, Regnaut S, Goudet J. Detecting the number of clusters of individuals using the software STRUCTURE: a simulation study. Molecular Ecology, 2005; 14(8): 2611-2620.

Holsinger KE, Weir BS. Genetics in geographically structured populations: defining, estimating and interpreting Fst. Nature Reviews Genetics. 2009; 10(9): 639-650.

Nei M. Genetic distance between populations. The American Naturalist. 1972; 106(949), 283-292. doi: 10.1086/282771

Excoffier L, Smouse PE, Quattro JM. Analysis of molecular variance inferred from metric distances among DNA haplotypes: Application to human mitochondrial DNA restriction data. Genetics. 1992; 131(2): 479–491. doi:10.1093/genetics/131.2.479

Cebeci Z, Gökçe G. Genetik parametreler ve genomik tahmin. 1. Basım. Ankara: Nobel Akademik Yayıncılık; 2023.

Van Raden PM. Efficient methods to compute genomic predictions. Journal of Dairy Science. 2008, 91(11): 4414–4423. doi:10.3168/jds.2007-0980

Astle W, Balding DJ. Population structure and cryptic relatedness in genetic association studies. Statist. Sci. 2009; 24(4): 451-471. doi: 10.1214/09-STS307

Yang J, Lee SH, Goddard ME, et al. GCTA: a tool for genome-wide complex trait analysis. The American Journal of Human Genetics. 2011; 88(1): 76-82.

Endelman JB, Jannink JL. Shrinkage estimation of the genomic relationship matrix. G3: Genes, Genomes. Genetics. 2012; 2(10): 1269–1313. doi: 10.1534/g3.112.004259

Vitti JJ, Grossman SR, Sabeti PC. Detecting natural selection in genomic data. Annual Review of Genetics. 2013, 47: 97-120. doi: 10.1146/annurev-genet-111212-133526

Voight BF, Kudaravalli S, Wen X, et al. A map of recent positive selection in the human genome. PLoS Biology. 2006, 4(3): e72. doi: 10.1371/journal.pbio.0040072

Sabeti PC, Varilly BF, Lohmueller J, et al. Genome-wide detection and characterization of positive selection in human populations. Nature. 2007; 449(7164): 913-918. doi: 10.1038/nature06250

Kim Y, Stephan W. Detecting a local selective sweep in continuous efficiency landscapes. Genetics 2002; 160(2): 765-777. doi: 10.1093/genetics/160.2.765

Tajima F (1989). Statistical method for testing the neutral mutation hypothesis by DNA polymorphism. Genetics. 1989; 123(3): 585-595. doi: 10.1093/genetics/123.3.585

Watterson GA. On the number of segregating sites in genetical models without recombination. Theoretical Pop. Biol.. 1975: 7(2): 256-276.

Fay JC, Wu CI. Hitchhiking under positive Darwinian selection. Genetics. 2000; 155(3): 1405-1413. doi: 10.1093/genetics/155.3.1405

Nielsen R, Slatkin M. An introduction to population genetics: theory and applications. Sunderland, MA: Sinauer Associates; 2013.

Durinck S, Moreau Y, Kasprzyk A, et al. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis. Bioinformatics. 2005; 21 (16): 3439-3440. doi: 10.1093/bioinformatics/bti525

Wright S. The genetical structure of populations. Annals of Eugenics. 1951; 15(1): 323-354. doi: 10.1111/j.1469-1809.1949.tb02451.x

Jost L. GST and its relatives do not measure differentiation. Molecular Ecology. 2008; 17(18): 4015-4026. doi: 10.1111/j.1365-294X.2008.03887.x

Sayfalar

397-428

Yayınlanan

5 Ekim 2026

Lisans

Lisans