Minimum message length is a general Bayesian principle for model selection and parameter estimation that is based on information theory. This paper applies the minimum message length principle to a small-sample model selection problem involving Poisson and geometric data models. Since MML is a Bayesian principle, it requires prior distributions for all model parameters. We introduce three candidate prior distributions for the unknown model parameters with both light- and heavy-tails. The performance of the MML methods is compared with objective Bayesian inference and minimum description length techniques based on the normalized maximum likelihood code. Simulations show that our MML approach with a heavy-tail prior distribution provides an excellent performance in all tests.
Global measures of peripheral blood DNA methylation have been associated with risk of some malignancies, including breast, bladder, and gastric cancer. Here, we examined genome‐wide measures of peripheral blood DNA methylation in prostate cancer and its non‐aggressive and aggressive disease forms.
Inference of complex hierarchical models is an increasingly common problem in modern Bayesian data analysis. Unfortunately, there are few computationally efficient and widely applicable methods for selecting between competing hierarchical models. In this paper we adapt ideas from the information theoretic minimum message length principle and propose a powerful yet simple model selection criteria for general hierarchical Bayesian models called MML-h. Computation of this criterion requires only that a set of samples from the posterior distribution be available. The flexibility of this new algorithm is demonstrated by a novel application to state-of-the-art Bayesian hierarchical regression estimation. Simulations show that the MML-h criterion is able to adaptively select between classic ridge regression and sparse horseshoe regression estimators, and the resulting procedure exhibits excellent robustness to the underlying structure of the regression coefficients.
The Bayesian horseshoe estimator is known for its robustness when handling noisy and sparse big data problems. This paper presents two extensions of the regular Bayesian horseshoe: (i) the grouped Bayesian horseshoe and (ii) the hierarchical Bayesian grouped horseshoe. The advantages of the proposed methods are their flexibility in handling grouped variables through extra shrinkage parameters at the group and within-group levels. We apply the proposed methods to the important class of additive models where group structures naturally exist, and we demonstrate that the grouped hierarchical Bayesian horseshoe has promising performance on both simulated and real data.
The horseshoe\(+\) estimator for Gaussian linear regression models is a novel extension of the horseshoe estimator that enjoys many favourable theoretical properties. We develop the first efficient Gibbs sampling algorithm for the horseshoe\(+\) estimator for linear and logistic regression models. Importantly, our sampling algorithm incorporates robust data models that naturally handle non-Gaussian data and are less sensitive to outliers. The resulting software implementation provides a powerful, flexible and robust tool for building prediction and classification models from potentially high-dimensional data and represents the state-of-the-art in Bayesian machine learning techniques.
Bayesian penalized regression techniques, such as the Bayesian lasso and the Bayesian horseshoe estimator, have recently received a significant amount of attention in the statistics literature. However, software implementing state-of-the-art Bayesian penalized regression, outside of general purpose Markov chain Monte Carlo platforms such as STAN, is relatively rare. This paper introduces bayesreg, a new toolbox for fitting Bayesian penalized regression models with continuous shrinkage prior densities. The toolbox features Bayesian linear regression with Gaussian or heavy-tailed error models and Bayesian logistic regression with ridge, lasso, horseshoe and horseshoe$+$ estimators. The toolbox is free, open-source and available for use with the MATLAB and R numerical platforms.
The genetic architecture of human reproductive behavior—age at first birth (AFB) and number of children ever born (NEB)—has a strong relationship with fitness, human development, infertility and risk of neuropsychiatric disorders. However, very few genetic loci have been identified, and the underlying mechanisms of AFB and NEB are poorly understood. We report a large genome-wide association study of both sexes including 251,151 individuals for AFB and 343,072 individuals for NEB. We identified 12 independent loci that are significantly associated with AFB and/or NEB in a SNP-based genome-wide association study and 4 additional loci associated in a gene-based effort. These loci harbor genes that are likely to have a role, either directly or by affecting non-local gene expression, in human reproduction and infertility, thereby increasing understanding of these complex traits.
Background: We have developed a genome-wide association study analysis method called DEPTH (DEPendency of association on the number of Top Hits) to identify genomic regions potentially associated with disease by considering overlapping groups of contiguous markers (e.g., SNPs) across the genome. DEPTH is a machine learning algorithm for feature ranking of ultra-high dimensional datasets, built from well-established statistical tools such as bootstrapping, penalized regression, and decision trees. Unlike marginal regression, which considers each SNP individually, the key idea behind DEPTH is to rank groups of SNPs in terms of their joint strength of association with the outcome. Our aim was to compare the performance of DEPTH with that of standard logistic regression analysis. Methods: We selected 1,854 prostate cancer cases and 1,894 controls from the UK for whom 541,129 SNPs were measured using the Illumina Infinium HumanHap550 array. Confirmation was sought using 4,152 cases and 2,874 controls, ascertained from the UK and Australia, for whom 211,155 SNPs were measured using the iCOGS Illumina Infinium array. Results: From the DEPTH analysis, we identified 14 regions associated with prostate cancer risk that had been reported previously, five of which would not have been identified by conventional logistic regression. We also identified 112 novel putative susceptibility regions. Conclusions: DEPTH can reveal new risk-associated regions that would not have been identified using a conventional logistic regression analysis of individual SNPs. Impact: This study demonstrates that the DEPTH algorithm could identify additional genetic susceptibility regions that merit further investigation. Cancer Epidemiol Biomarkers Prev; 25(12); 1619–24. ©2016 AACR.
Background:Global DNA methylation has been reported to be associated with urothelial cell carcinoma (UCC) by studies using blood samples collected at diagnosis. Using the Illumina HumanMethylation450 assay, we derived genome-wide measures of blood DNA methylation and assessed them for their prospective association with UCC risk.Methods:We used 439 case–control pairs from the Melbourne Collaborative Cohort Study matched on age, sex, country of birth, DNA sample type, and collection period. Conditional logistic regression was used to compute odds ratios (OR) of UCC risk per s.d. of each genome-wide measure of DNA methylation and 95% confidence intervals (CIs), adjusted for potential confounders. We also investigated associations by disease subtype, sex, smoking, and time since blood collection.Results:The risk of superficial UCC was decreased for individuals with higher levels of our genome-wide DNA methylation measure (OR=0.71, 95% CI: 0.54–0.94; P=0.02). This association was particularly strong for current smokers at sample collection (OR=0.47, 95% CI: 0.27–0.83). Intermediate levels of our genome-wide measure were associated with decreased risk of invasive UCC. Some variation was observed between UCC subtypes and the location and regulatory function of the CpGs included in the genome-wide measures of methylation.Conclusions:Higher levels of our genome-wide DNA methylation measure were associated with decreased risk of superficial UCC and intermediate levels were associated with reduced risk of invasive disease. These findings require replication by other prospective studies.
Common variants in 94 loci have been associated with breast cancer including 15 loci with genome-wide significant associations (P<5 × 10−8) with oestrogen receptor (ER)-negative breast cancer and BRCA1-associated breast cancer risk. In this study, to identify new ER-negative susceptibility loci, we performed a meta-analysis of 11 genome-wide association studies (GWAS) consisting of 4,939 ER-negative cases and 14,352 controls, combined with 7,333 ER-negative cases and 42,468 controls and 15,252 BRCA1 mutation carriers genotyped on the iCOGS array. We identify four previously unidentified loci including two loci at 13q22 near KLF5, a 2p23.2 locus near WDR43 and a 2q33 locus near PPIL3 that display genome-wide significant associations with ER-negative breast cancer. In addition, 19 known breast cancer risk loci have genome-wide significant associations and 40 had moderate associations (P<0.05) with ER-negative disease. Using functional and eQTL studies we implicate TRMT61B and WDR43 at 2p23.2 and PPIL3 at 2q33 in ER-negative breast cancer aetiology. All ER-negative loci combined account for ∼11% of familial relative risk for ER-negative disease and may contribute to improved ER-negative and BRCA1 breast cancer risk prediction. Oestrogen negative breast cancer is associated with a poor prognosis. In this study, the authors perform a meta-analysis of 11 breast cancer genome-wide association studies and identify four new loci associated with oestrogen negative breast cancer risk. These findings may aid in stratifying patients in the clinic.
Identifying genetic variants with pleiotropic associations can uncover common pathways influencing multiple cancers. We took a two-staged approach to conduct genome-wide association studies for lung, ovary, breast, prostate and colorectal cancer from the GAME-ON/GECCO Network (61,851 cases, 61,820 controls) to identify pleiotropic loci. Findings were replicated in independent association studies (55,789 cases, 330,490 controls). We identified a novel pleiotropic association at 1q22 involving breast and lung squamous cell carcinoma, with eQTL analysis showing an association with ADAM15/THBS3 gene expression in lung. We also identified a known breast cancer locus CASP8/ALS2CR12 associated with prostate cancer, a known cancer locus at CDKN2B-AS1 with different variants associated with lung adenocarcinoma and prostate cancer and confirmed the associations of a breast BRCA2 locus with lung and serous ovarian cancer. This is the largest study to date examining pleiotropy across multiple cancer-associated loci, identifying common mechanisms of cancer development and progression.
Nema pronađenih rezultata, molimo da izmjenite uslove pretrage i pokušate ponovo!
Ova stranica koristi kolačiće da bi vam pružila najbolje iskustvo
Saznaj više