Logo
User Name

Enes Makalic.

Društvene mreže:

Luca Li, Elaheh Zarean, D. L. Li, Yiyue Zhu, Shuai Li, E. Makalic, C. McLean, G. Giles, R. Milne et al.

Breast cancer remains a major challenge to public health. Biomarkers may be useful to improve prediction of breast cancer survival. Several epigenetic markers of cell division and ageing based on DNA methylation have been proposed. In this study, we measured these epigenetic markers in breast tumours and assessed their prognostic value. We used genome-wide DNA methylation data measured in 1992 breast cancer tumours from the Melbourne Collaborative Cohort Study and publicly available datasets. We calculated four markers of cell division (epiTOC2, stemTOC, MiAge and CellDRIFT), two markers of chronological age (Horvath age and BTEC), and four markers of biological age (PhenoAge, GrimAge, MRscore and DunedinPACE). Cox regression models were used to assess the associations of age-adjusted epigenetic markers with 5-year overall survival, with adjustment for clinical variables. Effect modification by estrogen receptor (ER) status and molecular subtype was also investigated. After adjustment for age and stratification by study, higher levels of cell division markers were associated with poorer survival (e.g., epiTOC2: per one-standard-deviation increase, hazard ratio [HR] = 1.14, 95% CI: 1.03-1.27), whereas higher chronological age markers were linked to better prognosis (e.g., Horvath age: HR = 0.70, 95% CI: 0.61-0.82). Epigenetic markers of biological age showed variable associations. These associations were partly explained by the main clinicopathological variables at diagnosis and varied across subtypes. Our study revealed associations of several epigenetic markers with breast cancer survival. The associations were quite weak, suggesting these markers may have limited prognostic value. The varying associations observed among subtypes may reflect underlying biological differences that should be further investigated.

Elaheh Zarean, Shuai Li, E. Makalic, R. Milne, G. Giles, C. McLean, M. Southey, P. Dugué

BACKGROUND Tumour DNA methylation is a potentially valuable marker of breast cancer survival, but previous studies have been limited by relatively small sample sizes. This study aimed to use a large sample size to identify survival-associated DNA methylation markers and develop a methylation-based signature predictive of breast cancer survival. METHODS We used DNA methylation data from 2,157 breast tumours collected in the Melbourne Collaborative Cohort Study (MCCS) and nine publicly available datasets. An epigenome-wide association study (EWAS) was conducted for five-year overall survival (N = 1,992 cases). Subgroup analyses were carried out by estrogen receptor status. Pathway enrichment analyses were conducted to identify related biological pathways. Elastic net Cox regression was used to develop a signature of survival from DNA methylation data and clinical characteristics, trained on publicly available datasets and tested in the MCCS (N = 425) in addition to the main clinical characteristics. A set of triple-negative tumours (N = 165) with disease-free survival as the outcome was used for additional replication. RESULTS We identified 2,535 CpGs showing an association (P < 1 × 10- 7) with survival in age-adjusted models, including 281 after adjustment for estrogen receptor (ER) status. In ER-positive tumours, there were 389 associations, of which 315 were absent from the primary EWAS. A 228-CpG signature showed a strong association with survival in the MCCS after adjustment for clinical characteristics: per SD, HR = 1.7, 95%CI: 1.2-2.3, P = 0.001, with a C-index of 0.79, compared with a C-index of 0.75 without methylation information (C-index increase: 0.04, 95%CI: 0.01-0.10). The signature was also predictive of disease-free survival beyond clinical characteristics in triple-negative tumours (HR per SD = 1.4, 95% CI: 1.1-1.8). CONCLUSION Our findings revealed numerous CpG sites where tumour DNA methylation was associated with breast cancer survival. A substantial improvement in the accuracy of five-year overall survival prediction was achieved by adding a 228-CpG epigenetic score to the main clinicopathological variables. These results suggest that DNA methylation markers may provide additional prognostic information beyond established clinicopathological factors.

Astrid K. M. Stubbusch, Caitlin Welsh, Lianxin Li, Y. Katayama, Edward M. Giles, Manh Vu, E. Makalic, Vanessa R. Marcelino, Samuel C. Forster et al.

Molecular hydrogen (H2) and hydrogen sulfide (H2S) are central gut metabolites that shape microbial metabolism and affect host health. In Crohn’s disease (CD), the shift in microbiota composition (‘dysbiosis’) is associated with intestinal accumulation of these gases, but the responsible microbes remain poorly resolved. Here, we analysed 4,644 bacterial and archaeal species-level genomes from the Unified Human Gastrointestinal Genome Collection to identify H2-cycling microbes, assessed their prevalence in ca. 1,700 stool metagenomes from healthy and diseased individuals, and validated their activity using culture-based incubations of stool isolates and biopsy samples. Approximately half of all species encoded H2-producing abilities, with acetate- and propionate-forming fermenters such as Phocaeicola and Bacteroides dominating healthy cohorts, whereas comparatively few taxa, including Escherichia and Megamonas, encoded H2 consuming abilities. In CD, H2 producers became more abundant but less diverse, favouring species with multiple H2-evolving hydrogenases and more fermentation routes, especially Clostridium and Enterocloster species. Consistently, isolates enriched in CD produced H2 faster and at higher concentrations than health-associated isolates. Increased H2S-producing capacity in CD was driven mainly by these H2-producing fermenters carrying anaerobic sulfite reductases (Asr), rather than sulfate-reducing bacteria, and was supported by elevated H2S production in Asr-positive isolates, likely providing an additional electron sink. These findings provide a species-resolved view of gut gas metabolism and implicate metabolically flexible fermenters in excessive gas and sulfide production in gut disorders.

E. Makalic, Daniel F. Schmidt

We introduce entropic strict minimum message length (SMML), a risk-sensitive generalization of strict minimum message length coding. The proposed criterion replaces expected two-part codelength under the prior predictive distribution with an exponential certainty equivalent, thereby defining a one-parameter family of coding rules that interpolates between Bayesian average-case coding and worst-case minimax coding. We show that ordinary SMML is recovered in the risk-neutral limit, while the extreme risk-sensitive limit yields a minimax codelength criterion. Applying the same entropic soft maximum to regret relative to the oracle maximum likelihood codelength recovers the normalized maximum likelihood (NML) minimax-regret principle. We further prove that entropic SMML admits a variational characterization as a Kullback--Leibler-regularized worst-case expected codelength, giving it a PAC--Bayes-type interpretation. We establish joint \(n\)--\(\tau\) asymptotics that identify how the risk parameter must scale with sample size in order to recover Bayesian average-case, intermediate robust, and worst-case minimax coding behavior. For regular exponential families, the fixed-codebook partition remains affine in sufficient-statistic space, while the codepoints satisfy a tilted moment-matching condition and admit an interpretation as tilted Bregman centroids. These results position entropic SMML as an information theoretic bridge between MML, PAC--Bayes, and MDL.

E. Makalic, Daniel F. Schmidt

Strict minimum message length (SMML) is an information-theoretic coding principle that represents a continuous statistical model by a finite set of assertions and a partition of the sample space. We show that the SMML objective decomposes into assertion entropy and conditional cross-entropy, balancing the cost of identifying an assertion against the cost of encoding data under the assigned model. For any fixed partition, the optimal codepoint for each cell is the model distribution that minimises Kullback–Leibler (KL) divergence from the data distribution restricted to that cell. Using the local Fisher–Rao geometry of regular parametric models, we show that, under a high-resolution LAN-scale regime, SMML partitions are asymptotically the pullback, through the maximum-likelihood estimator, of weighted Fisher–Rao Voronoi tessellations in parameter space, with assertion probabilities appearing as additive weights. For regular canonical exponential families, SMML codepoints satisfy a moment-matching condition and admit an interpretation as KL/Bregman centroids, while exact SMML cells are pullbacks of convex polyhedra in sufficient-statistic space. Together, these results show that SMML induces a natural information-geometric quantisation linking entropy-based coding, KL projection, and divergence-based Voronoi geometry.

E. Makalic, Daniel F. Schmidt

The Wallace--Freeman estimator is a classical minimum message length estimator whose relationship with likelihood-based asymptotic theory has not been fully developed. We show that, in regular parametric models, the Wallace--Freeman criterion is equivalent, up to constants, to a penalised likelihood criterion with penalty weight \(n^{-1}\). This representation places the estimator within the standard theory of penalised M-estimation and yields existence, consistency, an asymptotic linear expansion, and asymptotic normality under regularity conditions. We further derive the first-order difference between the Wallace--Freeman estimator and the maximum likelihood estimator, showing that it is an explicit \(O(n^{-1})\) shift determined by the gradient of the Wallace--Freeman penalty. Combining this expansion with the Cox--Snell formula gives a first-order bias expansion for the Wallace--Freeman estimator. The result clarifies its relationship with maximum likelihood, Jeffreys-prior penalisation, and Firth-type bias reduction. We illustrate the theory for the Weibull model, where the penalty modifies the leading bias of the maximum likelihood estimator of the shape parameter.

Helen M L Frazer, John L Hopper, T. Nguyen, M. Elliott, Katrina M. Kunicki, Osamah M. Al-Qershi, Daniel F. Schmidt, E. Makalic, Shuai Li et al.

BACKGROUND Artificial intelligence (AI)-based algorithms are being implemented in breast screening to detect breast cancers on mammographic images. We aimed to apply an epidemiological approach to demonstrate how a cancer detection algorithm can be leveraged as an intermediate-term predictor of breast cancer (current and 4-year risk) to deliver greater risk-based personalisation in screening mammography. METHODS In this population cohort study, we used detection scores from an AI cancer detection algorithm (BRAIx AI Reader), which was calibrated using a training dataset of 397 648 women aged 40 years to 97 years from women who screened at BreastScreen Victoria, Australia between Jan 1, 2016, and Dec 31, 2017, to create a woman-specific mammography-based score for breast cancer risk, the BRAIx risk score. Subsequently, the BRAIx risk score was evaluated on an independent test dataset of women from BreastScreen Victoria, Australia, comprising a random population cohort of 96 348 women who screened from Jan 1, 2016, to Dec 31, 2017, aged 40 years to 74 years, and an independent, external dataset from woman screened at Karolinska University Hospital, Stockholm, Sweden. We applied logistic regression, using the BRAIx risk score to estimate risks of invasive breast cancers on the test dataset: (1) detected at cohort entry (n=525); and (2) for women given an all clear, diagnosed during the next 4 years either at future screens (n=790) or during intervals between screens (n=308). We also trained full multivariate risk models (logistic regression and elastic net) using the training dataset and evaluated their predictive performance on the test and external validation data, with assessment of familial aspects of the BRAIx risk score achieved with inference about causation from examining changes in regression coefficients in an innovative statistical analysis framework. FINDINGS In both Australian and Swedish test datasets, the BRAIx risk score predicted cancer detection at cohort entry and future cancer risk (all p<0·0001). The BRAIx risk score was the strongest tested explanatory factor for cancer detection at cohort entry (odds ratio 13·80 [95% CI 9·54-20·80] in Australian data; 8·89 [3·19-37·49] in Swedish data) and for intermediate-term cancer risk (2·29 [2·13-2.47] in Australian data; 2·15 [1·85-2·50] in Swedish data). We found that adding a thresholded binary version of the BRAIx risk score significantly improved model fit (p<2·2 × 10-16, Australian and Swedish data) and women with BRAIx risk scores of more than 2 were significantly at many-fold increased risk of intermediate-term cancer than women below that threshold (12·34 [7·33-20·91], Australia; 44·7 [11·9-184·9], Sweden; p<0·0001). For the top 2% of women given an all clear with the highest BRAIx risk score, the probability of a cancer diagnosis within 4 years was 9·7%. The BRAIx risk score explained 23% of why family history predicts 4-year risk (p<0·0001). After fitting the BRAIx risk score in a multivariate model, mammographic density was no longer significantly associated with breast cancer risk in the Australian test data (p>0·05) and became associated with lower risk for intermediate-term cancer in the external Swedish test dataset (0·83 [0·73-0·95]). INTERPRETATION The BRAIx risk score is a strong intermediate-term predictor of breast cancer (current to 4-year risk). Calibrating the score on a training dataset produces population-specific probabilities for calculating individual-specific risk scores for screening clients based on their mammogram images. These risk scores enable future development of personalised screening pathways to transform population breast cancer screening and save lives. Identification of women given an all clear but at very high risk, similar to those carrying BRCA1 and BRCA2 mutations, could reveal insights into both familial and non-familial causes of breast cancer. FUNDING Australian Government Medical Research Future Fund, the Ramaciotti Foundation, the National Breast Cancer Foundation, Cancer Australia, and the National Health and Medical Research Council.

J. Meneses, C. Tejos, E. Makalic, Sergio Uribe

Liver proton density fat fraction (PDFF), the ratio between fat-only and overall proton densities, is an extensively validated biomarker associated with several diseases. In recent years, numerous deep learning-based methods for estimating PDFF have been proposed to optimize acquisition and post-processing times without sacrificing accuracy, compared to conventional methods. However, the lack of interpretability and the often poor generalizability of these DL-based models undermine the adoption of such techniques in clinical practice. In this work, we propose an Artificial Intelligence-based Decomposition of water and fat with Echo Asymmetry and Least-squares (AI-DEAL) method, designed to estimate both proton density fat fraction (PDFF) and the associated uncertainty maps. Once trained, AI-DEAL performs a one-shot MRI water-fat separation by first calculating the nonlinear confounder variables, R2∗ and off-resonance field. It then employs a weighted least squares approach to compute water-only and fat-only signals, along with their corresponding covariance matrix, which are subsequently used to derive the PDFF and its associated uncertainty. We validated our method using in vivo liver CSE-MRI, a fat-water phantom, and a numerical phantom. AI-DEAL demonstrated PDFF biases of 0.25% and -0.12% at two liver ROIs, outperforming state-of-the-art deep learning-based techniques. Although trained using in vivo data, our method exhibited PDFF biases of -3.43% in the fat-water phantom and -0.22% in the numerical phantom with no added noise. The latter bias remained approximately constant when noise was introduced. Furthermore, the estimated uncertainties showed good agreement with the observed errors and the variations within each ROI, highlighting their potential value for assessing the reliability of the resulting PDFF maps.

Ye Zhang, Amalia Karahalios, A. Win, E. Makalic, Alex Boussioutas, D. Buchanan, S. Schmit, N. Samadder, Finlay A. Macrae et al.

Abstract Background Being able to estimate the risk of metachronous disease in a patient with colorectal cancer (CRC) could enable risk-appropriate surveillance. The aim of this study was to develop a risk-prediction model to estimate individual 10-year risk of metachronous disease following a CRC diagnosis. Methods A population-based cohort of patients with CRC was recruited soon after diagnosis between 1997 and 2012 from the United States, Canada, and Australia. Cox regression with the least absolute shrinkage and selection operator penalization was used to identify factors that predicted the risk of a new primary CRC diagnosed at least 1 year after the initial CRC diagnosis. Potential predictors included demography, anthropometry, lifestyle factors, comorbidities, personal and family cancer history, medication use, age at diagnosis, and pathological features of the first CRC. Internal validation through bootstrapping was used to evaluate the discrimination and calibration. Results We included 6085 CRC cases; 138 (2.3%) of these cases were diagnosed with metachronous disease over a median of 12 years (IQR = 5-17 years). Metachronous CRC risk was predicted by body mass index; smoking status; level of physical activity; family history of cancer and synchronous CRC; stage, grade, histological type, and DNA mismatch repair status; and age at diagnosis of the first CRC. The model was valid with a C statistic of 0.65 (95% CI = 0.63 to 0.68) and a calibration slope of 0.873 (SD = 0.087). Conclusions Metachronous CRC can be predicted with reasonable accuracy using a prediction model that consists of clinical variables collected as part of routine practice.

Max Schuran, Benjamin Goudey, G. Dite, E. Makalic

Abstract Polygenic risk scores (PRS) combine the effects of multiple genetic variants to predict an individual’s genetic predisposition to a disease. PRS typically rely on linear models, which assume that all genetic variants act independently. They often fall short in predictive accuracy and are not able to explain the genetic variability of a trait to the full extent. There is growing interest in applying deep learning neural networks to model PRS given their ability to model non-linear relationships and strong performance in other domains. We conducted a survey of the literature to investigate how neural networks model PRS. We categorize deep learning-based approaches by their underlying architecture, highlighting their modeling assumptions, likely strengths and potential weaknesses of the architectures. Several categories of neural network architectures exhibited promising signs for the improvement of PRS’ predictive power, namely sequence-based architectures, graph neural networks and those that incorporated biological knowledge. Additionally, the use of latent representations in autoencoders has improved predictive performance across diverse ancestries. However, a lack of existing model benchmarks on consistent datasets and phenotypes makes it challenging to understand the extent to which different architectures improve performance. Interpretability of deep learning-based PRS is also challenging with great care required when inferring causation. To address these challenges, we suggest the establishment and adherence to reporting standards and benchmarks to aid the development of deep learning-based PRS to find quantifiable trends in neural network architectures.

...
...
...

Pretplatite se na novosti o BH Akademskom Imeniku

Ova stranica koristi kolačiće da bi vam pružila najbolje iskustvo

Saznaj više