Logo

Publikacije (195)

Nazad
Luca Li, Elaheh Zarean, D. L. Li, Yiyue Zhu, Shuai Li, E. Makalic, C. McLean, G. Giles et al.

Breast cancer remains a major challenge to public health. Biomarkers may be useful to improve prediction of breast cancer survival. Several epigenetic markers of cell division and ageing based on DNA methylation have been proposed. In this study, we measured these epigenetic markers in breast tumours and assessed their prognostic value. We used genome-wide DNA methylation data measured in 1992 breast cancer tumours from the Melbourne Collaborative Cohort Study and publicly available datasets. We calculated four markers of cell division (epiTOC2, stemTOC, MiAge and CellDRIFT), two markers of chronological age (Horvath age and BTEC), and four markers of biological age (PhenoAge, GrimAge, MRscore and DunedinPACE). Cox regression models were used to assess the associations of age-adjusted epigenetic markers with 5-year overall survival, with adjustment for clinical variables. Effect modification by estrogen receptor (ER) status and molecular subtype was also investigated. After adjustment for age and stratification by study, higher levels of cell division markers were associated with poorer survival (e.g., epiTOC2: per one-standard-deviation increase, hazard ratio [HR] = 1.14, 95% CI: 1.03-1.27), whereas higher chronological age markers were linked to better prognosis (e.g., Horvath age: HR = 0.70, 95% CI: 0.61-0.82). Epigenetic markers of biological age showed variable associations. These associations were partly explained by the main clinicopathological variables at diagnosis and varied across subtypes. Our study revealed associations of several epigenetic markers with breast cancer survival. The associations were quite weak, suggesting these markers may have limited prognostic value. The varying associations observed among subtypes may reflect underlying biological differences that should be further investigated.

Elaheh Zarean, Shuai Li, E. Makalic, R. Milne, G. Giles, C. McLean, M. Southey, P. Dugué

BACKGROUND Tumour DNA methylation is a potentially valuable marker of breast cancer survival, but previous studies have been limited by relatively small sample sizes. This study aimed to use a large sample size to identify survival-associated DNA methylation markers and develop a methylation-based signature predictive of breast cancer survival. METHODS We used DNA methylation data from 2,157 breast tumours collected in the Melbourne Collaborative Cohort Study (MCCS) and nine publicly available datasets. An epigenome-wide association study (EWAS) was conducted for five-year overall survival (N = 1,992 cases). Subgroup analyses were carried out by estrogen receptor status. Pathway enrichment analyses were conducted to identify related biological pathways. Elastic net Cox regression was used to develop a signature of survival from DNA methylation data and clinical characteristics, trained on publicly available datasets and tested in the MCCS (N = 425) in addition to the main clinical characteristics. A set of triple-negative tumours (N = 165) with disease-free survival as the outcome was used for additional replication. RESULTS We identified 2,535 CpGs showing an association (P < 1 × 10- 7) with survival in age-adjusted models, including 281 after adjustment for estrogen receptor (ER) status. In ER-positive tumours, there were 389 associations, of which 315 were absent from the primary EWAS. A 228-CpG signature showed a strong association with survival in the MCCS after adjustment for clinical characteristics: per SD, HR = 1.7, 95%CI: 1.2-2.3, P = 0.001, with a C-index of 0.79, compared with a C-index of 0.75 without methylation information (C-index increase: 0.04, 95%CI: 0.01-0.10). The signature was also predictive of disease-free survival beyond clinical characteristics in triple-negative tumours (HR per SD = 1.4, 95% CI: 1.1-1.8). CONCLUSION Our findings revealed numerous CpG sites where tumour DNA methylation was associated with breast cancer survival. A substantial improvement in the accuracy of five-year overall survival prediction was achieved by adding a 228-CpG epigenetic score to the main clinicopathological variables. These results suggest that DNA methylation markers may provide additional prognostic information beyond established clinicopathological factors.

Astrid K. M. Stubbusch, Caitlin Welsh, Lianxin Li, Y. Katayama, Edward M. Giles, Manh Vu, E. Makalic, Vanessa R. Marcelino et al.

Molecular hydrogen (H2) and hydrogen sulfide (H2S) are central gut metabolites that shape microbial metabolism and affect host health. In Crohn’s disease (CD), the shift in microbiota composition (‘dysbiosis’) is associated with intestinal accumulation of these gases, but the responsible microbes remain poorly resolved. Here, we analysed 4,644 bacterial and archaeal species-level genomes from the Unified Human Gastrointestinal Genome Collection to identify H2-cycling microbes, assessed their prevalence in ca. 1,700 stool metagenomes from healthy and diseased individuals, and validated their activity using culture-based incubations of stool isolates and biopsy samples. Approximately half of all species encoded H2-producing abilities, with acetate- and propionate-forming fermenters such as Phocaeicola and Bacteroides dominating healthy cohorts, whereas comparatively few taxa, including Escherichia and Megamonas, encoded H2 consuming abilities. In CD, H2 producers became more abundant but less diverse, favouring species with multiple H2-evolving hydrogenases and more fermentation routes, especially Clostridium and Enterocloster species. Consistently, isolates enriched in CD produced H2 faster and at higher concentrations than health-associated isolates. Increased H2S-producing capacity in CD was driven mainly by these H2-producing fermenters carrying anaerobic sulfite reductases (Asr), rather than sulfate-reducing bacteria, and was supported by elevated H2S production in Asr-positive isolates, likely providing an additional electron sink. These findings provide a species-resolved view of gut gas metabolism and implicate metabolically flexible fermenters in excessive gas and sulfide production in gut disorders.

E. Makalic, Daniel F. Schmidt

We introduce entropic strict minimum message length (SMML), a risk-sensitive generalization of strict minimum message length coding. The proposed criterion replaces expected two-part codelength under the prior predictive distribution with an exponential certainty equivalent, thereby defining a one-parameter family of coding rules that interpolates between Bayesian average-case coding and worst-case minimax coding. We show that ordinary SMML is recovered in the risk-neutral limit, while the extreme risk-sensitive limit yields a minimax codelength criterion. Applying the same entropic soft maximum to regret relative to the oracle maximum likelihood codelength recovers the normalized maximum likelihood (NML) minimax-regret principle. We further prove that entropic SMML admits a variational characterization as a Kullback--Leibler-regularized worst-case expected codelength, giving it a PAC--Bayes-type interpretation. We establish joint \(n\)--\(\tau\) asymptotics that identify how the risk parameter must scale with sample size in order to recover Bayesian average-case, intermediate robust, and worst-case minimax coding behavior. For regular exponential families, the fixed-codebook partition remains affine in sufficient-statistic space, while the codepoints satisfy a tilted moment-matching condition and admit an interpretation as tilted Bregman centroids. These results position entropic SMML as an information theoretic bridge between MML, PAC--Bayes, and MDL.

E. Makalic, Daniel F. Schmidt

Strict minimum message length (SMML) is an information-theoretic coding principle that represents a continuous statistical model by a finite set of assertions and a partition of the sample space. We show that the SMML objective decomposes into assertion entropy and conditional cross-entropy, balancing the cost of identifying an assertion against the cost of encoding data under the assigned model. For any fixed partition, the optimal codepoint for each cell is the model distribution that minimises Kullback–Leibler (KL) divergence from the data distribution restricted to that cell. Using the local Fisher–Rao geometry of regular parametric models, we show that, under a high-resolution LAN-scale regime, SMML partitions are asymptotically the pullback, through the maximum-likelihood estimator, of weighted Fisher–Rao Voronoi tessellations in parameter space, with assertion probabilities appearing as additive weights. For regular canonical exponential families, SMML codepoints satisfy a moment-matching condition and admit an interpretation as KL/Bregman centroids, while exact SMML cells are pullbacks of convex polyhedra in sufficient-statistic space. Together, these results show that SMML induces a natural information-geometric quantisation linking entropy-based coding, KL projection, and divergence-based Voronoi geometry.

E. Makalic, Daniel F. Schmidt

The Wallace--Freeman estimator is a classical minimum message length estimator whose relationship with likelihood-based asymptotic theory has not been fully developed. We show that, in regular parametric models, the Wallace--Freeman criterion is equivalent, up to constants, to a penalised likelihood criterion with penalty weight \(n^{-1}\). This representation places the estimator within the standard theory of penalised M-estimation and yields existence, consistency, an asymptotic linear expansion, and asymptotic normality under regularity conditions. We further derive the first-order difference between the Wallace--Freeman estimator and the maximum likelihood estimator, showing that it is an explicit \(O(n^{-1})\) shift determined by the gradient of the Wallace--Freeman penalty. Combining this expansion with the Cox--Snell formula gives a first-order bias expansion for the Wallace--Freeman estimator. The result clarifies its relationship with maximum likelihood, Jeffreys-prior penalisation, and Firth-type bias reduction. We illustrate the theory for the Weibull model, where the penalty modifies the leading bias of the maximum likelihood estimator of the shape parameter.

Helen M L Frazer, John L Hopper, T. Nguyen, M. Elliott, Katrina M. Kunicki, Osamah M. Al-Qershi, Daniel F. Schmidt, E. Makalic et al.

BACKGROUND Artificial intelligence (AI)-based algorithms are being implemented in breast screening to detect breast cancers on mammographic images. We aimed to apply an epidemiological approach to demonstrate how a cancer detection algorithm can be leveraged as an intermediate-term predictor of breast cancer (current and 4-year risk) to deliver greater risk-based personalisation in screening mammography. METHODS In this population cohort study, we used detection scores from an AI cancer detection algorithm (BRAIx AI Reader), which was calibrated using a training dataset of 397 648 women aged 40 years to 97 years from women who screened at BreastScreen Victoria, Australia between Jan 1, 2016, and Dec 31, 2017, to create a woman-specific mammography-based score for breast cancer risk, the BRAIx risk score. Subsequently, the BRAIx risk score was evaluated on an independent test dataset of women from BreastScreen Victoria, Australia, comprising a random population cohort of 96 348 women who screened from Jan 1, 2016, to Dec 31, 2017, aged 40 years to 74 years, and an independent, external dataset from woman screened at Karolinska University Hospital, Stockholm, Sweden. We applied logistic regression, using the BRAIx risk score to estimate risks of invasive breast cancers on the test dataset: (1) detected at cohort entry (n=525); and (2) for women given an all clear, diagnosed during the next 4 years either at future screens (n=790) or during intervals between screens (n=308). We also trained full multivariate risk models (logistic regression and elastic net) using the training dataset and evaluated their predictive performance on the test and external validation data, with assessment of familial aspects of the BRAIx risk score achieved with inference about causation from examining changes in regression coefficients in an innovative statistical analysis framework. FINDINGS In both Australian and Swedish test datasets, the BRAIx risk score predicted cancer detection at cohort entry and future cancer risk (all p<0·0001). The BRAIx risk score was the strongest tested explanatory factor for cancer detection at cohort entry (odds ratio 13·80 [95% CI 9·54-20·80] in Australian data; 8·89 [3·19-37·49] in Swedish data) and for intermediate-term cancer risk (2·29 [2·13-2.47] in Australian data; 2·15 [1·85-2·50] in Swedish data). We found that adding a thresholded binary version of the BRAIx risk score significantly improved model fit (p<2·2 × 10-16, Australian and Swedish data) and women with BRAIx risk scores of more than 2 were significantly at many-fold increased risk of intermediate-term cancer than women below that threshold (12·34 [7·33-20·91], Australia; 44·7 [11·9-184·9], Sweden; p<0·0001). For the top 2% of women given an all clear with the highest BRAIx risk score, the probability of a cancer diagnosis within 4 years was 9·7%. The BRAIx risk score explained 23% of why family history predicts 4-year risk (p<0·0001). After fitting the BRAIx risk score in a multivariate model, mammographic density was no longer significantly associated with breast cancer risk in the Australian test data (p>0·05) and became associated with lower risk for intermediate-term cancer in the external Swedish test dataset (0·83 [0·73-0·95]). INTERPRETATION The BRAIx risk score is a strong intermediate-term predictor of breast cancer (current to 4-year risk). Calibrating the score on a training dataset produces population-specific probabilities for calculating individual-specific risk scores for screening clients based on their mammogram images. These risk scores enable future development of personalised screening pathways to transform population breast cancer screening and save lives. Identification of women given an all clear but at very high risk, similar to those carrying BRCA1 and BRCA2 mutations, could reveal insights into both familial and non-familial causes of breast cancer. FUNDING Australian Government Medical Research Future Fund, the Ramaciotti Foundation, the National Breast Cancer Foundation, Cancer Australia, and the National Health and Medical Research Council.

J. Meneses, C. Tejos, E. Makalic, Sergio Uribe

Liver proton density fat fraction (PDFF), the ratio between fat-only and overall proton densities, is an extensively validated biomarker associated with several diseases. In recent years, numerous deep learning-based methods for estimating PDFF have been proposed to optimize acquisition and post-processing times without sacrificing accuracy, compared to conventional methods. However, the lack of interpretability and the often poor generalizability of these DL-based models undermine the adoption of such techniques in clinical practice. In this work, we propose an Artificial Intelligence-based Decomposition of water and fat with Echo Asymmetry and Least-squares (AI-DEAL) method, designed to estimate both proton density fat fraction (PDFF) and the associated uncertainty maps. Once trained, AI-DEAL performs a one-shot MRI water-fat separation by first calculating the nonlinear confounder variables, R2∗ and off-resonance field. It then employs a weighted least squares approach to compute water-only and fat-only signals, along with their corresponding covariance matrix, which are subsequently used to derive the PDFF and its associated uncertainty. We validated our method using in vivo liver CSE-MRI, a fat-water phantom, and a numerical phantom. AI-DEAL demonstrated PDFF biases of 0.25% and -0.12% at two liver ROIs, outperforming state-of-the-art deep learning-based techniques. Although trained using in vivo data, our method exhibited PDFF biases of -3.43% in the fat-water phantom and -0.22% in the numerical phantom with no added noise. The latter bias remained approximately constant when noise was introduced. Furthermore, the estimated uncertainties showed good agreement with the observed errors and the variations within each ROI, highlighting their potential value for assessing the reliability of the resulting PDFF maps.

Ye Zhang, Amalia Karahalios, A. Win, E. Makalic, Alex Boussioutas, D. Buchanan, S. Schmit, N. Samadder et al.

Abstract Background Being able to estimate the risk of metachronous disease in a patient with colorectal cancer (CRC) could enable risk-appropriate surveillance. The aim of this study was to develop a risk-prediction model to estimate individual 10-year risk of metachronous disease following a CRC diagnosis. Methods A population-based cohort of patients with CRC was recruited soon after diagnosis between 1997 and 2012 from the United States, Canada, and Australia. Cox regression with the least absolute shrinkage and selection operator penalization was used to identify factors that predicted the risk of a new primary CRC diagnosed at least 1 year after the initial CRC diagnosis. Potential predictors included demography, anthropometry, lifestyle factors, comorbidities, personal and family cancer history, medication use, age at diagnosis, and pathological features of the first CRC. Internal validation through bootstrapping was used to evaluate the discrimination and calibration. Results We included 6085 CRC cases; 138 (2.3%) of these cases were diagnosed with metachronous disease over a median of 12 years (IQR = 5-17 years). Metachronous CRC risk was predicted by body mass index; smoking status; level of physical activity; family history of cancer and synchronous CRC; stage, grade, histological type, and DNA mismatch repair status; and age at diagnosis of the first CRC. The model was valid with a C statistic of 0.65 (95% CI = 0.63 to 0.68) and a calibration slope of 0.873 (SD = 0.087). Conclusions Metachronous CRC can be predicted with reasonable accuracy using a prediction model that consists of clinical variables collected as part of routine practice.

Max Schuran, Benjamin Goudey, G. Dite, E. Makalic

Abstract Polygenic risk scores (PRS) combine the effects of multiple genetic variants to predict an individual’s genetic predisposition to a disease. PRS typically rely on linear models, which assume that all genetic variants act independently. They often fall short in predictive accuracy and are not able to explain the genetic variability of a trait to the full extent. There is growing interest in applying deep learning neural networks to model PRS given their ability to model non-linear relationships and strong performance in other domains. We conducted a survey of the literature to investigate how neural networks model PRS. We categorize deep learning-based approaches by their underlying architecture, highlighting their modeling assumptions, likely strengths and potential weaknesses of the architectures. Several categories of neural network architectures exhibited promising signs for the improvement of PRS’ predictive power, namely sequence-based architectures, graph neural networks and those that incorporated biological knowledge. Additionally, the use of latent representations in autoencoders has improved predictive performance across diverse ancestries. However, a lack of existing model benchmarks on consistent datasets and phenotypes makes it challenging to understand the extent to which different architectures improve performance. Interpretability of deep learning-based PRS is also challenging with great care required when inferring causation. To address these challenges, we suggest the establishment and adherence to reporting standards and benchmarks to aid the development of deep learning-based PRS to find quantifiable trends in neural network architectures.

M. V. Perini, Eunice G. Lee, M. Fink, G. Starkey, Osamu Yoshino, Bartholomew Mckay, R. Furtado, E. Makalic et al.

We aim to compare the incidence and risk factors for biliary anastomotic stricture (BAS) in patients undergoing orthotopic liver transplant (OLT) with and without transcystic externalised trans‐anastomotic biliary stenting.

Daniel F. Schmidt, E. Makalic

We consider the problem of exact maximum likelihood estimation of potentially high‐order ( p>50 ) autoregressive models. We propose an extremely fast coordinate‐wise algorithm for fitting autoregressive models. This fast algorithm exploits several properties of the negative log‐likelihood when parameterised in terms of partial autocorrelations. We consider extensions to learning a single autoregressive model from multiple time series and to the more general case of regressions with autoregressive residuals. An implementation of the coordinate‐wise descent algorithm is shown to be the orders of magnitude faster than competing algorithms and appears to be the fastest known algorithm for maximum likelihood estimation of autoregressive models.

D. Karmakar, Lochana Mendis, E. Keenan, M. Palaniswami, R. Hastie, E. Makalic, Fiona C. Brownfoot

Cardiotocography (CTG) is essential for monitoring high-risk pregnancies, yet perinatal asphyxia prediction accuracy remains limited to 50–55%. Regions of artifacts (missing valid signals)-including signal processing aberrations-possibly contribute to this limitation, highlighted by 40% of FDA reports on intrapartum stillbirths. This cohort study applied causal inference to two digitized CTG databases, analyzing 36,792 labor episodes (>36 weeks) at a tertiary Australian hospital (2010–2021) and externally validating on a Czech dataset (n = 552).High rates of missing valid signals (>30% fetal heart rate signal dropout or >1% maternal-fetal heart rate coincidence) was independently associated with asphyxia (aOR 1.47, 95% CI 1.19–1.81); dropout >30% showing stronger link (aOR 1.58, 95% CI 1.13–2.20 Australian dataset; aOR 2.30, 95% CI 1.08–4.91 Czech dataset). Risk of asphyxia increased with higher dropout (>37.45%, aOR 2.21 Australian dataset; >34.01%, aOR 4.08 Czech dataset). Integrating measures of missing valid signals into fetal monitoring algorithms may improve decision-making and neonatal outcomes.

Elaheh Zarean, Shuai Li, E. Wong, E. Makalic, R. Milne, G. Giles, C. McLean, M. Southey et al.

Tumour DNA methylation has been investigated as a potential marker for breast cancer survival, but findings often lack replication across studies. This study sought to replicate previously reported associations for individual CpG sites and multi-CpG signatures using an Australian sample of 425 women with breast cancer from the Melbourne Collaborative Cohort Study (MCCS). Candidate methylation sites (N = 22) and signatures (N = 3) potentially associated with breast cancer survival were identified from five prior studies that used The Cancer Genome Atlas (TCGA) methylation dataset, which shares key characteristics with the MCCS: comparable sample size, tissue type (formalin-fixed paraffin-embedded; FFPE), technology (Illumina HumanMethylation450 array), and participant characteristics (age, ancestry, and disease subtype and severity). Cox proportional hazard regression analyses were conducted to assess associations between these markers and both breast cancer-specific survival and overall survival, adjusting for relevant participant characteristics. Our findings revealed partial replication for both individual CpG sites (9 out of 22) and multi-CpG signatures (2 out of 3). These associations were maintained after adjustment for participant characteristics and were stronger for breast cancer-specific mortality than for overall mortality. In fully-adjusted models, strong associations were observed for a CpG in PRAC2 (per standard deviation [SD], HR = 1.67, 95%CI: 1.24–2.25) and a signature based on 28 CpGs developed using elastic net (per SD, HR = 1.48, 95%CI: 1.09–2.00). While further studies are needed to confirm and expand on these findings, our study suggests that DNA methylation markers hold promise for improving breast cancer prognostication.

Nema pronađenih rezultata, molimo da izmjenite uslove pretrage i pokušate ponovo!

Pretplatite se na novosti o BH Akademskom Imeniku

Ova stranica koristi kolačiće da bi vam pružila najbolje iskustvo

Saznaj više