List of Papers Browse by Subject Areas Author List
Abstract
Multimodal alignment of histopathology encoders with transcriptomic and genomic data has been shown to significantly improve performance in downstream diagnostic tasks. Hematological cytology is unique in that visual single-cell evaluation is often paired with cytogenetics and molecular genetics for blood cancer diagnosis. In this study, we present a framework to align single white blood cell images with chromosomal aberrations (karyotype) and somatic mutations from targeted gene panels. Our training strategy follows a two-stage approach: (i) self-supervised, vision-only pretraining of a transformer aggregator using an iBOT head on a cohort of over 1500 patients, and (ii) genetic alignment via supervised contrastive loss on acute myeloid leukemia patients. Our genetically aligned patient encoder improves hematological diagnostic tasks, outperforming slide-level histopathology foundation models. Additionally, the model provides off-the-shelf retrieval capabilities for diseases and genetic alterations. Incorporating genetic data into patient encoders increases the quality of patient representations, providing a framework that aligns with clinical diagnostic workflows and paves the way for future multimodal hematology-specific AI. Source code and model weights are available at https://github.com/marrlab/GenBloom.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/5759_paper.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to the Code Repository
https://github.com/marrlab/GenBloom
https://huggingface.co/MarrLab/GenBloom
Link to the Dataset(s)
Patient embeddings used in this study: https://huggingface.co/datasets/MarrLab/DinoBloom_hemato_embeddings
BibTex
@InProceedings{DasMuh_Genetically_MICCAI2026,
author = { Dasdelen, Muhammed Furkan AND Ozlugedik, Fatih AND Looser, Ilaria AND Umer, Rao Muhammad AND Pohlkamp, Christian AND Marr, Carsten},
title = { { Genetically Aligned Patient Representations Improve Hematological Diagnosis } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 16885},
month = {September},
page = {pending}
}
Reviews
Review #1
- Please describe the contribution of the paper
This paper proposes a hematology-specific patient representation framework that aligns peripheral blood smear morphology with cytogenetic and molecular genetic information. The method uses a two-stage training strategy: first, patient-level image pretraining based on a transformer aggregator over single-cell embeddings, and second, supervised multimodal alignment of slide, karyotype, and mutation representations in a shared embedding space. The authors show improvements on downstream hematological classification tasks.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
—The paper addresses a clinically relevant and underexplored problem in computational hematology by aligning blood smear morphology with cytogenetic and molecular information. The proposed multimodal setting is well motivated. —- overall good quality of work —- The manuscript is well prepared with good writing except a few places in the Methods
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The main weakness is the lack of clarity and organization in the Methods section. —In Section 2.4, EMA is introduced without explanation. Some symbols are introduced only partially or without explicit description at first use. —. The notation is also confusing because the symbol lambda is reused for different loss weights in Sections 2.4 and 2.5.. While not mathematically wrong, this reuse reduces clarity and should be avoided. —The manuscript does not clearly define how the BCE loss is formulated, whether it is applied separately for mutation and karyotype reconstruction, or how it is averaged across samples and dimensions.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
overall good quality of work, interesting and innovative idea.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #2
- Please describe the contribution of the paper
The authors propose a multimodal learning framework (GenBloom) that aligns single white blood cell morphology with underlying genetic alterations, including chromosomal aberrations and somatic mutations. The method leverages self-supervised vision encoders (DINOv2/iBOT) with a modified Vision Transformer architecture to aggregate single-cell embeddings into patient-level representations, which are then aligned with cytogenetic and molecular features. Using a large-scale cohort with matched imaging and genomic data, the study systematically evaluates the proposed framework against domain-specific pathology foundation models under different initialization strategies. The model is further assessed in bidirectional tasks, including predicting genetic alterations from histology and linking slide-level features to mutation profiles, demonstrating the potential to couple morphology with genomic characteristics.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
1.This study leverages a relatively large cohort in the field with rich multimodal data, including single-cell images and matched cytogenetic and molecular profiles, and includes an independent test set. The evaluation across multiple publicly available tasks further supports the robustness of the framework. 2.The proposed method demonstrates improved and more balanced performance across datasets compared to baseline foundation model approaches. The experimental design is reasonably comprehensive, including multiple setups and ablations, and highlights the advantage of the method even with comparatively less data used in the self-supervised pretraining stage. 3.The paper presents clear figures and well-structured experimental settings, which help improve the readability of the methodology and make the overall workflow and evaluation easier to follow.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.The performance on out-of-domain tasks appears limited, and the lack of strong or well-matched baselines in this setting makes it difficult to fully assess the generalizability of the proposed framework. 2.The formulation of the patient-level embedding is not fully justified. Although the task is inherently multimodal, the representation appears to be primarily derived from single-cell image features, raising questions about how effectively multimodal information is integrated. 3.The comparison with existing methods may not be entirely fair. In particular, benchmarking against cell-level foundation models would provide a more appropriate and convincing evaluation than comparisons with models trained on lower-magnification (e.g., 20×) histopathology data.
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
While the paper presents an interesting multimodal framework and is supported by a relatively strong experimental setup and dataset, there are several concerns that limit its overall impact. In particular, the generalizability of the method is not sufficiently demonstrated, as performance on out-of-domain tasks remains limited and lacks strong baseline comparisons. Additionally, the design choice of constructing patient-level embeddings primarily from single-cell image features is not fully justified given the multimodal nature of the problem. The fairness of comparisons is also a concern, as more relevant cell-level or domain-specific foundation models are not included.
- Reviewer confidence
Somewhat confident (2)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #3
- Please describe the contribution of the paper
GenBloom is a hematology-specific slide-level encoder that aligns single white blood cell images with cytogenetic (karyotype) and molecular genetic (somatic mutation) data. Training is two-stage: (1) self-supervised DINOv2/iBOT pretraining of a ViT aggregator on ~794K single-cell images from 1,634 patients, (2) supervised contrastive alignment of slide embeddings with karyotype and mutation profiles on 146 AML patients. The model outperforms larger histopathology foundation models (GigaPath, PRISM, TITAN) on hematological classification and retrieval tasks, and enables cross-modal retrieval between morphology and genetic profiles.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
S1.Novel and clinically well-motivated problem. Hematology is underserved in computational pathology, and the tight coupling between blood smear morphology and genetic alterations in leukemia is a natural fit for multimodal alignment.
S2.Strong methodology. The two-stage design (self-supervised pretraining then supervised contrastive alignment) is principled. The reconstruction loss (L_BCE) to prevent representational collapse is a thoughtful design choice, validated in the ablation (Table 3: removing it drops S->K MRR by 36%).
S3.Outperforms much larger models with less data. GenBloom’s small ViT (~6 layers) outperforms GigaPath (86.3M params, 171K WSIs), PRISM (99M, 587K WSIs), and TITAN (42.1M, 335K WSIs) on hematology tasks despite training on orders of magnitude less data. This highlights the value of domain-specific pretraining.
S4.Proper statistical analysis. Wilcoxon signed-rank tests with Bonferroni correction, 1,000 bootstrap iterations for retrieval, Friedman test for overall ranking.
S5.Cross-modal retrieval is a genuinely useful capability. Retrieving genetic profiles from morphology (and vice versa) has direct clinical utility for triage and prioritization of confirmatory genetic testing.
S6.Good reproducibility commitment. Code and model weights promised on GitHub. Carbon footprint reported (~7.26 kg CO2eq). Three evaluation datasets are publicly available.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
W1.Very small sample sizes throughout. Genetic alignment uses only 146 training patients. Test sets are tiny: AML-Hehr subtypes have 6-9 test patients per class, APL-AML has 12 test patients, AMH has 20.With such small n, even the reported statistical tests have limited power, and results may be unstable.
W2.Cross-modal retrieval performance is modest in absolute terms. Top-1 accuracy for slide-to-karyotype is 0.09-0.14 (Table 1). While significantly above random, these numbers are far from clinically actionable. The authors should discuss what level of retrieval performance would be needed for practical utility.
W3.Genetic alignment is limited to AML only. It is unclear whether the alignment approach generalizes to other hematological malignancies (MDS, MPN, lymphoma) that are present in the pretraining set. The out-of-domain evaluation (APL-AML, AMH) tests generalization of the visual encoder, not of the genetic alignment.
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
1.Report confidence intervals for classification balanced accuracy, not just retrieval metrics. The small test sets make point estimates unreliable. 2.Discuss what top-1 retrieval accuracy of 0.09-0.14 means practically. At what threshold would cross-modal retrieval become clinically useful? 3.Consider evaluating genetic alignment on at least one non-AML hematological malignancy to test generalizability of the alignment approach.
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(5) Accept — should be accepted, independent of rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
This paper addresses an underexplored and clinically important problem: aligning hematological morphology with genetic data. The methodology is sound, the statistical analysis is rigorous, and the results demonstrate clear value of domain-specific pretraining over much larger general-purpose models. The cross-modal retrieval capability is novel and clinically relevant. The main concerns are the small sample sizes (particularly for genetic alignment and test sets), modest absolute retrieval performance, and limitation to AML only. Despite these, the novelty, clinical relevance, and methodological rigor place this marginally above the acceptance threshold in my opinion.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Author Feedback
We thank the reviewers for their constructive feedback and recognition of our contributions, including the clinical relevance of genetic alignment in hematology (R1, R3), use of large scale single cell dataset (R2) and baseline comparisons with proper statistical analysis (R2, R3). Below we discuss additions to the camera ready to address the reviewer comments.
R1: Notations – We now expand the explanations of some of the terms that were previously not explained sufficiently, including EMA and L_BCE, and use separate lambdas in the formulas to increase clarity. EMA refers to the Exponential Moving Average of weights. L_BCE is the binary cross-entropy loss between each real and predicted genetic mutation, including both karyotype and molecular genetics, and is used to preserve the biological information of embeddings while aligning them with the patient representations.
R2, R3: Out-of-domain evaluation is missing, test datasets are small – Hematopathology suffers from a lack of large and extensive benchmarks, especially at the patient level. We tried to include all publicly available test datasets in our evaluations. Across all datasets, we ensured that the test sets were held out and never seen by the model. In the APL-AML dataset, the staining also differs significantly from that of the other datasets.
R2: Baseline evaluations may not be fair – We acknowledge that slide-level models (e.g., TITAN) were mainly trained on histopathology images. However, to the best of our knowledge, no hematology slide-level encoder is currently available to include in the evaluations. We trained our slide-level model using DinoBloom (Koch et. al., 2024), which is the only available single-cell foundation model. As a fair comparison, we added simple mean pooling of DinoBloom embeddings (see Figure 2a).
R3: Discussing retrieval capabilities – Although retrieval between image and genomic embeddings is promising, its clinical utility requires extensive evaluation and validation across different diseases. In this work, we present, to the best of our knowledge, the first proof-of-concept cross-retrieval framework in hematopathology, paving the way for future genetics-aware smear analysis.
R3: Cross-retrieval generalization to other diseases – Our genetic alignment cohort consisted of AML patients. To test whether genetic-to-image retrieval can be used in other diseases, we included myeloproliferative neoplasm (MPN) patients in the gene retrieval task for JAK2 (186 were diagnosed with MPN and 10 with MDS/MPN out of 213 patients, Table 2). We will highlight this point in the text more clearly.
Meta-Review
Meta-review #1
- Your recommendation
Provisional Accept
- Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.
The reviews are consistently positive, with two weak-accept and one accept recommendation. The paper addresses a clinically relevant and underexplored problem in computational hematology by aligning blood smear morphology with cytogenetic and molecular genetic information. The reviewers recognize several strengths, including the hematology-specific patient representation framework, two-stage pretraining and genetic alignment strategy, use of a large single-cell image cohort, comparisons with pathology foundation models, statistical analysis, and the clinically meaningful cross-modal retrieval setting.
The main concerns are the relatively small sample size for genetic alignment and test sets, limited absolute cross-modal retrieval performance, incomplete demonstration of generalization beyond AML, and the need for clearer methodological details and more directly comparable cell-level baselines. These issues should be addressed or discussed in revision.
