List of Papers Browse by Subject Areas Author List
Abstract
Deep learning denoisers achieve remarkable restoration quality, yet their high computational cost limits deployment in resource-constrained settings. While post-training quantization (PTQ) enables efficient deployment, conventional PTQ methods degrade fidelity under ultra-low precision, causing frequency-selective degradation in reconstructed outputs. Observing that quantization errors concentrate in high-frequency subbands, we propose a dual-stage PTQ framework based on stationary wavelet transform (SWT) that explicitly penalizes subband-wise quantization distortions. Extensive experiments demonstrate consistent improvements over existing methods across diverse architectures. Our method effectively balances fidelity with substantial reductions in model size, supporting deployment-oriented medical image denoising. The code is available at https://github.com/unikohoho/Quantization.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/5014_paper.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to the Code Repository
https://github.com/unikohoho/Quantization
Link to the Dataset(s)
N/A
BibTex
@InProceedings{KoUni_FrequencyAware_MICCAI2026,
author = { Ko, Uni AND Lee, Yumi AND Park, Juneyoung AND Choi, Jang-Hwan},
title = { { Frequency-Aware Post-Training Quantization for Medical Image Denoising } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 16890},
month = {September},
page = {pending}
}
Reviews
Review #1
- Please describe the contribution of the paper
This paper presents a frequency-aware post-training quantization framework for medical image denoising. The key idea is that conventional spatial-domain PTQ objectives under ultra-low precision under-penalize high-frequency distortion, which is especially harmful for diagnostically important structures. The authors support this claim with stationary wavelet transform (SWT) analysis showing that the HH subband suffers the largest quantization distortion, and then propose a dual-stage PTQ pipeline consisting of tensor-level coarse calibration and wavelet-guided output-level refinement. The method is designed to be architecture-agnostic and is evaluated across CNN, transformer, and hybrid denoisers on CT and fluoroscopy datasets. The experimental results suggest substantial fidelity gains over several PTQ baselines, especially at 2-bit and 4-bit precision, while still achieving practical compression for deployment.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The paper identifies a concrete and well-motivated failure mode of existing PTQ for medical denoising, namely the mismatch between spatial-domain calibration and frequency-selective degradation under ultra-low precision. The method is technically coherent: the SWT-based diagnosis directly motivates the wavelet-guided refinement objective, rather than appearing as an unrelated add-on. The evaluation is reasonably broad for this problem setting, covering two medical imaging scenarios and three representative denoising backbones, with results at 2/4/8 bits. The empirical gains are strong, including up to 5.54 dB PSNR improvement over baselines and up to 8× model compression, and the ablation study shows that the second-stage wavelet refinement contributes substantial additional gains. The paper also considers deployment, including ONNX Runtime export and an anonymized project link, which is a practical strength.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The novelty is meaningful but somewhat incremental. The paper builds on established PTQ baselines such as BRECQ, PTQ4SR, and 2DQuant, while also drawing on wavelet/subband ideas that have already appeared in low-dose CT restoration literature. In that sense, the main novelty lies more in the task-specific integration of frequency-aware refinement into PTQ for medical denoising than in a fundamentally new quantization principle.
Several important design choices are empirical, especially the band weights used in the wavelet-guided loss. The paper would be stronger with a sensitivity analysis for these weights and for the calibration-set size, since both may affect the stability and practical applicability of the method.
The deployment claim is not fully supported quantitatively. The text states that there is no additional latency overhead compared with FP32, but the deployment table reports model size and compression ratio rather than actual runtime, throughput, or memory measurements. This makes it difficult to assess the real deployment benefit beyond storage reduction.
The evaluation focuses mainly on image fidelity metrics such as PSNR/SSIM and qualitative examples. Given the medical setting, it would be more convincing to include a stronger task-oriented or reader-oriented analysis showing that diagnostically relevant structures are indeed better preserved after quantization.
The paper would benefit from a clearer comparison or discussion against stronger deployment-oriented alternatives, such as quantization-aware training or mixed-precision strategies. This would help clarify when the proposed PTQ approach is preferable in practice.
The quality of Fig. 1 should be improved. In the right panel, the stacked image examples contain visible artifacts that look like residual watermark-like patterns or incompletely edited/generated image content. This raises concerns about figure preparation and visual reliability. The authors should replace or clean these images and ensure that all visual materials are properly anonymized, traceable, and publication-ready.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(2) Reject — should be rejected, independent of rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
After reading the paper, my overall impression of the technical idea is positive. The proposed frequency-aware PTQ framework is practically motivated, and the use of wavelet/subband information for preserving high-frequency structures in medical denoising is reasonable. However, I have serious concerns about the reliability and preparation quality of the manuscript, which ultimately prevents me from recommending acceptance.
First, Fig. 1 appears insufficiently reliable. In the right panel, the stacked image examples contain visible residual artifacts and unnatural visual patterns, which look like incompletely edited watermark-like content or AI-assisted/generated image artifacts. This raises concerns about the preparation and traceability of the visual materials.
Second, the bibliography contains multiple obvious citation errors that resemble hallucination-like or automatically generated reference mistakes. For example, the citation of “Strategies for reducing radiation dose in CT” has an incorrect author list. The CTformer reference, “Convolution-free Token2Token Dilated Vision Transformer for Low-dose CT Denoising,” appears to have incorrect year/volume information and an incomplete author list. The Restormer reference includes an author not listed in the official CVPR paper. SwinIR is cited as a CVPR paper, whereas it was published in ICCV Workshops. Ref. 22 also appears to mix the title, SPIE volume, and page range from different mammographic-structure references. These are not minor formatting issues; they indicate substantial problems in citation accuracy and reduce confidence in the care taken in preparing the manuscript.
In addition, the related work is not well positioned for a MICCAI audience. Despite being submitted to MICCAI and addressing medical image denoising, the paper does not appear to cite any MICCAI work, nor does it sufficiently connect the contribution to the broader MICCAI literature on medical image restoration, efficient deployment, model compression, or clinically oriented evaluation.
Therefore, although the method itself is interesting and the experimental results appear promising, the figure artifacts, multiple hallucination-like reference errors, and weak positioning within the MICCAI literature make me unable to recommend this paper for acceptance in its current form.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #2
- Please describe the contribution of the paper
The core contribution of this paper lies in introducing the Stationary Wavelet Transform into the optimization process of post-training quantization. Experimental results show that this approach significantly improves both subjective and objective performance.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The core contribution of this paper lies in performing optimization for Post-Training Quantization in the frequency domain, rather than directly in the image domain, by leveraging wavelet transform.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The limitations of this paper are as follows:
1.The core contribution appears to lie in transforming the input image during the optimization stage and performing loss computation in the transformed domain. This modification seems relatively simple, with limited implementation difficulty and workload. 2.The relationship between Stage 1 and Stage 2 is not clearly explained. In particular, the paper does not provide sufficient analysis on why the Tensor-Level Coarse Calibration can offer a stable initialization. 3.The authors claim that quantization in the spatial domain leads to the loss of high-frequency information. However, from the qualitative results, the compared methods seem to mainly suffer from reduced denoising capability rather than poor edge preservation, while the proposed method introduces some degree of edge blurring. The authors should provide further explanation for this observation.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission has provided an anonymized link to the source code, dataset, or any other dependencies.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The experimental results of this paper are strong. However, the overall workload appears limited, and there are some issues in the methodological description. Therefore, my recommendation is Weak Accept.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #3
- Please describe the contribution of the paper
The authors propose a frequency-aware PTQ framework based on stationary wavelet transform (SWT). The method adopts a two-stage pipeline: (i) standard tensor-level calibration, followed by (ii) wavelet-guided output refinement that introduces subband-wise error weighting, with increased emphasis on high-frequency components. The approach is evaluated across multiple architectures (CNN, Transformer, hybrid) and medical datasets (CT and fluoroscopy), demonstrating consistent improvements over strong PTQ baselines, particularly in ultra-low-bitwidth regimes.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The proposed solution is simple yet meaningful: incorporating frequency-domain awareness via wavelet decomposition and reweighting reconstruction errors. The two-stage design is logical and integrates naturally with existing PTQ pipelines. The proposed method achieves substantial improvements over prior work, including gains of up to ~5.5 dB PSNR in fluoroscopy and consistent advantages across architectures. Notably, the approach remains robust in extremely low-bit settings (e.g., 2-bit), where most PTQ methods degrade significantly. This is a strong empirical validation of the core idea.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The frequency weighting scheme (e.g., λ coefficients for subbands) appears empirically chosen. The paper would benefit from either a principled derivation or a more thorough sensitivity analysis to demonstrate robustness to these hyperparameters. Experiments are confined to medical image denoising. While the motivation is domain-specific, it remains unclear whether the proposed method generalizes to other restoration tasks (e.g., super-resolution, deblurring) or even to non-medical image domains.
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission has provided an anonymized link to the source code, dataset, or any other dependencies.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
PTQ and frequency awareness are widely adopted things. However, the paper presented sufficient comparison with the two good methods.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Author Feedback
[R1: references/visual reliability] We verified the bibliography against original publication records and will correct the metadata errors flagged by R1: Strategies for reducing radiation dose in CT author list; CTformer year/volume/author list; Restormer author list; SwinIR venue (ICCV Workshop, not CVPR); and Ref.22 title/SPIE volume/page mismatch. We will also fix capitalization, missing pages, and inaccurate entries across the reference list. These are metadata errors, not unsupported citations. Fig.1 was manually prepared from actual saved experimental outputs, not AI-generated, watermarked, or synthetic content.
[R1/meta: deployment/clinical claims] Table 2 reports only model size and compression; it does not support latency, throughput, or memory claims. We will remove “without incurring additional latency overhead compared to FP32” and limit the deployment claim to ONNX-exported model-size compression: 3.88–4.00x for INT8 and 7.47–8.00x for INT4.We will also remove reader-level or diagnostic claims not supported by PSNR/SSIM. “Diagnostically critical structures” will be revised to reconstruction-fidelity/frequency-distortion wording.
[R1/R3: empirical weights/calibration size] The subband weights lambda_LL=0.2, lambda_HF=0.3, lambda_HH=0.5 were fixed across datasets/backbones to reflect the HH-dominant distortion in Fig.2; they were not tuned per test set and are not claimed to be optimal. We will add a compact robustness analysis for representative weights and calibration-set sizes. The claim will be robustness within reasonable settings, not a principled optimum.
[R2: Stage 1/Stage 2 relation] Stage 1 and Stage 2 optimize the same clipping-bound parameters Φ. Stage 1 provides a stable initialization under ultra-low-bit quantization by reconstructing quantized weights and activations to better align with the FP32 distributions before output-level refinement. Starting from this initialization Φ(0), Stage 2 refines Φ using the SWT-based output loss to correct the remaining frequency-selective distortion. Thus, Stage 1 stabilizes local quantization, while Stage 2 focuses on reducing output-level subband distortion. This design is supported by the ablation study, where Stage1+2 consistently outperforms Stage1 alone across all backbones and datasets, and Fig.5 shows reduced standardized HH error.
[R2: high-frequency/edge wording] High-frequency degradation should not be interpreted solely as edge disappearance. In our setting, quantization-induced high-frequency distortion may appear as residual noise amplification, texture corruption, or attenuation of fine structures depending on the backbone architecture and bit-width. We will replace “preserves fine structural details with comparatively reduced distortion” with wording that more precisely describes reduced high-frequency quantization distortion. This also addresses the reviewer’s observation that the visual degradation may reflect reduced denoising capability rather than only edge loss.
[R1/R2/R3: novelty and alternatives] The contribution is not SWT itself, but identifying frequency-selective failure of spatial PTQ for medical denoising and turning it into a two-stage, architecture-agnostic PTQ objective. QAT and mixed precision target different regimes. QAT requires retraining and usually full training data; our setting freezes the pretrained denoiser and uses only calibration samples. Mixed precision can improve fidelity but introduces layer-wise bit allocation and hardware/policy complexity; we intentionally test the stricter uniform W/A low-bit setting to isolate frequency-aware PTQ.
[Scope/MICCAI positioning] We will position the method as medical image denoising PTQ, not general restoration or clinical validation. We will add MICCAI-relevant related work on low-dose imaging, medical restoration, efficient deployment, and state extension to other restoration tasks and reader/task-level evaluation as future work.
Meta-Review
Meta-review #1
- Your recommendation
Provisional Accept
- Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.
This paper presents a technically sound and empirically well-supported methodological contribution on post-training quantization for medical image denoising. The strengths of the submission are clear. The paper identifies a plausible failure mode of standard spatial-domain PTQ under ultra-low precision, namely disproportionate degradation of high-frequency components, and translates that observation into a coherent two-stage frequency-aware refinement strategy. The experimental validation is broad for the claimed scope, covering two medical denoising settings, three backbone families, multiple PTQ baselines, and 2/4/8-bit quantization. The reported gains at 2-bit and 4-bit are substantial, and the ablation provides evidence that the second stage contributes meaningfully beyond the initial tensor-level calibration. The main concerns are more limited in scope. The band-weighting scheme is empirically chosen without a robustness study, the deployment discussion overstates what is actually shown because latency is claimed without runtime measurements, and the evaluation remains focused on reconstruction fidelity rather than task- or reader-level medical assessment. In addition, the bibliography contains several metadata inaccuracies that should be corrected, and the MICCAI-specific positioning could be stronger. After checking the manuscript, however, some reviewer concerns carry limited weight: the comment that the relationship between Stage 1 and Stage 2 is unclear is only partially supported because the manuscript explicitly defines Stage 1 as the initialization for Stage 2 and reports their respective contributions in the ablation, while the concern about suspicious artifacts in Fig. 1 is not substantiated by the submitted figure. Accordingly, I assign limited weight to those criticisms in the overall assessment.
