Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness
Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
- Published
- Sep 23, 2026 — 13:37 UTC
Problem
Federated imputation methods have predominantly focused on value-level missingness, with limited evaluation of feature-level missingness. This paper addresses this gap by proposing a novel approach that enables effective imputation in federated settings, particularly when features are missing across different clients. The work is presented as a preprint and has not undergone peer review.
Method
The authors introduce Fed-ReMasker, an adaptation of the ReMasker masked autoencoder tailored for federated learning environments. This architecture allows multiple centers to collaboratively impute features that may not have been observed locally, leveraging shared knowledge across the participating centers. The method is evaluated on both synthetic datasets, which include linear and nonlinear relationships, and real-world tabular datasets, specifically clinical data. The training compute requirements are not specified in the paper.
Results
The results indicate that Fed-ReMasker achieves an imputation error of 93.2% at the value level and 96.7% at the feature level, although the paper does not specify the baselines for comparison. The method demonstrates robustness to client heterogeneity, outperforming all baselines across 36 value-level scenarios. In feature-level scenarios, it outperformed each baseline in at least 35 out of 36 scenarios. Additionally, when compared to a centralized model trained on pooled data, Fed-ReMasker is within 3.0% on average, indicating its effectiveness in a federated context.
Limitations
The authors do not report any limitations in their work. However, the lack of specified baselines for the reported imputation errors may hinder the ability to fully assess the performance of Fed-ReMasker against existing methods.
Why it matters
The implications of this work are significant for downstream applications in federated learning, particularly in domains where data privacy is paramount, such as healthcare. By enabling effective feature-level imputation, Fed-ReMasker can enhance the utility of federated datasets, allowing for more robust machine learning models that can operate effectively even in the presence of missing data.
By Callan Zhang · Sep 23, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
