Abstract:
Recent advances in remote sensing have expanded Earth observation through an increasing number of sensing modalities, where each sensor provides distinct spectral, spatial, and temporal information. Effective analysis of these data depends on learning representations that integrate complementary information across modalities while supporting multiple learning objectives. Multi-task multi-modal learning in remote sensing remains challenging due to mismatched resolutions, variable data quality, and limited or heterogeneous annotation. Existing approaches typically enforce a single shared feature representation across all tasks and modalities without explicitly encoding the full set of available annotations. This design assumes uniform task compatibility and often yields latent spaces that fail to reflect task relationships or preserve meaningful structure derived from supervision. This dissertation introduces an annotation-driven continuous latent space framework for multi-task multi-modal representation learning. The central premise is that all information known about a sample should be reflected in the geometric position of its embedding, such that distances correspond to task-relevant similarity across modalities. A continuous triplet loss formulation is developed to construct this latent space using mixed discrete and continuous annotations. Extensive experiments on multiple remote sensing datasets demonstrate that the proposed approach produces more informative and better aligned representations than state-of-the-art multi-task multi-modal learning methods.
Links:

Citation:
M. Zhou, "Continuous Latent Spaces for Multi-Task Multi-Modal Learning," Ph.D Thesis, Gainesville, FL, 2026.
@book{zhou2026continuous,
title={Continuous Latent Spaces for Multi-Task Multi-Modal Learning},
author={Zhou, Meilun},
year={2026},
publisher={University of Florida}
}