NASA and IBM Research have released the NASA-IBM Lunar Foundation Model, described as one of the first open-source foundation models for lunar science. It was trained on nearly 2 million tile bundles drawn from 17 years of Lunar Reconnaissance Orbiter observations and other missions. On polar ice prediction it cut error by up to 22 percent against the SwinV2-B baseline. The release matters because it turns a vast archive of unlabeled Moon data into a reusable base for research with few labeled examples.

NASA and IBM release open lunar model trained on 17 years of orbiter data

How the corpus and model were built

The training corpus, SomBench, contains nearly 2 million tile bundles across 11 modalities and two spatial scales. About 1 million high-resolution images come from the Narrow Angle Camera at roughly 1 meter per pixel, while just under 964,000 multispectral images come from the Wide Angle Camera at 100 meters per pixel. In total it combines more than 30 spatially aligned data layers from nine instruments and four missions. The split between training, validation and test sets follows map zones rather than random tiles to prevent leakage.

The bulk of observations comes from the Lunar Reconnaissance Orbiter, whose data volume exceeds all other NASA planetary missions combined, according to NASA. The collection adds gravity data from GRAIL, hydrogen data from Lunar Prospector and mineralogy data from JAXA Kaguya/SELENE probe. The architecture builds on TerraMind, a multimodal Earth observation model, but researchers trained it from scratch instead of fine-tuning. Each tile receives imaging geometry as explicit context, including illumination angles, sun position and tile extent.

Feeding lighting geometry directly addresses a central problem of lunar imagery, where appearance depends more on sun angle than on surface properties. The model learns jointly from high-resolution and coarse imagery in one training run, capturing fine detail and broad context together. A technique called FlexiViT allows the same trained model to work with different image patch sizes without retraining. Each data layer gets its own processing path, while baseline models treat inputs as stacked channels. That design alone helped even a randomly initialized control beat five of seven baselines on ice prediction.

What this means for applied AI teams

For organizations that build on geospatial or scientific data, the project shows how to make large unlabeled archives usable with limited labeling budgets. The model matched or beat baselines on crater detection at 100-meter and 1-meter scales, polar ice prediction and segmentation of Irregular Mare Patches. Coarse-scale crater detection improved by nearly 19 percent over SwinV2-B while using only half the training data. Small teams gain a public starting point on Hugging Face, GitHub and TerraTorch, while larger groups can adapt the same weights across tasks.

The limits are clear and relevant for procurement and planning. Generation tests showed latitude and longitude errors of dozens of degrees, and elevation shapes could carry shifted absolute heights, so the model does not fit absolute geodetic positioning. On meter-scale craters and Irregular Mare Patches, results roughly tied the strongest baselines within run-to-run variance. Some test datasets are small, and controlled experiments isolating each innovation are still pending. Buyers should ask which downstream task, resolution and labeling volume were validated before assuming transfer.

Lighter adaptation looks practical: LoRA fine-tuning kept pace with full fine-tuning across tasks and did better on crater detection, with full tuning ahead only on the two smallest tasks. The work continues the NASA-IBM AI for Science collaboration under a Space Act Agreement since early 2022, following the Prithvi model from August 2023 for flood and wildfire mapping. A marker to watch is uptake of the released pretraining datasets and benchmarks in new lunar studies. Broader adoption would confirm demand for domain foundation models of this type.