- What happened: NASA and IBM have released a lunar foundation model, developed using 17 years of observations from the Lunar Reconnaissance Orbiter (LRO) and other data, along with datasets and code for additional training and evaluation.
- Why it matters: The model handles data from different instruments and resolutions in a single framework, and suggests it may be useful for tasks such as crater detection even with a small amount of labeled data.
- What to watch: The ice evaluation measures how well the model reproduces an existing prospectivity map. Field measurements and independent validation in other regions will be needed to determine how useful it is for actual exploration.
On September 10, 2026, NASA and IBM released the "NASA-IBM Lunar Foundation Model," which uses 17 years of observations from the Lunar Reconnaissance Orbiter (LRO) as its main training data. It is a lunar foundation model trained on imagery, terrain, mineral composition, imaging conditions and more, and it can be adapted through additional training for tasks such as detecting craters and estimating areas where ice is likely to exist. The effort aims to make it easier to use data from different spacecraft and instruments together.
However, the figures presented in the announcement, "about 22% improvement in ice estimation and about 19% in crater detection," are based on different evaluation metrics. In particular, the ground truth used in the ice evaluation is not the amount of ice measured on site, but an existing prospectivity map created from expert knowledge.
Aligning roughly 2 million sets of observations
Having a large volume of lunar imagery does not mean it can be used as-is to train AI. Even for the same location, instruments, pixel sizes and sun positions differ. To combine imagery with terrain, mineral composition and other data, the position of each observation on the Moon must be precisely aligned.
According to IBM's official announcement, the data foundation includes more than 30 layers of spatial data collected from nine instruments across four missions. In addition to LRO, it draws on data from GRAIL, which observed the gravity field, Lunar Prospector, and Japan's lunar orbiter Kaguya (SELENE). It integrates mineral composition obtained from Kaguya's spectral observations and terrain data from LRO's laser altimeter, LOLA. This allows information that cannot be obtained from camera images alone to be treated as observations of the same location.
"SomBench," built by the research team, is a training and evaluation dataset made by cutting aligned observation data into small tiles and pairing corresponding data. According to the paper, it contains about 964,000 sets derived from the Wide Angle Camera (WAC) and about 1 million derived from the Narrow Angle Camera (NAC). The "roughly 2 million sets" refers to the number of cropped and combined data sets, not 2 million original images.
WAC images have a resolution of about 100 m per pixel and NAC images about 1 m per pixel, so the spatial scales they capture differ greatly. Moreover, the NAC data is limited to locations where a corresponding high-resolution terrain model exists. It is not a dataset covering the entire Moon at 1 m resolution.
For pretraining, the model uses 11 types of information: nine inputs treated as images plus imaging conditions and other data. Not all 30 original layers are loaded as images at the same resolution; some auxiliary data is converted into values such as per-tile averages before input. Aligning coarse-resolution observations with high-resolution images does not make the original observations themselves higher in resolution.
The way training and evaluation data were split also shows some care. If tiles were assigned randomly, images of the same location taken at different times could appear in both the training and evaluation sets, meaning the model could effectively see familiar terrain again during evaluation. The team therefore split data by region so that observations of the same place do not span training and evaluation. Geographic alignment is used not only to link different observations but also to avoid overstating performance.
Learning how the sun's angle changes the Moon's appearance
The model's basic architecture adopts the design of TerraMind, a foundation model for Earth observation. However, it was not created by fine-tuning an Earth-trained model on lunar data; it was pretrained from scratch using lunar observations.
During training, part of the input is hidden, and the model is asked to predict the hidden portions from the remaining information. By converting imagery, terrain, observation conditions and other data into small units by type, the model can learn relationships among different observations without people labeling craters or geology on every image. It is then trained further with labeled data for each application.
On the Moon, the angle of the sun changes how shadows appear on the same terrain. It is difficult to tell from imagery alone whether a dark area reflects a difference in material or simply the way light falls. This model takes the angles of the sun and camera as known conditions, including the factors that change appearance as part of its learning material.
The handling of wide-angle and narrow-angle images is another feature. Rather than overlaying them on the same tile as input, the model uses data groups that each preserve their own resolution in the same training process, updating shared model parameters. The design lets one model handle observation scales that differ by roughly 100 times. For additional training, it also uses FlexiViT, a technique that allows the unit into which images are divided to be changed flexibly.
These mechanisms are meant to address differences in observation conditions and uses. However, comparative experiments isolating how much feeding in imaging angles and training on mixed resolutions each contributed to performance gains are listed in the paper as future work.
What the ice "prospectivity map" actually represents
NASA's announcement also envisions using the model to investigate lunar volatiles and resources. Narrowing down where ice is likely to exist could help in planning future exploration. But it is important to be careful about what the current evaluation actually measures.
In the ice task, the model estimates a "prospectivity" value from 0 to 1 from 240 m-per-pixel data covering the areas near both lunar poles. It uses eight types of input: slope, slope direction, the depth at which ice can stably exist, maximum temperature, permanently shadowed regions, shadow density, distance to shadow, and terrain curvature.
The prospectivity map used as ground truth here was also created by combining the same eight types of information based on expert knowledge. It uses "fuzzy logic," which evaluates conditions such as low temperature and shade in stages and integrates them. The test is therefore not an external validation that identifies ice actually measured at unknown locations, but a measure of how well the model can reproduce a map built from existing rules.
This approach is also used in fields such as mineral exploration on Earth. In July 2025, the U.S. Geological Survey (USGS) released a map narrowing down where water is likely to exist at Mons Mouton near the lunar south pole, based on conditions such as shade and slope. It is a method for deciding where to investigate first when on-site surveys are insufficient.
The description of the ice dataset also states that the weighting used in the reference map includes assumptions and may change as more on-site observations become available. High values do not represent the amount or purity of ice, and cannot simply be read as the "probability that ice will be found."
Even if AI reproduces the reference map with high accuracy, that does not prove that the assumptions used to create the map are correct.
The model's value lies in allowing the work of narrowing down candidate sites from many types of observation data to be handled through a common framework. Determining whether ice actually exists at those sites, and whether it is in a usable amount or state, will require other means of observation and on-site measurement.
What the "about 22%" and "about 19% improvement" mean
The basis for the performance evaluation is the non-peer-reviewed paper by Paolo Fraccaro and colleagues, "Multimodal-Multiresolution Foundation Model for Lunar Remote Sensing". The published figures come from the research team's own benchmarks, and the table below shows averages over five runs with different random conditions. It does not mean five independent on-site observations were made.
| Task and metric | Lunar foundation model | Comparison model | How to read the result |
|---|---|---|---|
| Ice prospectivity: RMSE (lower is better) | 0.0293 | SwinV2-B: 0.0377 | Error against the reference map is about 22.3% smaller |
| Wide-angle craters: mAP (higher is better; both use 50% of training data) | 0.2541 | SwinV2-B: 0.2313 | About 9.9% higher |
| Narrow-angle craters: mAP (higher is better) | 0.1543 | SwinV2-B: 0.1552 | Roughly equivalent |
| Irregular volcanic terrain: IoU (higher is better) | 0.5709 | ConvNeXtV2-B: 0.5687 | Roughly equivalent given run-to-run variation |
For the lunar foundation model, the team tried several approaches for each task, including training the whole model further, training only some parameters, and freezing the feature-extraction part, and adopted the best-performing setting. In the comparison with other models, the data splits within each task, loss functions and output-head structures were matched, while optimization conditions such as learning rates were set to suit each model group. It is not a comparison in which all conditions were completely identical.
The "about 22%" reported for ice corresponds to the root mean square error (RMSE) falling from 0.0377 to 0.0293: (0.0377 − 0.0293) ÷ 0.0377 is about 22.3%. For mean absolute error (MAE), the drop from 0.0256 to 0.0197 is about 23.0%.
Both are the proportion by which prediction error decreased, and do not mean that "the probability of discovering ice rose by 22 points."
The "about 19% improvement" announced for wide-angle crater detection is a comparison using AP@75; the improvement in mAP when using the same amount of training data is about 9.9%.
In Table 4 of the paper, comparing conditions in which both models use 50% of the training data, AP@75 rises from 0.1862 to 0.2213. 0.2213 ÷ 0.1862 − 1 is about 18.9%, corresponding to the "about 19%" in the announcement. mAP, meanwhile, rises from 0.2313 to 0.2541, an improvement of about 9.9%.
AP@75 is a metric that sets a strict standard for how closely a detected region must overlap the ground-truth region, while mAP averages results across multiple overlap thresholds. Even for the same detection results, the improvement rate differs depending on the metric used.
Nor should the explanation of "a 19% improvement with half the data" be read as meaning the lunar foundation model used half the data while the comparison model used all of it.
That said, the results do suggest that a small amount of labeled data may be used effectively. In wide-angle crater detection, the lunar foundation model's mAP of 0.2541 with 50% of the training data exceeded the 0.2420 SwinV2-B obtained with 100% of the training data. For this task at least, the result suggests the model may reduce the labeling work done by people.
On the other hand, in narrow-angle crater detection and volcanic terrain classification, the model did not clearly outperform strong comparison models. The roughly 3% improvement in volcanic terrain that IBM cites is also a figure from the comparison with SwinV2-B, and should be considered separately from the difference with the best-performing comparison model shown in the table.
Using the released model in real research
Using the released model weights and code, researchers can try additional training tailored to their own research questions. In crater detection, even LoRA, which freezes most of the model and trains only a small number of added parameters, showed competitive performance.
However, simply freezing the feature-extraction part completely reduced crater-detection performance. Adjustment suited to the application is still necessary.
In the ice task, even a model with the same architecture but no pretraining outperformed many of the comparison models. The model structure itself, which processes each input type differently, may be contributing to performance, so not all of the improvement can be attributed to pretraining on 17 years of lunar data.
The breadth of the evaluation also has limits. The ice test covers 25 tiles and the volcanic terrain test 10, which is small in scale. The model card also says use in actual mission operations, such as certifying landing sites, has not been validated.
For generated terrain, there are reports of cases where the shape is similar but absolute elevation values are off, and where generated latitude and longitude deviate from the actual location by several tens of degrees. Being able to learn relationships among observation data is a separate matter from whether the output can be used directly as survey values.
The model weights and the code for additional training and inference are released under the Apache-2.0 license. Alongside configuration files, training datasets and evaluation tasks were also released. However, the GitHub description states explicitly that pretraining code is not included.
As a result, while research and additional training using the released weights can begin, the model's pretraining process itself cannot be fully reproduced under the same conditions.
Going forward, if independent studies can confirm that performance holds in different regions and lighting conditions, and ice candidate sites can be checked against new observations, the advantages of a shared foundation model will become clearer. For the ice evaluation in particular, an important condition for moving closer to real exploration decisions is not only improving the model's predictive accuracy but also updating the assumptions of the reference map itself through on-site measurements.
