Back to Blog

What a Random Forest Does With Your Map Layers, and Why It Can Mislead You

You have a stack of rasters (gridded map layers): geology, magnetics, radiometrics, stream-sediment geochemistry, distance to faults. You also have a short list of known occurrences. The practical question is how to combine them into a ranked map of where to work next without the result being a black box you cannot defend. After reading this you should be able to explain what a random forest and a neural network do with those layers, set up a basic prospectivity model, and test whether its output means anything.

The principle: learning an association, not a geological process

In plain terms, a supervised model looks at the places where you already know there is mineralisation, and at places where there is not, and finds combinations of layer values that separate the two. It then scores every other cell in your study area by how closely it resembles the known occurrences. The model has no knowledge of fluid flow, structure or metamorphism. It learns only statistical resemblance.

Your geological reasoning therefore goes into the inputs. The mineral-systems approach asks what supplied the metals and fluids, what pathways moved them, and what trapped them. Each layer should act as a proxy for one of those questions. Fault proximity is a proxy for pathways. A magnetic low over altered rock suggests magnetite destruction. Potassium enrichment in radiometrics suggests potassic alteration. A pathfinder element anomaly (an element that travels with the target metal) indicates a dispersion halo.

Random forests

A decision tree asks a series of yes/no questions about the layers, such as "distance to shear zone below 800 m?" followed by "potassium above a given value?". Each question splits the training cells into purer groups, usually measured by Gini impurity, which is low when a group is mostly one class. A single tree overfits, meaning it memorises noise. A random forest grows hundreds of trees. Each tree sees a bootstrap sample (a random draw with replacement) of the training cells, and at each split it may consider only a random subset of the layers. The forest averages the votes into a probability. The randomness decorrelates the trees, so their individual errors tend to cancel.

Neural networks

A neural network (here a multilayer perceptron) multiplies each input layer by a weight, sums the results, and passes them through a non-linear activation function. It repeats this across one or more hidden layers and ends in a probability. Training adjusts the weights by gradient descent to reduce prediction error. Networks can capture smooth, complex interactions, but they need scaled inputs and careful regularisation, and they are harder to interpret. Convolutional networks, which read neighbourhoods of pixels rather than single cells, need far more labelled examples than most licence-scale projects have.

Workflow, step by step

  1. Define the target and the system. Choose one deposit type, such as orogenic gold in a greenstone belt. Mixing deposit types gives the model contradictory examples.
  2. Build a common grid. Resample all layers to one cell size and projection. Choose the cell size to match your coarsest critical layer, not your finest.
  3. Assemble evidence layers. Free or common sources include national geological survey maps and aeromagnetic and radiometric grids, Sentinel-2 and Landsat imagery for alteration indices, and the SRTM or Copernicus DEM for terrain and lineaments. Derive features from them: reduced-to-pole magnetics, first vertical derivative, K/Th ratios, distance-to-fault rasters, and geochemical anomalies levelled by catchment.
  4. Prepare the labels. Convert occurrences to positive cells. Choosing negatives is harder, because the literature notes that datasets usually hold only positive and unlabelled examples, not verified negatives (Airola et al.). Sample background cells away from known deposits and treat them as "unlabelled", not "barren".
  5. Handle imbalance. Positives are typically a tiny fraction of cells. Approaches include balanced bootstrap sampling inside the forest, which improves minority-class prediction error in severely imbalanced data, and oversampling methods such as SMOTE, which are used widely in this field.
  6. Split the data spatially. Hold out whole blocks of ground, not random cells (see the limitations section).
  7. Train and tune. In QGIS, extract raster values to points. Then use Python (scikit-learn) or R (ranger, randomForest) to fit the model. Tune the number of trees, the number of layers tried per split, and minimum leaf size. For a network, tune the hidden layer size and regularisation.
  8. Predict and map. Apply the model to every cell, export the probability raster, and rank it by the share of area covered, not by raw probability.
  9. Interpret. Check which layers drive the result using permutation importance (the drop in performance when one layer is shuffled) or SHAP values (a per-cell attribution of each layer's contribution).

Worked example (illustrative, hypothetical numbers)

Suppose a 30 × 30 km greenstone-belt area gridded at 200 m, giving 22,500 cells. You have 25 known gold occurrences and six layers: distance to major shear, magnetic gradient, K/Th ratio, an iron-oxide index from Sentinel-2, a stream-sediment arsenic anomaly, and lithology class. You take 250 random background cells at least 1 km from any occurrence.

You fit a forest of 500 trees. With random 10-fold cross-validation, the area under the ROC curve (AUC, where 0.5 is chance and 1.0 is perfect separation) comes out at 0.91. That looks excellent. You then divide the area into 5 km blocks and hold out whole blocks in turn. The AUC drops to 0.72. In this made-up case the first figure was inflated because neighbouring cells share almost identical values, so the model was effectively tested on its own training data. The 0.72 is the honest estimate.

Permutation importance then shows distance to shear and arsenic dominating, with the iron-oxide index contributing almost nothing. That matches the geology, so you keep it. If lithology had dominated instead, you would check whether it merely traces where earlier explorers sampled. Finally, you find that the top 10% of the area by probability contains 9 of the 25 occurrences when blocks are held out. That is useful, but modest, and it should be reported as such.

Common mistakes and limitations

  • Spatial leakage. Spatial autocorrelation can result in overoptimistic results, if not corrected for. Use spatial block cross-validation. Recent work applies nested spatial cross-validation with SHAP and permutation importance for this reason.
  • Treating unlabelled cells as barren. Unexplored ground is not evidence of absence. The model may learn "where nobody has looked" instead of "where mineralisation is absent".
  • Exploration bias. Known occurrences cluster near roads, outcrop and past licences. Layers that correlate with access will score well for the wrong reason.
  • Tiny positive sets. With a few dozen positives, a neural network can overfit badly. Studies in this field routinely work with very small labelled sets (scarcity of labeled data and severe class imbalance are common), so stay sceptical of any single accuracy figure.
  • Misleading accuracy. With heavy imbalance, a model that predicts "barren" everywhere scores high overall accuracy. Report AUC from spatial validation, plus how many held-out deposits fall in the top slice of area.
  • Point labels versus real bodies. A deposit is recorded as one point but covers several cells, and local controls may not appear in regional layers.
  • Regional models, local targets. Probability maps rank areas for follow-up. They do not locate drill targets.

Checks to run: compare against a simple baseline such as distance to the nearest fault. Re-run with different random seeds and see whether the top areas stay stable. Remove one layer at a time. Have a geologist review the highest-scoring cells in the field or in the data before trusting them.

Key points to remember

  • Models learn resemblance to known occurrences. Your geological reasoning enters through the choice of layers, labels and deposit type.
  • A random forest averages many decorrelated trees and is a robust, interpretable starting point. A neural network is more flexible but needs more data and care.
  • Negatives are usually unlabelled ground, not proven barren ground.
  • Validate with spatial blocks. Random splits overstate skill.
  • Check importance rankings against geology and use the output to prioritise fieldwork, not replace it.

Sources

About Orex — Orex is a mineral exploration intelligence platform based in Mwanza, Tanzania. We combine satellite remote sensing, elevation-derived structural analysis and open geoscience data to help explorers, licence holders and investors focus their fieldwork on the ground that matters.

Want to see fault structures and intersection targets on your area of interest — for free? Install GoldRadar on your phone or desktop: it maps lineaments and automatically flags fault intersections derived from satellite elevation data, giving you a structural framework for preliminary exploration before you spend a dollar on the ground.

Related Articles