autoRA: An Algorithm For Automatic Delineation Of Reference Areas To Support Optimized Soil Sampling.
Digital Soil Mapping; Predictive Modeling; Sampling Strategies.
The autoRA algorithm is a data-driven methodology designed to delineate Reference Areas (RAs) that capture critical soil-forming factors, enhancing digital soil mapping (DSM) workflows. By leveraging Gower’s Dissimilarity Index, autoRA systematically determines optimal target area size and spatial resolution configurations, balancing predictive performance and cost-effectiveness. To evaluate its efficacy, autoRA was tested in three distinct scenarios.
In the first scenario, conducted in Rio de Janeiro and Florida, we examined the effect of spatial resolution in calculating Gower’s Index and the impact of target area size on model performance. The predictive model was developed to estimate a Simulated Theoretical Surface (STS) using a single iteration. The optimal Reference Area Model (RAM) with a 50% target area and a 10-pixel block size achieved Euclidean Distance (ED) values (0.15 in Rio de Janeiro and 0.38 in Florida), closely approximating results from exhaustive sampling while reducing costs by approximately 61% and 63%, respectively.
In the second scenario, autoRA was applied to an already mapped soil unit classification in Sátiro Dias, Bahia, Brazil. As this study had an RA manually delineated by a specialist and soil samples collected according to a complete soil profile framework, autoRA was applied using the same covariates as the specialist. We tested RAs at 10%, 20%, 30%, 40%, and 50%, intersecting the proposed RAs with the actual sample locations. The inner samples were used for model training, while the outer samples validated the extrapolation of predictions. At 40% RA coverage, the prediction error using autoRA was lower than that of manually delineated RA. Additionally, manual and autoRA-generated RA soil class maps were compared against 100 iterations of Random Forest modeling using the conventional DSM approach, where datasets were randomly split (70% training, 30% validation). This iterative validation confirmed the robustness of autoRA’s predictions.
The third scenario explored different proportions of pixels classified as low or high Gower’s Dissimilarity. The same inner-training, outer-validation approach was applied. The results demonstrated that focusing on high-dissimilarity regions allowed for the reduction of sampled areas while maintaining predictive accuracy. The Total Area (TA) approach was also tested with sample sizes ranging from 100 to 1,000. The TA model developed using 100 iterations of 70%-30% splits showed that reducing the sample size resulted in a 20% ED improvement compared to the 1,000-sample benchmark. The 800-sample dataset was identified as the optimal benchmark for TA modeling. To test the effect of mosaicking high- and low-dissimilarity areas, we maintained the 20% ED improvement threshold, using 800 as a new benchmark. The results demonstrated that a 600-sample dataset focused within the 40% RA delineated by autoRA produced ED metrics comparable to the TA-STS predicted model.
Sensitivity analysis using STS simulations allowed for extensive testing of parameter configurations, ultimately confirming that the most effective autoRA settings prioritize high-dissimilarity regions. As important as determining the number of points for an accurate prediction is deciding where to place them. The autoRA approach answers this question by simulating multiple configurations, providing expectations of accuracy and cost while honoring the knowledge of landscape dissimilarity and avoiding redundant sampling.