Description
Statistical and Machine Learning Engine for Long-Term Natural Resource Management Data.
Description
A comprehensive toolkit for statistical and machine learning-based analysis of long-term Natural Resource Management (NRM) datasets. Integrates formula-driven approaches, statistical inference, and machine learning (ML) models for advanced analytics. Modules cover trend and structural analysis (Mann-Kendall test, slope estimation, Chow test, structural break detection), multivariate system modelling (Partial Least Squares (PLS), Structural Equation Modelling (SEM)), response curve optimisation, time-series forecasting (Autoregressive Integrated Moving Average (ARIMA), hybrid models), panel data and treatment effects (Difference-in-Differences (DiD), causal machine learning), uncertainty and sensitivity analysis (bootstrap, Monte Carlo, Bayesian), and automated model selection and performance comparison. Designed for long-term datasets covering soil, water, crop, and climate domains. Key references: Mann and Kendall (1945) <doi:10.2307/1907187>; Sen (1968) <doi:10.1080/01621459.1968.10480934>; Bai and Perron (2003) <doi:10.1002/jae.659>; Rosseel (2012) <doi:10.18637/jss.v048.i02>; Croissant and Millo (2008) <doi:10.18637/jss.v027.i02>.
README.md
NRMstatsML
NRMstatsML is a comprehensive R package providing a statistical and machine learning engine for long-term Natural Resource Management (NRM) datasets. It integrates formula-driven approaches, statistical inference, and machine learning for reproducible analytics across soil, water, crop, and climate domains.
Installation
# Install from CRAN
install.packages("NRMstatsML")
Core Modules
| Module | Key Functions | Description |
|---|---|---|
| trendML | nrm_trend(), nrm_mann_kendall(), nrm_sens_slope(), nrm_structural_break() | Monotonic trend tests, slope estimation, structural break detection |
| multiSysML | nrm_multivariate(), nrm_pls(), nrm_sem() | OLS, PLS, Structural Equation Modelling |
| responseML | nrm_response_curve(), nrm_optimize_input() | Yield-response curves and input optimisation |
| tsML | nrm_forecast(), nrm_arima() | ARIMA/SARIMA forecasting with prediction intervals |
| panelML | nrm_panel(), nrm_did() | Fixed/random effects, difference-in-differences |
| uncertaintyML | nrm_uncertainty(), nrm_bootstrap(), nrm_monte_carlo() | Bootstrap, Monte Carlo, Bayesian uncertainty |
| autoML | nrm_automl(), nrm_benchmark() | Automated model selection and benchmarking |
Quick Start
library(NRMstatsML)
# Load synthetic example data
data(nrm_example)
# 1. Validate data
nrm_data_check(nrm_example)
# 2. Trend analysis
trend <- nrm_trend(nrm_example, time_var = "year", value_var = "crop_yield")
print(trend)
nrm_plot(trend)
# 3. Yield-response curve
rc <- nrm_response_curve(nrm_example, input_var = "N",
response_var = "crop_yield", type = "quadratic")
opt <- nrm_optimize_input(rc, price_ratio = 0.6)
print(opt)
# 4. Forecast next 5 years
fc <- nrm_forecast(nrm_example, value_var = "crop_yield", horizon = 5)
nrm_plot(fc)
# 5. Bootstrap uncertainty of mean yield
bs <- nrm_bootstrap(nrm_example,
stat_fn = function(d) mean(d$crop_yield),
n_iter = 1000)
print(bs)
Design Principles
- Consistent API: every module follows
nrm_<verb>()naming. - CRAN compliant: no global state, explicit imports, documented datasets.
- Reproducible: set
seedarguments for deterministic results. - Modular: use individual module functions or the high-level wrappers.
- ggplot2-based visualisation: all
nrm_plot()methods return editableggplotobjects.
Citation
citation("NRMstatsML")
License
GPL (>= 3). See the GNU General Public License for details.