Fire Prediction Montesinho Natural Park - ML Classification Model Weekly 2020 - 03/2025
This repository contains the machine learning classification part of the project. In this part, we focus on predicting fire weeks in Montesinho Natural Park using weekly environmental and weather-related data.
The datasets cover the period from January 2020 to March 30th, 2025. The project explores the use of different classification models to identify weeks associated with fire occurrence based on environmental and meteorological conditions.
- The collection of data for this project is documented in the repository FireFlow.
- The creation of features, such as lagged and rolling variables, and the creation of the datasets are documented in the repository FireFeatures.
The main objectives of this repository are:
- Explore relationships between variables and identify highly correlated features. Simplify the datasets accordingly.
- Address class imbalance in the fire/no-fire classification problem.
- Develop and compare several machine learning classification models.
- Evaluate model performance using appropriate classification metrics.
The project evaluates several classification algorithms, including:
- Decision Tree
- Random Forest
- Logistic Regression
- XGBoost
Different approaches to handling class imbalance are explored, including oversampling techniques.
The dataset contains environmental and weather-related variables, together with lagged and rolling features designed to capture conditions over previous periods.
Feature selection includes correlation analysis and statistical feature selection. Highly correlated features are examined to reduce unnecessary redundancy while retaining relevant information for the classification models. The temporal features month and season are kept in the analysis.
The target variable represents whether a week is classified as a fire week based on the maximum temperature-related variable (max_T21).
The classification threshold used in the final analysis is:
- Fire:
max_T21 ≥ 325 - No fire:
max_T21 < 325
The max_T21 variables used to define the target are excluded from the predictor variables to avoid target leakage.
The models are evaluated using cross-validation and a separate test dataset covering 01/01/2024 – 30/03/2025.
Performance is assessed using metrics including:
- Precision
- Recall
- F1-score
- Balanced accuracy
- Confusion matrix
- Accuracy
Particular attention is given to the performance of the fire class, as correctly identifying fire weeks is an important aspect of the classification problem.
FinalClassifModel-fv.ipynb— final notebook containing the classification analysis and model evaluation.- The
datafolder contains the datasetsmontesinho_week_20_23.csvandmontesinho_week_24_25.csvused for the analysis. .gitignore— files and folders excluded from Git tracking.
This part of the project was developed in Python using libraries including:
- pandas
- NumPy
- scikit-learn
- imbalanced-learn
- XGBoost
- matplotlib
- seaborn
The analysis uses weekly data from Montesinho Natural Park covering the period from 01/01/2020 to 30/03/2025.
The training data is contained in montesinho_week_20_23.csv, while the test data is contained in montesinho_week_24_25.csv.
The datasets were created using the FireFeatures repository and datasets shared through Zenodo: https://doi.org/10.5281/zenodo.17610275
The data include environmental and weather-related variables used to develop predictors for the fire classification models.
The main analysis is available in:
FinalClassifModel-fv.ipynb
The notebook contains the data preparation, exploratory analysis, feature selection, model development, cross-validation, and final model evaluation.
AI Disclosure: Parts of the code in this repository were written with the assistance of ChatGPT, followed by manual review and testing.