Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Storm Prediction: Modeling U.S. Property Damage from Severe Weather

This is a machine learning project modeling structural property damage from severe weather events using 76 years of historical NOAA storm records (1950–2025). The project uses a Generalized Linear Model framework (Tweedie Regressor) benchmarked against Ordinary Least Squares and a Decision Tree Regressor to handle a target variable that is both heavily zero inflated and extremely right skewed. It covers large scale data aggregation, data cleaning, exploratory data analysis, feature engineering, model comparison, and coefficient based feature importance analysis. All datasets, the full notebook, and a fully presentable, rendered final report were graded and recieved an A letter grade.

Table of Contents

Why Did I Build This?

I built this as a final project for an upper division course (PSTAT 100: Data Science Concepts and Analysis) during my time as a Statistics and Data Science major at UCSB, alongside my my friend and fellow student, Ivan Gladnikov. The inspiration was personal rather than academic. Severe storms during the 2026 Winter Quarter caused destructive localized flooding across Santa Barbara County, requiring full structural remodels of residential properties in our college town of Isla Vista. This included our own homes, leading to full remodling due to water damage, leaks, and mold. That experience motivated us to ask a data driven question: can objective storm characteristics actually predict the scale of financial property damage, and which meteorological events are the biggest drivers of it?

Results

Model Mean Absolute Error (MAE) Minimum Prediction Tweedie Deviance Deviance Explained
Baseline (Dummy) $610,823.45 Global Mean 4,251.76 0.00%
OLS Linear Regression $725,871.56 -$4,301,743.12 N/A N/A
Decision Tree $487,052.70 $0.00 N/A N/A
Tweedie GLM $657,258.62 $0.00 2,102.76 50.54%

Key Finding

  • The Tweedie GLM (power = 1.5, log link) was model most mathematically appropriate for the data's structure, correctly bounding predictions at $0 while explaining 50.54% of the Tweedie deviance
  • OLS Linear Regression mathematically collapsed, producing an impossible minimum prediction of -$4,301,743.12 and underperforming even the naive mean baseline
  • The Decision Tree achieved the lowest raw MAE ($487,052.70) but, as a "black box" model, offered no interpretable coefficients for answering the actual research question
  • Coastal and tropical events dominate the damage hierarchy: Hurricanes carry a 2,525x damage multiplier over baseline, followed by Storm Surge/Tide (1,056x), Tsunami (144x), Tropical Storm (126x), and Coastal Flood (119x)
  • Approximately 75% of all recorded storm events resulted in exactly $0 of property damage, which was the central statistical challenge the entire modeling approach was built around

Outline & Analysis

For the full Jupyter Notebook, see FinalReportNotebook. For a rendered, readable version, see Rendered Report. A brief outline is below:

Part 1: Abstract & Introduction

  • Research question: can specific storm attributes predict financial property damage, and which meteorological events drive it most?
  • Motivated by real, localized flood damage in Isla Vista, CA during the 2026 Winter Quarter
  • Dataset: 76 years of NOAA Storm Events Database records (NCEI, 2026), combined into 2,012,975 rows across 51 original columns

Part 2: Methodology

  • Standard OLS assumes normally distributed residuals, which fails outright against a target that is ~75% zero and extremely right-skewed for the remainder
  • Dropping zero-damage events to "fix" this would introduce severe selection bias, so a Generalized Linear Model with a Tweedie distribution and log link was used instead
  • Power parameter set to $p = 1.5$, corresponding to a Compound Poisson-Gamma distribution (Rahim, 2020), designed to handle an exact probability mass at zero alongside a continuous, long right tail
  • Fit using the newton-cholesky solver for memory efficiency and fast convergence at scale (2M+ rows)
  • Benchmarked against OLS (parametric baseline) and a Decision Tree Regressor (nonparametric alternative)

Part 3: Data

  • Raw data: separate annual CSVs (events, fatalities, locations) for each year from 1950–2025, merged into one unified dataframe
  • Acknowledged reporting bias: 1950–1954 captured tornadoes only, 1955–1995 added thunderstorm wind and hail, and 1996–present captures all 48 standardized event types via modern radar — meaning some of the rise in recorded events over time reflects better observation technology, not just more storms
  • DAMAGE_PROPERTY was originally stored as alphanumeric strings (e.g., '10.00K', '15.00M'); a custom parser converted these to numeric floats, with missing values imputed as $0.0 per NOAA's own recording convention for non-damaging events
  • Engineered DURATION_HOURS, DECADE, and SEASON features from unified datetime fields
  • Final feature set narrowed from 59 engineered columns down to 7 predictors (STATE, EVENT_TYPE, SEASON, FLOOD_CAUSE, DURATION_HOURS, DECADE) after dropping highly-missing columns (e.g., tornado F-scale, 96%+ missing) and administrative identifiers
  • Target variable (DAMAGE_PROPERTY_NUM) was log-transformed (LOG_DAMAGE) purely for EDA visualization, justifying the choice of a log-link GLM for modeling

Part 4: Results & Feature Importance

  • Tweedie GLM explained 50.54% of deviance with a valid $0.00 minimum prediction and fully interpretable, exponentiated coefficients
  • Top 10 drivers of property damage (exponentiated coefficients): Hurricane (Typhoon) 2,525.34x, Storm Surge/Tide 1,056.20x, Tsunami 143.67x, Tropical Storm 126.00x, Coastal Flood 119.36x, Tornado 60.37x, Ice Storm 59.05x, Wildfire 47.04x, Flood 36.07x, Lakeshore Flood 24.73x
  • A DummyRegressor baseline (predicts the global mean for every event) was used to mathematically confirm the Tweedie GLM's real predictive value over naive guessing
  • Standard metrics like $R^2$ and RMSE were explicitly rejected as inappropriate for this distribution; Tweedie Deviance Explained and MAE were used instead, since RMSE would be distorted by the extreme right tail of billion-dollar storms

Part 5: Conclusions & Limitations

  • Standard linear regression is fundamentally inadequate for zero-inflated, highly skewed financial damage data; a Tweedie GLM is the mathematically appropriate tool
  • Coastal and tropical events are confirmed as the most destructive meteorological forces to structural property, directly contextualizing the Isla Vista flooding that motivated the project
  • Historical reporting bias limits conclusions to associative trends over time rather than strict causal claims
  • Future work: incorporate geospatial parameters (e.g., distance to coast, county-level zoning) to improve variance explained and provide more actionable insight for real estate and disaster preparedness in high-risk coastal states

Rendered Report

For the full rendered HTML version with all code, output, and visualizations, click below (opens via htmlpreview.github.io since GitHub doesn't render raw HTML files directly):

Dependencies

import pandas as pd
import numpy as np
import os
from google.colab import drive
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.model_selection import train_test_split
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.linear_model import TweedieRegressor, LinearRegression
from sklearn.tree import DecisionTreeRegressor
from sklearn.dummy import DummyRegressor
from sklearn.metrics import mean_absolute_error, mean_tweedie_deviance

License

© 2026 Ryan Fabrick & Ivan Gladnikov. All rights reserved. This project may not be reused, adapted, or submitted for academic coursework without explicit written permission from the authors.

Authors

Ryan Fabrick

Ivan Gladnikov

Acknowledgements & References


Built with ❤️ for UCSB

This project demonstrates my interest in machine learning, applied data science, and predictive modeling. It was completed as a final exam for an upper division undergraduate course in Data Science Concepts and Analysis (PSTAT 100), taught by Professor John Inston. This recieved an A, letter grade.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages