Quantifying natural-hazard and climate risk and turning it into financial loss numbers.
I'm a Statistics graduate moving into catastrophe and climate risk modeling. I'm currently doing an MSc in Atmospheric Sciences at NIT Rourkela, which gives me the hazard-science side to pair with the statistics and economics I already work in.
The short version of what I'm after: a meteorologist can model the cyclone but can't price the loss; a quant can price the loss but can't model the cyclone. I'm building toward doing both - the hazard × exposure × vulnerability → loss chain that insurers, reinsurers, and climate-risk teams actually use.
- Statistics & extreme-value theory - my core. Risk is a tail-probability problem, so this is the part I lean on most.
- Economics & finance - my minor. Turning a hazard into a loss number, a premium, a capital charge.
- Hazard science - from the MSc (tropical cyclones, floods, extreme heat, climate scenarios).
- Python / geospatial / SQL / ML - from coursework, a research internship in optimization, and the catastrophe modeling work below.
These are deliberately split the way a catastrophe risk team is split - a portfolio risk view, exposure data management, and event response. Each one consumes the one before it, and they cross-reference each other, including where one found a defect in another.
odisha-cyclone-risk - end-to-end catastrophe risk model for Bay of Bengal tropical cyclones over the coastal Odisha belt. Full hazard × exposure × vulnerability → loss chain in CLIMADA: 2,754-event stochastic catalogue validated against Cyclone Fani and Phailin, LitPop exposure, an OSDMA-derived vulnerability curve, OEP/AEP curves from a 100,000-year Year Loss Table, and CAT XL layer pricing.
Headline finding: vulnerability specification drives an ~11.4× spread in average annual loss - larger than climate intensification and exposure uncertainty combined - and propagates directly into reinsurance pricing. GPD tail extrapolation was tested and rejected as unsupported by the data.
odisha-exposure-quality - the insured portfolio the risk view runs on, and what data-quality defects do to it. A 5,000-location synthetic Odisha portfolio with 10,001 controlled errors injected across 39 rules, a PostgreSQL + PostGIS detection engine evaluated against ground truth, and a materiality framework that separates "a detector fired" from "this changes the answer".
99.81% recall at 90.46% precision across 11,035 flags. Accumulation eligibility moved portfolio TIV from ₹1.465T to ₹145.7B and linked AAL from ₹15.84B to ₹1.81B, while a single unresolved location anomaly raised spatial concentration (HHI) from 0.224 to 0.728. Every treatment is FIX / REFER / QUARANTINE / ASSUME with an audit trail, never silent imputation.
odisha-event-response - the same risk view run forward in time under forecast uncertainty. What did the model support at 72, 48, 24 and 12 hours before landfall? 500-member perturbed-track ensembles for four historical cyclones, 9,500 CLIMADA wind fields, with the along-track and cross-track error split derived from IMD's published position and landfall statistics rather than assumed.
Headline finding: attachment probability rose monotonically for the one cyclone that reached the portfolio and fell monotonically for the three that missed, across all 19 panels, with no reversal. The decision-relevant signal is the slope against lead time, not the level at any single forecast issue. 72 hours of warning with zero false alarms across a wide band of notification thresholds, while the P75 reserve peaked at 3.68× the realised loss.
It also found a defect in its own ensemble generator, and one in the project before it. Ensemble tracks ended at landfall, so a member displaced backward along track had every position offshore and reported exactly zero loss for a reason that had nothing to do with the storm - indistinguishable, in a loss file, from a genuine miss. I found it by reconstructing the latent random draws and testing the sign balance among zero-loss members: 13.7 standard deviations one-sided. Fixing it changed every number downstream. The before-and-after diagnostic is committed alongside the result.
lasalgaon-onion-dss - a decision-support system modeling price-crash risk for a commodity market, combining statistics with economic reasoning. Closest in spirit to the tail-risk / loss modeling above.
uidai-operational-dashboard - an operational analytics dashboard modeling district-level stress from real administrative data.
(Other repositories include a research internship in combinatorial optimization - Python + Gurobi - where the focus was rigorous, honestly-validated results.)
I'd rather be honest about what's finished and what isn't:
- Contemporaneous forecast skill. The event-response work runs all four cyclones on 2020-2024 IMD skill so they're comparable, but three of them predate that window. Phailin's 24h landfall error in 2013 was 3.5× today's. Rerunning each storm under the skill of its own era asks a different and better question: what would a team actually have faced at the time?
- Coast-following exposure. The synthetic portfolio's locations sit inside a bounding box rather than along the coastline, which the event-response project exposed. Regenerating it along the coast is the single change that would most alter those conclusions.
- OasisLMF - the open catastrophe-modeling framework, as a complement to CLIMADA.
- ML for Earth - applying machine learning to hazard problems, including the current generation of ML weather models.
Python (pandas, numpy, scipy, scikit-learn, xgboost) · CLIMADA · statistics & extreme-value theory · geospatial Python (geopandas, xarray, shapely, pyproj) · SQL / PostgreSQL / PostGIS · IBTrACS, LitPop, Natural Earth · optimization (Gurobi) · learning: OasisLMF · QGIS · PyTorch
Building consistently toward the intersection of climate science and financial risk. Open to conversations, collaborations, and pointers from anyone working in cat modeling or climate risk.