Skip to content

Feature: Basic Reliability Model Implementation for Performance Model Classes - #833

Merged
johnjasa merged 96 commits into
NatLabRockies:developfrom
RHammond2:feature/reliability
Sep 30, 2026
Merged

johnjasa merged 96 commits into
NatLabRockies:developfrom
RHammond2:feature/reliability

Conversation

@RHammond2

@RHammond2 RHammond2 commented Aug 11, 2026 •

Copy link
Copy Markdown
Collaborator

Integration of Basic Reliability Modeling

This is a timestep and simulation duration agnostic, WOMBAT-lite reliability model to calculate the availability of a system or collection of systems.

The PerformanceReliability class enables users to create iid models with separate distributions to simulate failure and maintenance related downtime events using either a Weibull distribution or fixed interval distribution

Features:

  • Standalone, non-OpenMDAO class that can be integrated as needed to existing performance models with minimal modification.
  • Ability to generate an arbitrary number of component-level downtime modes
  • WeibullReliability to sample a Weibull distribution for time to next downtime events like WOMBAT's failure model
  • FixedIntervalReliability to create evenly spaced downtime events like WOMBAT's maintenance model
  • Using "fractional" or "minimum" availability_type the system-level availability can be calculated as collection of systems (e.g., wind farm) modeled as components taking the average availability across components or a single system (e.g., 1 natural gas turbine) taking the minimum availability across components.
  • Vary the downtime length using a uniform distribution (UniformDowntime), lognormal distribution (LogNormalDowntime), or fixed duration (FixedDowntime) model.
  • Using burn_in control when in the lifetime of a system the events are taken from instead of repeatedly simulating the first year of operations.
  • Repeatable results with a fixed, large random seed.
  • Easy to subclass parent classes BaseDowntime and BaseReliability for creating new distributions for event frequency and durations.

Working example

from h2integrate.reliability import PerformanceReliability

config = {
    "simulation": {"dt": 3600, "n_timesteps": 8760},
    "use_reliability": True,
    "burn_in": 6.5,
    "availability_type": "fractional",
    "failure_model": "WeibullReliability",
    "maintenance_model": "FixedIntervalReliability",
    "failure_parameters": {
        "scale": 0.5,
        "shape": 1,
        "n_components": 3,
        "downtime": {
            "model": "FixedDowntime",
            "hours": 5,
            "n_components": 3,
        },
    },
    "maintenance_parameters":{
        "frequency": [0.25, 1, 4],
        "downtime": {
            "model": "UniformDowntime",
            "min_hours": 2,
            "max_hours": 5,
            "n_components": 3,
        },
    },
}

# we can also do the same thing for maintenance, the whole operation just hasn't been combined yet
reliability = PerformanceReliability.from_dict(config)
reliability.run()

# energy-based availability without ramping of downtime events
print(f"Failure-based availability: {reliability.failures.system_availability.sum() / failures.system_availability.size:.2%}")
print(f"Maintenance-based availability: {reliability.maintenance.system_availability.sum() / maintenance.system_availability.size:.2%}")
print(f"System-level availability: {reliability.availability.sum() / reliability.availability.size:.4%}")

Section 1: Type of Contribution

  • Feature Enhancement
    • Framework
    • New Model
    • Updated Model
    • Tools/Utilities
    • Other (please describe):
  • Bug Fix
  • Documentation Update
  • CI Changes
  • Other (please describe):

Section 2: Draft PR Checklist

  • Open draft PR
  • Describe the feature that will be added
  • Fill out TODO list steps
  • Describe requested feedback from reviewers on draft PR
  • Complete Section 8: New Model Checklist (if applicable)

TODO: (see other feedback/considerations for what I'm considering or other aspects I could be missing)

  • Get the model working
  • Burn in period (skipping forward x number of years) - suggestion from @cfrontin
  • Componentizable approach (matrix of failures vs single model)
  • Maintenance (fixed interval) and Failures (random)
    • Integrate failure and maintenance into overarching system availability
  • Enable randomized downtime
  • Tests
  • Better docs
  • Rethink approach to base models
  • Remove setup steps from individual models for easy addition of new models
  • Ramping up/down production from 0-1/1-0f
    • Will become a separate feature in a new PR
  • create array gt/ge validators outside of class
  • n_timesteps as an input to the model in place of using module level N_TIMESTEPS = 8760.
  • dt as an input to the model for duration of the simulation in place of N_TIMESTEPS = 8760.
  • allow user to input use_reliability for a performance model to toggle its usage or check that a model has been provided if using during model initialization.
  • Potentially convert hours to use a dt basis naming/input scheme

Type of Reviewer Feedback Requested (on Draft PR)

Structural feedback: Anything is welcome

Implementation feedback: Should the reliability slot into the performance in a more streamlined way? Any other feedback is welcome.

Other feedback/considerations for the finalized PR: It would be great to get feedback on the importance of the following items and any preferred approaches.

  • better control over random seeding - fixed unless there is a good reason to provide user-level control
  • uncertainty quantification - getting well ahead of ourselves
  • ramping before/after downtime events - better follow-on feature
  • passing n_timesteps from the plant configuration
  • Calculate availability at initialization or manually run model.calculate_availability()?

All of the above final considerations were implemented.

Section 3: General PR Checklist

  • PR description thoroughly describes the new feature, bug fix, etc.
  • Added tests for new functionality or bug fixes
  • Tests pass (If not, and this is expected, please elaborate in the Section 6: Test Results)
  • Documentation
    • Docstrings are up-to-date
    • Related docs/ files are up-to-date, or added when necessary
    • Documentation has been rebuilt successfully
    • Examples have been updated (if applicable)
  • CHANGELOG.md
    • At least one complete sentence has been provided to describe the changes made in this PR
    • After the above, a hyperlink has been provided to the PR using the following format:
      "A complete thought. [PR XYZ]((https://github.com/NatLabRockies/H2Integrate/pull/XYZ)", where
      XYZ should be replaced with the actual number.

Section 4: Related Issues

Section 5: Impacted Areas of the Software

Section 5.1: New Files

  • h2integrate/reliability/utilities.py
    • update_dimensions: Check n_components and any passed arguments to either broadcast the input arrays or update n_components for consistency within a model.
    • calculate_simulation_years: Calculates the length of the simulation period, in years based on dt and n_timesteps.
    • calculate_annual_timesteps: Calculates the rounded up number of timesteps in a year based on dt.
    • calculate_hourly_timesteps: Calculates the rounded up number of timesteps in an hour based on dt.
  • h2integrate/reliability/models.py
    • create_reliability_model: Match-case statement to initialize a new reliability model for PerformanceReliability
    • create_downtime_model: Match-case statement to initialize a new downtime model for BaseReliability
    • SimulationConfig: Standardizes the simulation parameters for reuse and computes additional values to handle varying timestep lengths and number
    • BaseDowntime: abstract base class providing the base attributes and required setup for implemented downtime models.
    • BaseReliability: abstract base class providing the base attributes and required setup for implemented reliability models.
    • PerformanceReliability: WOMBAT-lite style reliability to model both failure and maintenance type events.
    • FixedDowntime: Fixed inteval downtime duration model.
    • UniformDowntime: Uniform distribution-based downtime duration model.
    • LogNormalDowntime: Lognormal distribution-based downtime duration model.
    • WeibullReliability: Weibull distribution-based time to next downtime event model.
    • FixedIntervalReliability: Fixed interval time to next downtime event model.

Section 5.2: Modified Files

  • h2integrate/converters/natural_gas/natural_gas_cc_ct.py
    • NaturalGasPerformanceModel: Adds use_reliability (default False) and reliability_model (default None) to optionally apply the PerformanceReliability.availability to the natural_gas_demand as a limiting factor for production. The availability was not applied to the production because demand is the usage of natural gas to determine the amount of energy produced, so we limit the operational capacity of the natural gas power plant with minimal modification of the actual model.

Section 6: Additional Supporting Information

Only the natural gas model has an active implementation of the reliability modeling to limit the amount of modified code in a single PR. Any further modeling implementations can be made as follow-on PRs where more consideration can be made for the intricacies of the performance models.

Section 7: Test Results, if applicable

Tests pass.

Section 8 (Optional): New Model Checklist

  • Model Structure:
    • Follows established naming conventions outlined in docs/developer_guide/coding_guidelines.md
    • Used attrs class to define the Config to load in attributes for the model
      • If applicable: inherit from BaseConfig or CostModelBaseConfig
    • Added: initialize() method, setup() method, compute() method
      • If applicable: inherit from CostModelBaseClass
  • Integration: Model has been properly integrated into H2Integrate
    • Add the new model to the appropriate __init__.py file to ensure it is properly imported and used in supported_models.py
    • Added to supported_models.py
    • If a new commodity_type is added, update create_financial_model in h2integrate_model.py
  • Tests: Unit tests have been added for the new model
    • Pytest-style unit tests
    • Unit tests are in a "test" folder within the folder a new model was added to
    • If applicable add integration tests
  • Example: If applicable, a working example demonstrating the new model has been created
    • Input file comments
    • Run file comments
    • Example has been tested and runs successfully in test_all_examples.py
  • Documentation:
    • Write docstrings using the Google style
    • Model added to the main models list in docs/user_guide/model_overview.md
      • Model documentation page added to the appropriate docs/ section
      • <model_name>.md is added to the _toc.yml
    • Run generate_class_hierarchy.py to update the class hierarchy diagram in docs/developer_guide/class_structure.md

@elenya-grant
elenya-grant self-requested a review August 11, 2026 19:49
Comment thread h2integrate/reliability/models.py Outdated

@elenya-grant elenya-grant left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left some big-picture comments! Happy to chat through anything if you'd like! Thanks!

Comment thread h2integrate/reliability/models.py Outdated
Comment thread h2integrate/reliability/models.py Outdated
Comment thread h2integrate/reliability/models.py Outdated
Comment thread h2integrate/reliability/models.py Outdated
Comment thread h2integrate/reliability/models.py Outdated
to a quarterly downtime event and 0.25 is equivalent to an every 4 years downtime event.
For all events the timing of the first event will be sampled within the first year or
interval period to offset events from being based on January 1st in an 8760.
downtime (int | float | dict): Either fixed length of each downtime, in hours, or a

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

could we name this downtime_hrs or something? Also - can these models be updated to handle varying timesteps? It seems like a lot of logic is intended for hourly? If so - I think we should should make dt an input parameter too.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could be tricky, but I'll add it to the to do section.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The downtime durations are now individual models, with dt and n_timesteps fully incorporated.

merge_shared_inputs(self.options["tech_config"]["model_inputs"], "performance"),
additional_cls_name=self.__class__.__name__,
)
if self.options["tech_config"]["model_inputs"]["reliability"]:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need to use self.options["tech_config"]["model_inputs"].get("reliability", False) here?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not exactly, but that did get me thinking on a better way to use use_reliability, so I appreciate the question!

Comment on lines +71 to +72
self.reliability_model = None
self.use_reliability = False

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could these be parameters in the NaturalGasPerformanceConfig? So a user can input reliability_model and use_reliability? Where the __attrs_post_init__ checks that reliability_model is provided if use_reliability is True?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good question, it didn't dawn on me that a user could want to provide a definition, but not use it. Though it could be easier to iterate on a problem by simply turning it on/off instead of commenting out a whole section of the inputs. I'll also add this to the to do section.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I forgot to comment, but this is now included with control over the results application being provided in the performance model itself.

@RHammond2
RHammond2 marked this pull request as ready for review September 25, 2026 00:51
@RHammond2
RHammond2 requested a review from johnjasa September 25, 2026 00:52
@RHammond2 RHammond2 added enhancement New feature or request ready for review This PR is ready for input from folks and removed in progress labels Sep 25, 2026
@RHammond2
RHammond2 requested review from johnjasa and removed request for johnjasa September 25, 2026 17:47
@RHammond2 RHammond2 changed the title Reliability Prototype Feature: Basic Reliability Model Implementation for Performance Model Classes Sep 25, 2026
@RHammond2
RHammond2 requested review from johnjasa and removed request for johnjasa September 29, 2026 15:05
@johnjasa
johnjasa requested a review from kbrunik September 29, 2026 19:39

@johnjasa johnjasa left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this, Rob! I've pushed up small-ish changes directly, one about actually assigning the demand value within the NG component, the other changing the sampling spacing. Take a look and feel free to undo anything there.

Then I've left two comments that may elucidate changes for this PR, follow-on PRs, or never. Regardless, I'm approving so we can get this in by tomorrow!

Comment thread h2integrate/reliability/models.py Outdated


# generated from np.random.SeedSequence().entropy
rng = np.random.default_rng(279299947538423226929715083173412195503)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh wait, I do have a not-a-joke comment/question about this, though!

The module-level rng is shared across every model in the process, and PerformanceReliability.run() resamples on every compute(). As a result, availability changes between OpenMDAO iterations: two back-to-back run() calls gave mean availability 1.0 and then 0.9945. That could introduce noise when running an optimization with finite differences that might be tough to track down. Is there a different way to set this rng seed, or am I misunderstanding?

Comment on lines +565 to +568
return np.ceil(
rng.lognormal(self.mean, self.sigma, size=(self.mean.shape[0], 100))
/ self.simulation.n_timesteps_in_hour
).astype(int)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I'm reading the numpy docs correctly, this mean and sigma value are for the underlying normal distribution and not the literal distribution (i.e. hours): https://numpy.org/doc/2.2/reference/random/generated/numpy.random.Generator.lognormal.html#numpy.random.Generator.lognormal

Does that change how we should write this expression here? i.e. do we need to resample in some way? Maybe you already accounted for that but I thought to check!

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, I didn't match the docstrings to the NumPy docstrings, I'll update it.

@johnjasa
johnjasa enabled auto-merge September 30, 2026 15:30
@johnjasa
johnjasa merged commit c8d0e33 into NatLabRockies:develop Sep 30, 2026
12 checks passed
@RHammond2
RHammond2 deleted the feature/reliability branch September 30, 2026 16:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request ready for review This PR is ready for input from folks

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants